The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
51,110 characters
Local Projection Inference in High Dimensions
\onehalfspacing
\maketitle
\begin{abstract}
In this paper, we estimate impulse responses by local projections in high-dimensional settings. We use the desparsified (de-biased) lasso to estimate the high-dimensional local projections, while leaving the impulse response parameter of interest unpenalized. We establish the uniform asymptotic normality of the proposed estimator under general conditions. Finally, we demonstrate small sample performance through a simulation study and consider two canonical applications in macroeconomic research on monetary policy and government spending.
\bigskip
\noindent \textbf{Keywords:} {Local projections, Impulse response analysis, High-dimensional data, Honest inference, Lasso}
\end{abstract}
\section{Introduction}\label{sec:Introduction}
In this paper, we develop a simple approach for conducting valid inference on impulse responses in high-dimensional settings. We do so by pairing the estimation framework of local projections (LPs) with the inference framework of the desparsified (or debiased) lasso. Since their introduction by \citet{jorda2005estimation}, LPs have been widely used in macroeconomic research for studying the dynamic propagation of shocks through impulse response analysis, see e.g. \citet{angrist2018semiparametric}, \citet{ramey2018government}, \citet{ramey2016macroeconomic} and \citet{stock2018identification}.
As such, they have become
an increasingly used alternative to structural Vector AutoRegressions (SVAR) pioneered by \citet{sims1980macroeconomics},
see for instance \citet{ramey2016macroeconomic} or \citet{kilian2017structural} for an overview on this work.
While the methods have different finite-sample properties, SVARs and LPs are equivalent in population, as established by \cite{plagborg2021local} who show that the underlying impulse response estimands are the same for both. This implies that the well-documented necessity for SVARs to add assumptions in order to achieve structural identification of the impulse responses, is equally true for LPs. Indeed, \cite{plagborg2021local} show that the identification strategies typically used for SVARs have an equivalent implementation for LPs, and vice versa. While identification issues therefore do not affect the choice between SVAR or LP, finite-sample considerations about estimation and inference problem do. LP impulse responses are obtained by estimating only univariate linear regressions, and performing standard inference on (typically) a single parameter of interest across these univariate regressions. In contrast, SVARs require estimating the whole system of equations, and transforming these into the Vector Moving Average (VMA) representation from which the impulse responses can be derived. While the system estimation of the SVAR can still be done linearly, the inversion to the VMA representation is a nonlinear operation that renders the impulse response coefficients a complex nonlinear functions of \emph{all} VAR parameters. This makes inference considerably more cumbersome, even in
low-dimensional settings, where the complications and inaccuracies of applying the Delta method in finite samples have led to bootstrap inference becoming the norm \citep[see e.g. Chapter 12 of][and the references therein]{kilian2017structural}. These problems are exacerbated when the dimensionality of the system grows, making LPs our method of choice to obtain impulse responses in high dimensions.
We consider high-dimensional local projections (HDLPs) in a general time series framework where the number of regressors can grow faster than the sample size. Impulse response analysis quickly becomes high-dimensional in macroeconomic research. Even when considering impulse response analysis with few variables, the number of regressors in LPs or SVARs is often large due to the common practices of including many lags to control for autocorrelation (see e.g., \citealp{bernanke1998measuring}, \citealp{romer2004new} and \citealp{sims2006were}) or to robustify against (near) unit roots via lag augmentation as in \citet{olea2020local}. Similarly, additional regressors are often introduced to capture seasonal patterns in impulse responses (see e.g., quarter-dependent coefficients used in \citealp{blanchard2002empirical}), or to permit nonlinearities to produce state-dependent impulse responses (\citealp{koop1996impulse}; \citealp{ramey2018government}).
Additionally, one might be interested in impulse response analysis with many variables. Motivations of the latter include, amongst others, avoidance of contaminated measurements of policy innovations or informational deficiency with impulse response estimates that are distorted by omitted variable bias, the ability to use several observable measures of difficult-to-measure variables in theoretical models, or opportunities for researchers and policy makers to investigate the impact of shocks on a larger, possibly more disaggregated set of variables they care about (see e.g., \citealp{bernanke2005measuring}, \citealp{forni2014sufficient}, \citealp{stock2016DFM} and \citealp[Ch.~16]{kilian2017structural}).
Regular estimation approaches, however, become infeasible for high-dimensional settings. Traditional approaches to high-dimensionality include modelling commonalities between variables, such as through factor-augmented VARs (FAVAR) (\citealp{bernanke2005measuring}), dynamic factor models (DFM) (\citealp{forni2009opening}; \citealp{stock2016DFM}) or -- for panel structures -- global VARs (\citealp{ChudikPesaran16}); as well as Bayesian shrinkage methods (\citealp{banbura2010large}; \citealp{chan2020large}). Recently, sparse shrinkage methods such as the lasso have gained considerable popularity in econometrics for policy evaluation (\citealp{BCH14}) and (macroeconomic) time series analysis (\citealp{kock2020penalized}). Their use in impulse response analysis however has only scarcely been explored.
While several methods and theoretical results now exist for estimating sparse VAR models -- see e.g., \cite{BasuMichailidis15, KockCallot2015, masini2019regularized} and the references cited therein -- inference on impulse responses is complicated by two issues. First, sparse estimation techniques such as the lasso perform model selection, which induces issues with non-uniformity of limit results if this selection is ignored (\citealp{LeebPoetscher05}). Second, while several methods such as orthogonalization (\citealp{BCH14}) and debiasing, or desparsifying the lasso (\citealp{vandeGeer14}; \citealp{JavanmardMontanari14}) have been proposed to yield uniformly valid (or `honest') post-selection inference, the impulse response parameters are nonlinear functions of all estimated VAR parameters. This severely limits the applicability of existing post-selection inference methods which are typically designed for (relatively) low-dimensional parameters of interest that can be estimated directly. Indeed, to our knowledge, impulse response analysis in sparse HD-SVARs is only considered in \cite{krampe2022}, who construct a complex multi-step algorithm to overcome these complications. Instead, by casting the problem in the LP framework, we reduce the impulse response parameter(s) to a (directly estimable) low-dimensional object in the presence of high-dimensional nuisance parameters, which makes the standard post-selection tools available.
We develop HDLP inference based on the desparsified lasso of \cite{vandeGeer14} and its time series extension in \cite{adamek2021lasso}.
The latter has recently been empirically investigated in the context of LP estimation with instrumental variables (LP-IV) in contemporaneous work by \cite{Karapanagioti21}. Their approach not only differs in the focus on IV models, but they also apply the method of \cite{adamek2021lasso} `as is'. Instead, we tailor the approach of \cite{adamek2021lasso} specifically to high-dimensional LPs consisting of a small number of parameters of interest -- the dynamic response of a variable to a shock at a given horizon -- and many controls. In particular, we modify their approach by leaving the parameter of interest unpenalized during the estimation procedure to ensure it does not suffer from penalization bias. Such a setting is of more general relevance for treatment effect models consisting of a small number of variables whose effects are of interest combined with a large set of controls. We theoretically show that the combination of few unpenalized parameters with many penalized ones does not affect the asymptotic behaviour of the desparsified lasso.
Through simulation experiments, we show that our proposed estimator has considerably better coverage rates than the standard desparsified lasso in finite samples. Furthermore, on an empirically calibrated DFM (\citealp{li2021local}) we find that our proposed HDLP method shows competitive performance in a sparse DFM compared to a low-dimensional LP and a factor-augmented LP.
We consider two canonical macroeconomic applications and demonstrate the performance of the proposed desparsified-lasso based estimator for HDLPs in recovering structural impulse responses. Specifically, we first extend the work of \citet{bernanke2005measuring} on macroeconomic responses to a shock in monetary policy to a HDLP setting.
Second, we consider the work by \citet{ramey2018government} on state-dependent impulse responses to a shock in government spending, as also studied by \cite{Karapanagioti21} in their LP-IV setting, but we extend the original specification used in \citet{ramey2018government} to a HDLP specification with more lags
as well as a more complex state-dependent HDLP with interaction between different state variables.
The methods needed for these applications are implemented in the R package \texttt{desla} (\citealp{desla}), which also offers the methods of \citet{adamek2021lasso}.
The paper is organized as follows. Section \ref{sec:HDLP} introduces the model and our inferential procedure,
as well as our main result on the asymptotic normality of the desparsified lasso.
Section \ref{sec:Simulations} contains simulation studies that examine the small sample performance of the proposed desparsified lasso. Section \ref{sec:empiricalapplications} presents the results of the two macroeconomic applications. Section \ref{sec:conclusion} concludes.
A word on notation. For any $N$ dimensional vector $\boldsymbol{x}$, $\left\Vert \boldsymbol{x}\right\Vert_r=\left(\sum_{i=1}^{N}\left\vert x_i\right\vert^r\right)^{1/r}$
denotes the $l_r$-norm, with the convention that $\left\lVert\boldsymbol{x}\right\rVert_0 = \sum_{i} 1 (\left\lvertx_i\right\rvert>0)$. Depending on the context, $\sim$ denotes equivalence in order of magnitude of sequences, or equivalence in distribution.
\section{High-dimensional Local Projections}\label{sec:HDLP}
Consider the local projection regression
\begin{equation}\label{eq:localprojection}
y_{t+h}=\phi_{h}x_{t}+ \rho_{h}y_t + \boldsymbol{\eta}_h^\prime\boldsymbol{w}_{s,t}+\sum_{k=1}^K \boldsymbol{\delta}_{h,k}^\prime\boldsymbol{z}_{t-k}+u_{h,t}, \quad h = 0, 1, \ldots, h_{\max},
\end{equation}
where $\phi_{h}, \rho_{h}, \boldsymbol{\eta}_h, \boldsymbol{\delta}_{h,k}$ are the projection parameters, $u_{h,t}$ is the projection error and the set of variables $\boldsymbol{z}_t=\big(\boldsymbol{w}_{s,t}^\prime,y_t,x_t, \boldsymbol{w}_{f,t}^\prime\big)^\prime$ consist of the response $y_t$, the shock variable $x_t$, and the vectors of control variables split into the ``slow" variables $\boldsymbol{w}_{s,t}\in\mathds{R}^{n_s}$, and the ``fast" ones $\boldsymbol{w}_{f,t}\in\mathds{R}^{n_f}$ for identification purposes, as discussed below.
\footnote{We omit the constant, since we generally demean the data when using the desparsified lasso.}
Our single parameter of interest is $\phi_{h}$, the dynamic response at horizon $h$ of $y_t$ after an impulse in $x_t$. The LP impulse response function of $y_t$ with respect to $x_t$ is given by $\{ \phi_{h}\}_{h\geq 0}$ and can be obtained by estimating \cref{eq:localprojection} for
$h=0,1,\dots,h_{\max}$.
To permit causal interpretation of the impulse responses, structural assumptions about the data generating process are needed. Example \ref{ex:exampleSVMA} considers a Structural Vector Moving Average model (see Assumption 3 in \citealp{plagborg2021local}) as an example.
\begin{example}[SVMA]\label{ex:exampleSVMA}
\textnormal{Consider the Structural Vector Moving Average (SVMA)
\begin{equation}\label{eq:SVMA}
{\boldsymbol{z}_t}=\boldsymbol{\mu}+
{\boldsymbol{A}(L)}
{\boldsymbol{\epsilon}_t},~~\boldsymbol{A}(L)=\sum_{k=0}^\infty
{\boldsymbol{A}_{k}}L^k,~~\boldsymbol{\epsilon}_{t}\overset{i.i.d.}{\sim}N(\boldsymbol{0},\boldsymbol{I}),
\end{equation}
where $\boldsymbol{\epsilon}_t$ is a vector of structural shocks, $L$ is the lag operator, $\{\boldsymbol{A}_{k}\}_k$ is absolutely summable, and $\boldsymbol{A}(x)$ has full row rank for all complex scalars $x$ on the unit circle. }
\end{example}
While an assumption such as \Cref{ex:exampleSVMA} is necessary to recover the structural impulse responses, it is not a sufficient condition: $\boldsymbol{A}_0$ is not identified without further restrictions.
We focus on impulse responses identified by a (partially) recursive structure. This can be implemented in LPs by partitioning the variables in $\boldsymbol{z}_t$ into slow variables ($\boldsymbol{w}_{s,t}$) and fast variables ($\boldsymbol{w}_{f,t}$). The slow variables are predetermined with respect to $x_t$, while the fast variables can react contemporaneously to the shock variable $x_t$ and therefore only enter the equation with a lag. Equivalently, in the language of recursively identified SVARs, $\boldsymbol{w}_{s,t}$ is ordered above $x_t$, while $\boldsymbol{w}_{s,t}$ is ordered below. Note also that in \cref{eq:localprojection} we assume that $y_t$ is predetermined with respect to (ordered above) $x_t$; if the reverse were true one should remove $y_t$ from the right-hand side of \cref{eq:localprojection}.\footnote{In this situation, the model would be equivalent to equation (1) of \cite{plagborg2021local}.}
\Cref{eq:localprojection} can also be used outside of the recursive identification framework, for example when the structural shock is externally identified. Examples of the latter are the narrative monetary policy shock series of \cite{romer2004new}, or the military spending series of \cite{ramey2018government}. One can then directly take $x_t=\epsilon_{1,t}$ provided the structural shock is observed without measurement error. Given the exogeneity (and thus predeterminedness) of such shocks, all other variables would then be considered to be ``fast'' and no variable would be included contemporaneously; see e.g. \Cref{ex:ramey}.
Section 3.2 of \cite{plagborg2021local} provides detailed discussions on how other identification schemes, such as sign restrictions, can be imposed while estimating \cref{eq:localprojection}; these arguments generally remain valid in our high-dimensional setup as well. However, identification through instrumental variables
as in e.g.~\cite{stock2018identification}
becomes more complicated. Adapting IV estimation methods such as two-stage least-squares (\citealp[Section 3.3]{plagborg2021local}) to high dimensions requires a different setup for the desparsified lasso (see e.g.~\citealp{gold2020inference}), which has not yet been explored for time series models. This remains an interesting avenue for future research.
While the LP in \cref{eq:localprojection} has a single parameter of interest, it can easily be extended to allow for LPs with a small number of parameters of interest, see Example \ref{ex:ramey}.
\begin{example}[State-Dependent LP]\label{ex:ramey}
\textnormal{Consider the nonlinear state-dependent model of \cite{ramey2018government}:
\begin{equation}\label{eq:statedependent}
\small
y_{t+h}= I_{t-1}\left[\alpha_h+\phi_{A,h}x_{t}+\sum_{k=1}^K\boldsymbol{\delta}_{A,h,k}^\prime\boldsymbol{z}_{t-k} \right]
+(1-I_{t-1})\left[\phi_{B,h}x_{t}+\sum_{k=1}^K\boldsymbol{\delta}_{B,h,k}^\prime\boldsymbol{z}_{t-k} \right]+u_{h,t}.
\end{equation}
Here, $\phi_{A,h}$ and $\phi_{B,h}$ are the two responses of $y_t$ after an impulse in $x_t$, corresponding to two different states distinguished by the dummy variable $I_t$.}
\end{example}
To generally accommodate LP settings with a small number of parameters of interest and a large number of controls, we will use the more general notation
\begin{equation}\label{eq:DGP}
{y_t}=\underset{1\times H}{\boldsymbol{x}_{\mathcal{H},t}^\prime}\underset{H\times 1}{\boldsymbol{\beta}_\mathcal{H}}+\underset{1\times(N-H)}{\boldsymbol{x}_{-\mathcal{H},t}^\prime}\underset{(N-H)\times 1}{\boldsymbol{\beta}_{-\mathcal{H}}}+u_t,~~t=1,\dots,T,
\end{equation}
and in matrix form
\begin{equation*}
\boldsymbol{y}=\boldsymbol{X}_\mathcal{H}\boldsymbol{\beta}_\mathcal{H}+\boldsymbol{X}_{-\mathcal{H}}\boldsymbol{\beta}_{-\mathcal{H}}+\boldsymbol{u}.
\end{equation*}
We hereby separate the parameter vector $\boldsymbol{\beta}=[\boldsymbol{\beta}_\mathcal{H}^{\prime},\boldsymbol{\beta}_{-\mathcal{H}}^{\prime}]^{\prime}$ into two groups.
The first group consists of the $H$-dimensional set of parameters of interest $\boldsymbol{\beta}_\mathcal{H}$, corresponding to the variables $\boldsymbol{x}_{\mathcal{H},t}$, which will be unpenalized during the estimation procedure to ensure they are not affected by penalization bias. The second group consists of the large set of parameters $\boldsymbol{\beta}_{-\mathcal{H}}$, corresponding to control variables $\boldsymbol{x}_{-\mathcal{H},t}$, which will be penalized during the estimation procedure. We hereby use $\mathcal{H}$ to denote the set that indexes the $H$ variables of interest. Without loss of generality, we order these variables first in $\boldsymbol{X}=[\boldsymbol{X}_\mathcal{H},\boldsymbol{X}_{-\mathcal{H}}]$. Finally, since each horizon $h$ of the LPs is estimated separately, we suppress the dependence on $h$ in our notation.
\subsection{Local Projection Estimation} \label{subsec:estimator}
We start by estimating \cref{eq:DGP} with the lasso to obtain an initial estimate of $\boldsymbol{\beta}$, while leaving the parameters of interest $\boldsymbol{\beta}_{\mathcal{H}}$ unpenalized.
The first stage lasso is defined as
\begin{equation}\label{eq:problem}
\hat\boldsymbol{\beta}^{(L)}=[\hat{\boldsymbol{\beta}}_\mathcal{H}^{(L)\prime},\hat{\boldsymbol{\beta}}_{-\mathcal{H}}^{(L)\prime}]^\prime=\operatorname*{arg\,min}_{\boldsymbol{\beta}^*\in\mathds{R}^{N}}\left\lVert\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}^*\right\rVert_2^2/T+2\lambda\left\lVert\boldsymbol{W}\boldsymbol{\beta}^*\right\rVert_1,
\end{equation}
where $\boldsymbol{W}$ is an $N\times N$ diagonal matrix with $\boldsymbol{W}_{i,i}=0$ for $i\in \mathcal{H}$ and 1 otherwise.
The Frish-Waugh-Lovell-style result of \cite{yamada2017} ensures that the penalized and unpenalized estimates can be respectively obtained by
\begin{equation*}\begin{split}
\hat{\boldsymbol{\beta}}^{(L)}_{-\mathcal{H}}&=\operatorname*{arg\,min}_{\boldsymbol{\beta}^*\in\mathds{R}^{(N-H)}}\left\lVert\boldsymbol{M}_{(\boldsymbol{X}_\mathcal{H})}\boldsymbol{y}-\boldsymbol{M}_{(\boldsymbol{X}_\mathcal{H})}\boldsymbol{X}_{-\mathcal{H}}\boldsymbol{\beta}^*\right\rVert_2^2/T+2\lambda\left\lVert\boldsymbol{\beta}^*\right\rVert_1,\\
\hat\boldsymbol{\beta}_\mathcal{H}^{(L)}&=\hat\boldsymbol{\Sigma}_\mathcal{H}^{-1}\boldsymbol{X}_\mathcal{H}^\prime(\boldsymbol{y}-\boldsymbol{X}_{-\mathcal{H}}\hat\boldsymbol{\beta}_{-\mathcal{H}}^{(L)})/T,
\end{split}\end{equation*}
where the residual maker $\boldsymbol{M}_{(\boldsymbol{X}_\mathcal{H})}:=(I-\boldsymbol{X}_\mathcal{H}(\boldsymbol{X}_\mathcal{H}^\prime\boldsymbol{X}_\mathcal{H})^{-1}\boldsymbol{X}_\mathcal{H}^\prime)$, and $\hat\boldsymbol{\Sigma}_\mathcal{H}:=\frac{\boldsymbol{X}_\mathcal{H}^\prime\boldsymbol{X}_\mathcal{H}}{T}$.
We can therefore use regular lasso estimation on the transformed response $\boldsymbol{M}_{(\boldsymbol{X}_\mathcal{H})}\boldsymbol{y}$ and predictors $\boldsymbol{M}_{(\boldsymbol{X}_\mathcal{H})}\boldsymbol{X}_{-\mathcal{H}}$ to obtain $\hat{\boldsymbol{\beta}}^{(L)}_{-\mathcal{H}}$.
We next ``desparsify'' this estimate to allow for uniformly valid inference. The desparsified lasso estimator is defined as
\begin{equation}\label{eq:MDL}
\hat\boldsymbol{\beta}_\mathcal{H}:=\hat\boldsymbol{\beta}_\mathcal{H}^{(L)}+\hat\boldsymbol{\Theta}\boldsymbol{X}^\prime(\boldsymbol{y}-\boldsymbol{X}\hat\boldsymbol{\beta}^{(L)})/T,
\end{equation}
where $\hat\boldsymbol{\Theta}$ is an $H\times N$ submatrix of an approximate inverse of $\hat\boldsymbol{\Sigma}:=\frac{\boldsymbol{X}^\prime\boldsymbol{X}}{T}$. We obtain $\hat\boldsymbol{\Theta}$ from `nodewise' regressions, which estimate the linear projections of $\boldsymbol{x}_j$ onto $\boldsymbol{X}_{-j}$, where $\boldsymbol{x}_j$ is the $j$th column of $\boldsymbol{X}$ and $\boldsymbol{X}_{-j}$ is $\boldsymbol{X}$ without the $j$th column. We let
\begin{equation*}
\hat\boldsymbol{\gamma}_j=\operatorname*{arg\,min}_{\boldsymbol{\gamma}_j^*\in\mathds{R}^{N-1}}\left\lVert\boldsymbol{x}_j-\boldsymbol{X}_{-j}\boldsymbol{\gamma}_j^*\right\rVert_2^2/T+2\lambda_j\left\lVert\boldsymbol{\gamma}_j^*\right\rVert_1,
\end{equation*}
and construct $\hat\boldsymbol{\Theta}:=\hat\boldsymbol{\Upsilon}^{-2} \hat\boldsymbol{\Gamma}$, where
$\hat\boldsymbol{\Upsilon}^{-2}:=\text{diag}(1/\hat\tau^2_1,\dots 1/\hat\tau^2_H)$,
with $\hat\tau_j^2=\left\lVert\hat\boldsymbol{v}_j\right\rVert_2^2/T+\lambda_j\left\lVert\hat\boldsymbol{\gamma}_j\right\rVert_1, \hat\boldsymbol{v}_j=\boldsymbol{x}-\boldsymbol{X}\hat\boldsymbol{\gamma}_j$
and $\hat\boldsymbol{\Gamma}$ stacks the $\hat\boldsymbol{\gamma}_j$'s row-wise:
\begin{equation*}
\hat\boldsymbol{\Gamma}:=\left[\begin{array}{ccccc}
1 & \dots & -\hat\gamma_{1,H} & \dots & -\hat\gamma_{1,N}\\
\vdots & \ddots & \vdots & & \vdots \\
-\hat\gamma_{H,1} & \dots & 1 & \dots & -\hat\gamma_{H,N}
\end{array}\right].
\end{equation*}
The specific structure of the LPs consisting of the same set of regressors at each horizon $h$ allows for
computational and efficiency gains in obtaining the desparsified lasso.
The initial lasso in \cref{eq:problem} should be obtained for each horizon, but
the population linear projection coefficients $\boldsymbol{\gamma}_j$ in the nodewise regression
do not change with the horizon. It is therefore sufficient to compute the nodewise lasso estimator once. This can be done best at horizon zero where most time points are available for estimating the nodewise regression.
We therefore partly avoid the loss of efficiency
at further horizons where the most recent observation at each new horizon would be lost for estimation otherwise. As an example,
for the LP in \cref{eq:localprojection}, we
estimate
\begin{equation}\label{eq:nodewiseLP}
x_t=\boldsymbol{\psi}_0^\prime(y_t,{\boldsymbol{w}}_{s,t}^\prime)^\prime+\sum_{k=1}^K\boldsymbol{\psi}_k^\prime\boldsymbol{z}_{t-k} +v_{1,t},
\end{equation}
then store $\hat\boldsymbol{\Theta}$ constructed from $\hat\boldsymbol{\gamma}_{1}=\left[\hat\boldsymbol{\psi}_0^{\prime},\dots, \hat\boldsymbol{\psi}_{K}^{\prime}\right]^\prime$, and the residual vector $\hat\boldsymbol{v}_1$. The former
can then be re-used at each horizon in the
desparsified lasso estimator in \cref{eq:MDL}, while
the latter can be re-used in the estimation of the
long-run covariance matrix which enters the asymptotic distribution of the desparsified lasso, see Section \ref{subsec:inference}.
Finally, regarding the choice of tuning parameter $\lambda$ in the initial regression and $\lambda_j$ in the nodewise regressions,
we follow the approach of \citet{adamek2021lasso} who propose an iterative plug-in procedure that simulates the quantiles of an appropriate ``empirical process'' which should be dominated by $\lambda$; see their Section 5.1 for details.
\subsection{Local Projection Inference} \label{subsec:inference}
We next discuss the asymptotic properties of
$\hat\boldsymbol{\beta}_{\mathcal{H}}$. The assumptions needed for our analysis are listed in detail in Appendix A, and are discussed briefly below. Detailed discussions can be found in \cite{adamek2021lasso}, on whose approach our theory is built.
We need assumptions on the DGP (\cref{ass:DGP}), the sparsity of the parameter (\ref{ass:Sparsity}), regularity conditions for high-dimensional settings (\ref{ass:Eigenvalue}) and assumptions on the set of parameters on which inference is conducted (\ref{ass:HFinite}).
\cref{ass:DGP}\ref{ass:DGPStationary} requires that
the regressors are contemporaneously uncorrelated with the error term, and that the error terms
and the variables in $\boldsymbol{z}_{t}$ have finite unconditional moments up to a certain order. Crucially, \cref{ass:DGP}\ref{ass:DGPStationary} allows for serial correlation in the error terms, which is a typical feature of LP regressions.
\cref{ass:DGP}\ref{ass:DGPNED} assumes $\boldsymbol{z}_{t}$ to be near-epoch-dependent (NED). An NED process can be interpreted as a process that is well-approximated by a mixing process, but does not have to be mixing itself. Unlike mixing conditions, which may exclude even very simple autoregressive processes, NED conditions permit very general forms of dependence often encountered in macroeconomic research, see
\cite{adamek2021lasso} for a more detailed discussion.
Comparing our DGP assumptions
to the SVMA assumptions of Example \ref{ex:exampleSVMA},
one can verify that the moment condition is satisfied trivially under Gaussianity, and the NED assumption can be shown to hold when the VMA coefficients $\boldsymbol{A}_k$ converge to 0 sufficiently quickly; see Example 17.3 of \cite{Davidson02}.
If one would relax the Gaussianity assumption, a trade-off would occur between the number of moments and the allowed dependence -- the more moments, the more persistent the process may be.
Second, we impose a sparsity assumption on the parameter $\boldsymbol{\beta}$ (\cref{ass:Sparsity}), as typically required for lasso estimators.
We derive our results under weak sparsity, thereby recognizing that the true parameters are likely not exactly zero. In particular, we assume
that $\left\lVert\boldsymbol{\beta}\right\rVert_{r}^r$ is sufficiently small for some $0\leq r<1$.
While for $r=0$ our results boil down to
requiring exact sparsity
for larger values of $r$, we allow for many non-zero entries in $\boldsymbol{\beta}$ provided they are relatively small. The smaller $r$, the more restrictive this assumption therefore is. We also require weak sparsity in the nodewise regressions, which
amounts to assuming that the inverse population covariance matrix $\boldsymbol{\Sigma}^{-1}:=\left(\mathbb{E}\boldsymbol{X}^\prime\boldsymbol{X}/T\right)^{-1}$ is sparse in (induced) $r$-norm. Since \cref{eq:localprojection} has a similar regressor structure as a VAR, this can be shown to hold under conditions similar to the sparse VAR in Example 6 of \cite{adamek2021lasso}.
Next, we require that the minimum eigenvalue of $\boldsymbol{\Sigma}$ is bounded away from 0 (\cref{ass:Eigenvalue}).
This assumption ensures that the population covariance matrix of the regressors remains well-behaved with increasing $N$. Again, thanks to the similarity of \cref{eq:localprojection} to a VAR, lower bounds on the minimum eigenvalue can be derived using the results on page 6 of \citet{masini2019regularized}. Finally, we also require that the number of unpenalized parameters of interest to be bounded (\cref{ass:HFinite}). In LPs, this is trivially satisfied, for instance $H=1$ in \cref{eq:localprojection}, and $H=3$ in \cref{eq:statedependent}.\footnote{While the impulse responses are the two parameters of interest, we may also want to leave the dummy parameter unpenalized.}
Under these assumptions, \cref{thm:Inference} establishes the asymptotic normality of our desparsified lasso estimator, which allows for valid asymptotic inference.
\begin{theorem}\label{thm:Inference}
Under \cref{ass:DGP,ass:Sparsity,ass:Eigenvalue,ass:HFinite,ass:asymptoticrates}, for the model in \cref{eq:DGP}, we have
\begin{equation*}
\sup_{\underset{z\in\mathds{R}} {\boldsymbol{\beta} \in \boldsymbol{B}_N(r,s_r)}} \left\vert\mathbb{P}\left(\sqrt{T} \frac{\boldsymbol{r}' (\hat\boldsymbol{\beta}_\mathcal{H}-\boldsymbol{\beta}_\mathcal{H})}{\sqrt{\boldsymbol{r}' \hat\boldsymbol{\Upsilon}^{-2} \hat{\boldsymbol{\Omega}} \hat{\boldsymbol{\Upsilon}}^{-2} \boldsymbol{r}}} \leq z\right)-\Phi(z)\right\vert=o_p(1),
\end{equation*}
where $\Phi(\cdot)$ is the CDF of $N(0,1)$, $\boldsymbol{r}\in\mathds{R}^\mathcal{H}$ is chosen to test hypotheses of the form $\boldsymbol{r}^\prime \boldsymbol{\beta}_H=0$, and $\hat{\boldsymbol{\Omega}}$ is a consistent long-run covariance estimator.
\end{theorem}
Note that in LPs, the errors $u_{h,t}$ from \cref{eq:localprojection} are generally autocorrelated due to the leads of $y_t$ on the left-hand side. As in \cite{adamek2021lasso}, we therefore use an autocorrelation robust Newey-West long-run covariance estimator
$\hat\boldsymbol{\Omega}:=\sum_{l=1-Q_T}^{Q_T-1}\left(1-l/Q_T\right)\hat\boldsymbol{\Xi}(l)$,
with bandwidth $Q_T<T$, where
$\hat\boldsymbol{\Xi}(l):=\frac{1}{T-l}\left.\sum_{t=l+1}^{T}\hat\boldsymbol{q}_t\hat\boldsymbol{q}_{t-l}^{\prime}\right.$, $\hat\boldsymbol{q}_t:=\left(\hat v_{1,t}, \dots, \hat v_{H,t}\right)^\prime\hat u_t$.
To obtain our asymptotic results, we generally require that the bandwidth $Q_T$ grows as $T\to\infty$ at a sufficiently slow rate, see \cref{ass:asymptoticrates} in Appendix A for details. Regarding the choice of $Q_T$ in finite samples, we use the bandwidth estimator of \cite{andrews1991heteroskedasticity}.
For the LP in \cref{eq:localprojection}, Theorem \ref{thm:Inference}
permits asymptotic confidence intervals
$CI_h(\alpha):=\left[\hat\phi_h\pm z_{\alpha/2}\sqrt{T^{-1}\hat{\omega}_{ h}\hat{\tau}_{h}^{-4}}\right]$,
where $\hat\omega_{h}$ and $\hat\tau_{h}^{-2}$ are the scalar versions of $\hat\boldsymbol{\Omega}_h$ and $\hat\boldsymbol{\Upsilon}^{-2}_h$ respectively, and $z_{\alpha/2}:=\Phi^{-1}(1-\alpha/2)$. Note the dependence on $h$ in this result, since the residuals $\hat \boldsymbol{u}=\hat \boldsymbol{u}_{h}$ are different at each horizon, and $\hat\boldsymbol{\Omega}=\hat\boldsymbol{\Omega}_h$ and $\hat\boldsymbol{\Upsilon}^{-2}=\hat\boldsymbol{\Upsilon}_h^{-2}$ as well.
\section{Simulations}\label{sec:Simulations}
We perform a simulation study to analyse the finite sample performance of the desparsified lasso estimator.
In \cref{sec:Partial_penalization} we compare our proposed method with unpenalized parameter of interest to the standard desparsified lasso in a sparse structural VAR. In \cref{sec:DFM}, we study our proposed method in an empirically calibrated DFM.
\subsection{Sparse Structural VAR Model}\label{sec:Partial_penalization}
Consider the structural VAR with structural shocks $\boldsymbol{\epsilon}_t$
\begin{equation}\label{eq:sim}
\left(\begin{array}{c}
y_t\\
\boldsymbol{w}_{s,t}
\end{array}\right)=\underset{ P\times 1}{\boldsymbol{z}_t}=\sum_{k=1}^4\boldsymbol{A}_{k}\boldsymbol{z}_{t-k}+\boldsymbol{\epsilon}_t,\quad \boldsymbol{\epsilon}_t\overset{iid}{\sim}N(\boldsymbol{0},\boldsymbol{I}),
\end{equation}
where the autoregressive parameter matrices $\boldsymbol{A}_k$ are tapered Toeplitz matrices.
We analyse the response of $y_t$ to the first shock $\epsilon_{1,t}$ and investigate the coverage and width of 95\% point-wise confidence intervals corresponding to the impulse response parameters at horizons 1 to 10.
We consider different values for the number of variables $P=\{20, 40, 100\}$ and time series length $T =\{100, 200, 500\}$ and report results for 1000 replications. Details on the DGP and implementation of the HDLP are included in Appendix \ref{app:sparse_var}.
\begin{figure}[t]
\centering
\includegraphics[width=\textwidth]{"fig1".pdf}
\caption{Coverage rates of the standard desparsified lasso (red) and the proposed desparsified lasso with $\phi_{h}$ unpenalized (blue). Dashed lines indicate results for the second
DGP. }\label{fig:simresults}
\end{figure}
\Cref{fig:simresults} presents the coverage rates for the proposed desparsified lasso (blue), which leaves the impulse response parameters of interest unpenalized, in comparison to the standard desparsified lasso (red). We only display results for $P=40$, but the results for $P=20
$ and $P=100$ are very similar and included in \Cref{fig:S1}
in Appendix \ref{app:extra_figures},
alongside the interval widths in \Cref{fig:S2}.
Our proposed estimator attains good coverage rates close to the nominal level of 95\% for all combinations of $P$ and $T$.
For both DGPs, coverage rates are usually the lowest around 70-90\% for horizon 1, but revert to the nominal level for larger horizons. Good coverage at further horizons is to be expected, since the local projection coefficients will tend to become more sparse in stationary models.
Importantly, considerable gains are obtained compared to the standard desparsified lasso.
Especially at shorter horizons -- which are usually of most interest in practice -- coverage rates of the latter drop to 25\% in the first DGP and even more in the second. We also see in \Cref{fig:S2}
(Appendix \ref{app:extra_figures})
that the interval widths are broadly similar or even smaller than the ones for the standard desparsified lasso, meaning we obtain better coverage without losing power.
The poor coverage rates of the standard desparsified lasso occur since the parameters of interest are not selected in the initial lasso regression.
A simple alteration of the desparsified lasso that leaves this parameter unpenalized thus brings the coverage rates much closer to the nominal level. Note that the standard desparsified lasso has coverage exceeding our proposed estimator at further horizons; this is because the true impulse response becomes close to zero, where the lasso's bias towards zero is beneficial. This does not run counter to our conclusion, as we believe more uniform coverage over different parameter values is desirable.
\subsection{Empirically Calibrated Dynamic Factor Model}\label{sec:DFM}
We now examine the performance of our proposed estimator on an empirically calibrated dynamic factor model, inspired by \cite{lazarus2018har} and \cite{li2021local}.
Unlike the former works which use the quarterly dataset of \cite{stock2016DFM} to fit their DFM, we use
the monthly FRED-MD database (\citealp{mccracken2016fred}) which we analyse
in \Cref{sec:empiricalapplications}.
Implementation details on the simulation set-up are in Appendix \ref{sec:DFM_details},
an overview on the variables in the FRED-MD database in Appendix \ref{sec:data}.
We consider both a ``dense'' DFM where the factor loadings are obtained using principal components and a factor VAR is estimated by OLS, and a ``sparse'' DFM where the loadings are obtained using the WF-SOFAR estimator of \cite{uematsu2022estimation} and the VAR is estimated by the lasso.
The dense DFM is inherently difficult for our proposed estimator since the covariance structure of the variables is dense, whereas we assume it to be sparse in \Cref{ass:Sparsity}, and the DGP is likely to violate the sparsity of the nodewise regression coefficients.
However, bounds on the sparsity of the nodewise regressions can be obtained for sparse factor models and VARs, see Example 5 and 6 of \cite{adamek2021lasso}.
We therefore expect
better coverage properties of our proposed estimator is this setting.
The
WF-SOFAR estimator seems particularly appropriate to use since \cite{uematsu2022inference} find evidence of a sparse factor structure in the FRED-MD data.\footnote{The difference in sparsity between the dense and sparse DFMs is relatively limited. For details, see \Cref{fig:S3} in Appendix \ref{sec:DFM_details}.}
\begin{figure}[t]
\centering
\includegraphics[width=\textwidth]{"fig2".pdf}
\caption{Coverage rates and median interval widths
in the sparse DFM.}\label{fig:sparse_models}
\end{figure}
We study the impulse response of Industrial Production (IP) to the Federal Funds Rate (FFR).
Our HDLP method estimates \cref{eq:localprojection} with $y_t=\text{IP}_t$, $x_t=\text{FFR}_t$, the remaining 120 variables in $\boldsymbol{w}_{s,t}$, $\boldsymbol{w}_{f,t}$ empty, and 3 lags included.\footnote{Each local projection therefore has 488 regressors.}
We compare the performance of our proposal to two benchmarks.
The first benchmark is a factor-augmented local projection, labelled ``FA-LP''.
We favour FA-LP by including the true number of six factors, estimated by principal components from the 120 variables, and estimate the
FA-LP by OLS.
As this method matches the true DGP closely, we expect this benchmark to be a highly competitive standard.
Our second benchmark
is a 3-variable LP (estimated by OLS) with consumer price index as third variable as in \cite{bernanke2005measuring}. This benchmark, labelled ``LP", allows us to assess the gain in using high-dimensional data.
All models are estimated using the R package \texttt{desla} (\citealp{desla}), taking the plug-in constant equal to 0.4 for the HDLP.
In the dense DFM, see \Cref{fig:S4}
in Appendix C, our proposed HDLP method attains
coverage around 70\% for the first 5 periods, then steadily improves at further horizons where it reaches nominal coverage around horizon 20 for large sample sizes.
The benchmarks perform, however, considerably better as expected.
They perform similarly at $T=200$
with 90\% coverage, but diverge at larger sample sizes. The FA-LP model has improved coverage at larger samples, reaching nominal at $T=600$, whereas the coverage of the LP deteriorates at larger samples.
In the sparse DFM (\Cref{fig:sparse_models}), the performance of the HDLP considerably improves. Its coverage (top panels) is still below the benchmarks for $T=200$, stays around 85\% for all horizons, but it matches or exceeds the benchmarks at larger sample sizes. The performance of the FA-LP is similar to that of the dense DFM, and the LP benchmark maintains similar coverage to the FA-LP even at larger sample sizes; the 3-variable LP is likely a better approximation of the sparse DFM than the dense DFM.
Importantly though, we see across all sample sizes that the widths (bottom panels) of the FA-LP benchmark are 2-4 times larger than those of the HDLP, while the LP widths are also wider everywhere though to a lesser extent.
One should keep in mind that we set up the simulations advantageously to the FA-LP by using the true number of factors. The superior performance of the FA-LP, particularly for the dense model where our HDLP is expected to suffer, is thus hardly surprising.
Yet for the sparse DFM, our HDLP method outperforms both the idealized FA-LP and the 3-variable LP
with comparable or superior coverage
yet considerably tighter intervals.
To conclude our simulation experiments,
we investigate the sensitivity of our method to the choice of long-run covariance estimator. We compare the Newey-West (NW) estimator to
the fixed-$b$ Newey-West
and Equal-Weighted Cosine (EWC) estimator
suggested by \cite{lazarus2018har}.
In Appendix \ref{sec:other_LRVs}
we find that the first generally performs slightly worse than our default choice of Newey-West (likely due to our adaptive choice of bandwidth), whereas the second performs very well in terms of coverage but has confidence intervals around 4-6 times as wide as the ones from our NW estimator. In settings where our method has good coverage, such as the sparse DFM, EWC thus provides little benefit.
We therefore proceed with our default choice of NW estimator in the remainder of the paper, but these alternatives are available in the \texttt{desla} package.
\section{Structural Impulse Responses Estimated by HDLPs}\label{sec:empiricalapplications}
We apply the desparsified lasso for HDLPs to two canonical macroeconomic applications.
In \Cref{subsec:bernanke}, we build on the work by \cite{bernanke2005measuring} on monetary policy analysis,
in
\Cref{subsec:ramey} we build on the work by \cite{ramey2018government} on government spending.
All analyses are performed in \texttt{R}
using the package
\href{https://CRAN.R-project.org/package=desla}{\texttt{desla}}.
\subsection{Impulse Responses to a Shock in Monetary Policy} \label{subsec:bernanke}
We consider the work by \citet{bernanke2005measuring}, wherein they estimate impulse responses to a monetary policy shock identified by the federal funds rate in a factor-augmented SVAR.
We use HDLPs to perform the impulse response analysis and expand the dataset of the original paper
to the more recent and larger FRED-MD database \citep{McCrackenNg16}, thereby considering monthly data from January 1960 to
October 2008
for 122 macroeconomic variables. For a full description of the data, see Appendix \ref{sec:data}.
We estimate the HDLP in \cref{eq:localprojection}, where as response variables $y_t$ we take the federal funds rate (FFR), industrial production (IP) and the consumer price index (CPI). As shock variable of interest $x_t$ we consider the federal funds rate (FFR). Furthermore, we use the identification strategy of \citet{bernanke2005measuring}, which consists of classifying variables as ``slow" in $\boldsymbol{w}_{s,t}$ or ``fast" in $\boldsymbol{w}_{f,t}$, according to whether these variables can respond within the same month to an unexpected change in the FFR.\footnote{As discussed before, this strategy is equivalent to recursive (partial) identification in an SVAR, as shown in Example 1 of \citet{plagborg2021local}. In fact, this application can also be seen as a high-dimensional extension of \citet{christiano2005nominal}.}
For variables in FRED-MD that can be directly matched to their counterparts in the DRI/McGraw Hill Basic Economics Database, we use the same slow/fast classification as \citet{bernanke2005measuring}.
For new variables in FRED-MD, we apply the same general classification rule; i.e., variables in the categories Prices, Output \& Income, Labour Market, and Consumption are classified as slow, whereas those in the categories Interest \& Exchange Rates, Money \& Credit, Stock Market, and Housing as fast.\footnote{One exception in the Category Prices is the series ``OILPRICEx", which we classify as fast.} We apply the transformation codes of \citet{bernanke2005measuring} and \citet{McCrackenNg16} (for new variables) to remove stochastic trends. Using $K=13$ lags, the resulting HDLP consists of $N=1654$ regressors, while the time series length is $T=572$.\footnote{The exact number of variables in the model depends on which variable is taken as the response, and the effective number of observations depends on the horizon. These numbers are for the model where FFR is the response, at horizon 1.}
\begin{figure}[t]
\centering
\includegraphics[width=\textwidth]{"fig3".pdf}
\caption{Estimated impulse responses with 95\% confidence intervals of FFR, IP, and CPI to a monetary policy shock identified by the FFR for the HDLP and FAVAR.
}\label{fig:FAVAR v HDLP}
\end{figure}
We compare the HDLP impulse responses obtained from the desparsified lasso to the ones obtained from a 3-factor FAVAR as used in \citet{bernanke2005measuring}. Details about the FAVAR estimation are provided in Appendix \ref{app:FAVAR}.
\Cref{fig:FAVAR v HDLP} shows the impulse responses of FFR, IP, and CPI to a shock in the FFR of a size such that the FFR has unit response at horizon 0.
Since we use the first difference of IP and CPI in the model, we cumulate the impulse response which then corresponds to the response in the original (level) variables.\footnote{Cumulative impulse response functions for the HDLP are obtained by cumulating the dependent variable, i.e. taking $\sum_{\ell=0}^{h}y_{t+\ell}$ on the LHS of \cref{eq:localprojection}.
For the FAVAR, we take the cumulative sum of the regular impulse responses from previous horizons. To ensure that these impulse responses are comparable, we scale the FAVAR impulse responses by the response of the FFR at horizon 0.
} The comparable figure in \citet{bernanke2005measuring} is Figure 1.
The impulse responses from the FAVAR (bottom panel \Cref{fig:FAVAR v HDLP}) and
HDLP (top panel) are generally similar.
The impulse response of the FFR peaks at horizon 1 and then steadily declines to zero, which is in line with Figure 1 in \citet{bernanke2005measuring}.
For the HDLP, the response stays at a higher level for a longer period than for the FAVAR. The response of IP is also in line with \citet{bernanke2005measuring}, with the largest drop around horizon 20, before eventually returning to zero. Notably, the response obtained with the HDLP is considerably smaller than the one for the FAVAR. Finally for the response of CPI,
we see a more pronounced price puzzle with the HDLP. The response is positive and significant for 30 months, peaking at horizon 20. The FAVAR impulse response is largely in line with \citet{bernanke2005measuring}, with a small positive effect at early horizons, followed by a negative, though mainly insignificant, response after horizon 10.
\subsection{Impulse Responses to a Shock in Government Spending
} \label{subsec:ramey}
\cite{ramey2018government} estimate impulse responses to a shock in US government spending identified based on military spending news. In this paper, we first augment the authors' main LP specification with more lags as a robustness check, before considering an extended state-dependent HDLP specification with interacting states related to unemployment and recession conditions. We use the quarterly data
provided by the authors at \href{https://econweb.ucsd.edu/~vramey/research.html#govt}{\texttt{econweb.ucsd.edu/$\sim$vramey}} covering the period 1889Q1 to 2015Q4.
We estimate the state-dependent HDLP in \cref{ex:ramey}
for
real per capita GDP and government spending as $y_t$, while the shock variable $x_t$ is the military spending news shock.
We include taxes as controls in $\boldsymbol{w}_{f,t}$, while no variables are included in $\boldsymbol{w}_{s,t}$.\footnote{While \cite{ramey2018government} do not include taxes in their main analysis, they mention in footnote 11 that their results were robust to their inclusion. We prefer to include them as additional controls, following the VAR of \citet{blanchard2002empirical}.}
The state dummy variable $I_t$ distinguishes
between low and high unemployment states, with $I_t=1$ when unemployment in period $t$ is larger than 6.5\%.
For full details on the construction of the shock variable and the treatment of the other variables, see section II.B of \cite{ramey2018government}. We use $K=40$ lags, such that the resulting HDLP consists of $N=464$ regressors, while the time series length is $T=161$.\footnote{As in \Cref{subsec:bernanke}, these correspond to the model estimated at horizon 1.}
\begin{figure}[t]
\centering
\includegraphics[width=\textwidth]{"fig4".pdf}
\caption{Estimated impulse responses with 95\% confidence intervals of government spending and GDP to a government spending shock, in a model with 40 lags.}\label{fig:overall_RZ_40lags}
\end{figure}
Figure \ref{fig:overall_RZ_40lags} shows the impulse responses to a shock in government spending. Note that the military spending news variable is scaled by GDP such that the impulse responses can be seen as a reaction to a shock of size 1\% of GDP. The comparable figure in \cite{ramey2018government} is Figure 5. The general shape and magnitude of these impulse responses are similar to the LP model with only 4 lags used in \cite{ramey2018government}. In our analysis, the peak of the responses in the high unemployment state occurs 2-3 quarters earlier, and experiences a sharper drop after horizon 15. We also find that the impulse response is not significant at horizon 0 in the linear and low-unemployment state.
\begin{figure}[t]
\centering
\includegraphics[width=\textwidth]{"fig5".pdf}
\caption{Estimated Impulse responses with 95\% confidence intervals of government spending and GDP to a government spending shock.}\label{fig:2_states}
\end{figure}
Next to allowing the inclusion of more lags, the HDLP framework can easily be extended to allow for $S$ states.
The LP is then given by
\begin{equation*}
y_{t+h}=\sum_{i=1}^S \alpha_i I_{i,t-1} + \sum_{i=1}^S I_{i,t-1} \left[\phi_{i,h}x_{t}+\sum_{k=1}^K\boldsymbol{\delta}_{i,h,k}^\prime \boldsymbol{z}_{t-k}\right]+u_{h,t}.
\end{equation*}
In addition to the unemployment state dummy $I_t^{(U)}$, we consider recession dummy $I_t^{(R)}$.\footnote{As in \cite{ramey2018government}, we follow the NBER classification of recession periods.}
Letting them interact
results in $S=4$ distinct state dummies.
\Cref{fig:2_states} displays the impulse response functions with four states. The impulse responses out of recessions resemble the state-dependent model of \Cref{fig:overall_RZ_40lags}, both in shape and magnitude, but we see a different pattern in the recession state. During recessions with low unemployment, there is a negative impact response for both government spending and GDP, which becomes insignificant after a few periods. In the high unemployment state, in contrast, we see a large, positive, and persistent effect on both variables.
\section{Conclusion} \label{sec:conclusion}
In this paper, we propose a modified version of the desparsified lasso to estimate local projections in high dimensions.
The modification is simple, namely leaving the (small number of) dynamic impulse response parameter(s) of interest unpenalized in the initial lasso regression, yet provides considerable improvements in finite sample performance in terms of better coverage. The modified desparsified lasso estimator still comes with desirable asymptotic normality and uniformly valid asymptotic inference. Finally, we show how this method performs in a simulation using an empirically calibrated dynamic factor model and in canonical macroeconomic applications where we use the high-dimensional local projections framework to estimate structural impulse responses.
\section*{Acknowledgements}
We thank the
co-editor and two referees for their thorough review
which substantially improved the quality of the manuscript.
The first and second author were financially supported by the Dutch Research Council (NWO) under grant number 452-17-010.
Previous versions of this paper were presented at seminars at Aarhus University and ESSEC Business School, the NESG 2022 conference and the 2022 Maastricht Workshop on Dimensionality Reduction and Inference in High-Dimensional Time Series. We gratefully acknowledge comments by participants at these seminars and conferences. We also thank Otilia Boldea and Lenard Lieb for helpful discussions. Remaining errors are our own.
\bibliographystyle{chicago}
\bibliography{bibliography}