EconBase
← Back to paper

Local Projection Inference in High Dimensions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

51,110 characters · 12 sections · 98 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Local Projection Inference in High Dimensions

\onehalfspacing

abstractIn this paper, we estimate impulse responses by local projections in high-dimensional settings. We use the desparsified (de-biased) lasso to estimate the high-dimensional local projections, while leaving the impulse response parameter of interest unpenalized. We establish the uniform asymptotic normality of the proposed estimator under general conditions. Finally, we demonstrate small sample performance through a simulation study and consider two canonical applications in macroeconomic research on monetary policy and government spending. Keywords: {Local projections, Impulse response analysis, High-dimensional data, Honest inference, Lasso}

Introduction

In this paper, we develop a simple approach for conducting valid inference on impulse responses in high-dimensional settings. We do so by pairing the estimation framework of local projections (LPs) with the inference framework of the desparsified (or debiased) lasso. Since their introduction by jorda2005estimation, LPs have been widely used in macroeconomic research for studying the dynamic propagation of shocks through impulse response analysis, see e.g. angrist2018semiparametric, ramey2018government, ramey2016macroeconomic and stock2018identification. As such, they have become an increasingly used alternative to structural Vector AutoRegressions (SVAR) pioneered by sims1980macroeconomics, see for instance ramey2016macroeconomic or kilian2017structural for an overview on this work.

While the methods have different finite-sample properties, SVARs and LPs are equivalent in population, as established by plagborg2021local who show that the underlying impulse response estimands are the same for both. This implies that the well-documented necessity for SVARs to add assumptions in order to achieve structural identification of the impulse responses, is equally true for LPs. Indeed, plagborg2021local show that the identification strategies typically used for SVARs have an equivalent implementation for LPs, and vice versa. While identification issues therefore do not affect the choice between SVAR or LP, finite-sample considerations about estimation and inference problem do. LP impulse responses are obtained by estimating only univariate linear regressions, and performing standard inference on (typically) a single parameter of interest across these univariate regressions. In contrast, SVARs require estimating the whole system of equations, and transforming these into the Vector Moving Average (VMA) representation from which the impulse responses can be derived. While the system estimation of the SVAR can still be done linearly, the inversion to the VMA representation is a nonlinear operation that renders the impulse response coefficients a complex nonlinear functions of all VAR parameters. This makes inference considerably more cumbersome, even in low-dimensional settings, where the complications and inaccuracies of applying the Delta method in finite samples have led to bootstrap inference becoming the norm kilian2017structural. These problems are exacerbated when the dimensionality of the system grows, making LPs our method of choice to obtain impulse responses in high dimensions.

We consider high-dimensional local projections (HDLPs) in a general time series framework where the number of regressors can grow faster than the sample size. Impulse response analysis quickly becomes high-dimensional in macroeconomic research. Even when considering impulse response analysis with few variables, the number of regressors in LPs or SVARs is often large due to the common practices of including many lags to control for autocorrelation (see e.g., bernanke1998measuring, romer2004new and sims2006were) or to robustify against (near) unit roots via lag augmentation as in olea2020local. Similarly, additional regressors are often introduced to capture seasonal patterns in impulse responses (see e.g., quarter-dependent coefficients used in blanchard2002empirical), or to permit nonlinearities to produce state-dependent impulse responses (koop1996impulse; ramey2018government).

Additionally, one might be interested in impulse response analysis with many variables. Motivations of the latter include, amongst others, avoidance of contaminated measurements of policy innovations or informational deficiency with impulse response estimates that are distorted by omitted variable bias, the ability to use several observable measures of difficult-to-measure variables in theoretical models, or opportunities for researchers and policy makers to investigate the impact of shocks on a larger, possibly more disaggregated set of variables they care about (see e.g., bernanke2005measuring, forni2014sufficient, stock2016DFM and kilian2017structural).

Regular estimation approaches, however, become infeasible for high-dimensional settings. Traditional approaches to high-dimensionality include modelling commonalities between variables, such as through factor-augmented VARs (FAVAR) (bernanke2005measuring), dynamic factor models (DFM) (forni2009opening; stock2016DFM) or -- for panel structures -- global VARs (ChudikPesaran16); as well as Bayesian shrinkage methods (banbura2010large; chan2020large). Recently, sparse shrinkage methods such as the lasso have gained considerable popularity in econometrics for policy evaluation (BCH14) and (macroeconomic) time series analysis (kock2020penalized). Their use in impulse response analysis however has only scarcely been explored.

While several methods and theoretical results now exist for estimating sparse VAR models -- see e.g., BasuMichailidis15, KockCallot2015, masini2019regularized and the references cited therein -- inference on impulse responses is complicated by two issues. First, sparse estimation techniques such as the lasso perform model selection, which induces issues with non-uniformity of limit results if this selection is ignored (LeebPoetscher05). Second, while several methods such as orthogonalization (BCH14) and debiasing, or desparsifying the lasso (vandeGeer14; JavanmardMontanari14) have been proposed to yield uniformly valid (or `honest') post-selection inference, the impulse response parameters are nonlinear functions of all estimated VAR parameters. This severely limits the applicability of existing post-selection inference methods which are typically designed for (relatively) low-dimensional parameters of interest that can be estimated directly. Indeed, to our knowledge, impulse response analysis in sparse HD-SVARs is only considered in krampe2022, who construct a complex multi-step algorithm to overcome these complications. Instead, by casting the problem in the LP framework, we reduce the impulse response parameter(s) to a (directly estimable) low-dimensional object in the presence of high-dimensional nuisance parameters, which makes the standard post-selection tools available.

We develop HDLP inference based on the desparsified lasso of vandeGeer14 and its time series extension in adamek2021lasso. The latter has recently been empirically investigated in the context of LP estimation with instrumental variables (LP-IV) in contemporaneous work by Karapanagioti21. Their approach not only differs in the focus on IV models, but they also apply the method of adamek2021lasso `as is'. Instead, we tailor the approach of adamek2021lasso specifically to high-dimensional LPs consisting of a small number of parameters of interest -- the dynamic response of a variable to a shock at a given horizon -- and many controls. In particular, we modify their approach by leaving the parameter of interest unpenalized during the estimation procedure to ensure it does not suffer from penalization bias. Such a setting is of more general relevance for treatment effect models consisting of a small number of variables whose effects are of interest combined with a large set of controls. We theoretically show that the combination of few unpenalized parameters with many penalized ones does not affect the asymptotic behaviour of the desparsified lasso.

Through simulation experiments, we show that our proposed estimator has considerably better coverage rates than the standard desparsified lasso in finite samples. Furthermore, on an empirically calibrated DFM (li2021local) we find that our proposed HDLP method shows competitive performance in a sparse DFM compared to a low-dimensional LP and a factor-augmented LP.

We consider two canonical macroeconomic applications and demonstrate the performance of the proposed desparsified-lasso based estimator for HDLPs in recovering structural impulse responses. Specifically, we first extend the work of bernanke2005measuring on macroeconomic responses to a shock in monetary policy to a HDLP setting. Second, we consider the work by ramey2018government on state-dependent impulse responses to a shock in government spending, as also studied by Karapanagioti21 in their LP-IV setting, but we extend the original specification used in ramey2018government to a HDLP specification with more lags as well as a more complex state-dependent HDLP with interaction between different state variables. The methods needed for these applications are implemented in the R package desla (desla), which also offers the methods of adamek2021lasso.

The paper is organized as follows. Section (ref) introduces the model and our inferential procedure, as well as our main result on the asymptotic normality of the desparsified lasso. Section (ref) contains simulation studies that examine the small sample performance of the proposed desparsified lasso. Section (ref) presents the results of the two macroeconomic applications. Section (ref) concludes.

A word on notation. For any $N$ dimensional vector $\boldsymbol{x}$, $\left\Vert \boldsymbol{x}\right\Vert_r=\left(\sum_{i=1}^{N}\left\vert x_i\right\vert^r\right)^{1/r}$ denotes the $l_r$-norm, with the convention that $\left\lVert\boldsymbol{x}\right\rVert_0 = \sum_{i} 1 (\left\lvertx_i\right\rvert>0)$. Depending on the context, $\sim$ denotes equivalence in order of magnitude of sequences, or equivalence in distribution.

High-dimensional Local Projections

Consider the local projection regression

equation[equation omitted — 239 chars of source]

where $\phi_{h}, \rho_{h}, \boldsymbol{\eta}_h, \boldsymbol{\delta}_{h,k}$ are the projection parameters, $u_{h,t}$ is the projection error and the set of variables $\boldsymbol{z}_t=\big(\boldsymbol{w}_{s,t}^\prime,y_t,x_t, \boldsymbol{w}_{f,t}^\prime\big)^\prime$ consist of the response $y_t$, the shock variable $x_t$, and the vectors of control variables split into the “slow" variables $\boldsymbol{w}_{s,t}\in\mathds{R}^{n_s}$, and the “fast" ones $\boldsymbol{w}_{f,t}\in\mathds{R}^{n_f}$ for identification purposes, as discussed below. \footnote{We omit the constant, since we generally demean the data when using the desparsified lasso.} Our single parameter of interest is $\phi_{h}$, the dynamic response at horizon $h$ of $y_t$ after an impulse in $x_t$. The LP impulse response function of $y_t$ with respect to $x_t$ is given by $\{ \phi_{h}\}_{h\geq 0}$ and can be obtained by estimating (ref) for $h=0,1,\dots,h_{\max}$.

To permit causal interpretation of the impulse responses, structural assumptions about the data generating process are needed. Example (ref) considers a Structural Vector Moving Average model (see Assumption 3 in plagborg2021local) as an example.

example[SVMA] Consider the Structural Vector Moving Average (SVMA) \begin{equation} {\boldsymbol{z}_t}=\boldsymbol{\mu}+ {\boldsymbol{A}(L)} {\boldsymbol{\epsilon}_t}, \boldsymbol{A}(L)=\sum_{k=0}^\infty {\boldsymbol{A}_{k}}L^k, \boldsymbol{\epsilon}_{t}\overset{i.i.d.}{\sim}N(\boldsymbol{0},\boldsymbol{I}), \end{equation} where $\boldsymbol{\epsilon}_t$ is a vector of structural shocks, $L$ is the lag operator, $\{\boldsymbol{A}_{k}\}_k$ is absolutely summable, and $\boldsymbol{A}(x)$ has full row rank for all complex scalars $x$ on the unit circle.

While an assumption such as (ref) is necessary to recover the structural impulse responses, it is not a sufficient condition: $\boldsymbol{A}_0$ is not identified without further restrictions. We focus on impulse responses identified by a (partially) recursive structure. This can be implemented in LPs by partitioning the variables in $\boldsymbol{z}_t$ into slow variables ($\boldsymbol{w}_{s,t}$) and fast variables ($\boldsymbol{w}_{f,t}$). The slow variables are predetermined with respect to $x_t$, while the fast variables can react contemporaneously to the shock variable $x_t$ and therefore only enter the equation with a lag. Equivalently, in the language of recursively identified SVARs, $\boldsymbol{w}_{s,t}$ is ordered above $x_t$, while $\boldsymbol{w}_{s,t}$ is ordered below. Note also that in (ref) we assume that $y_t$ is predetermined with respect to (ordered above) $x_t$; if the reverse were true one should remove $y_t$ from the right-hand side of (ref).\footnote{In this situation, the model would be equivalent to equation (1) of plagborg2021local.}

(ref) can also be used outside of the recursive identification framework, for example when the structural shock is externally identified. Examples of the latter are the narrative monetary policy shock series of romer2004new, or the military spending series of ramey2018government. One can then directly take $x_t=\epsilon_{1,t}$ provided the structural shock is observed without measurement error. Given the exogeneity (and thus predeterminedness) of such shocks, all other variables would then be considered to be “fast” and no variable would be included contemporaneously; see e.g. (ref).

Section 3.2 of plagborg2021local provides detailed discussions on how other identification schemes, such as sign restrictions, can be imposed while estimating (ref); these arguments generally remain valid in our high-dimensional setup as well. However, identification through instrumental variables as in e.g. stock2018identification becomes more complicated. Adapting IV estimation methods such as two-stage least-squares (plagborg2021local) to high dimensions requires a different setup for the desparsified lasso (see e.g. gold2020inference), which has not yet been explored for time series models. This remains an interesting avenue for future research.

While the LP in (ref) has a single parameter of interest, it can easily be extended to allow for LPs with a small number of parameters of interest, see Example (ref).

example[State-Dependent LP] Consider the nonlinear state-dependent model of ramey2018government: \begin{equation} y_{t+h}= I_{t-1}\left[\alpha_h+\phi_{A,h}x_{t}+\sum_{k=1}^K\boldsymbol{\delta}_{A,h,k}^\prime\boldsymbol{z}_{t-k} \right] +(1-I_{t-1})\left[\phi_{B,h}x_{t}+\sum_{k=1}^K\boldsymbol{\delta}_{B,h,k}^\prime\boldsymbol{z}_{t-k} \right]+u_{h,t}. \end{equation} Here, $\phi_{A,h}$ and $\phi_{B,h}$ are the two responses of $y_t$ after an impulse in $x_t$, corresponding to two different states distinguished by the dummy variable $I_t$.

To generally accommodate LP settings with a small number of parameters of interest and a large number of controls, we will use the more general notation

equation[equation omitted — 293 chars of source]

and in matrix form

equation*[equation* omitted — 167 chars of source]

We hereby separate the parameter vector $\boldsymbol{\beta}=[\boldsymbol{\beta}_\mathcal{H}^{\prime},\boldsymbol{\beta}_{-\mathcal{H}}^{\prime}]^{\prime}$ into two groups. The first group consists of the $H$-dimensional set of parameters of interest $\boldsymbol{\beta}_\mathcal{H}$, corresponding to the variables $\boldsymbol{x}_{\mathcal{H},t}$, which will be unpenalized during the estimation procedure to ensure they are not affected by penalization bias. The second group consists of the large set of parameters $\boldsymbol{\beta}_{-\mathcal{H}}$, corresponding to control variables $\boldsymbol{x}_{-\mathcal{H},t}$, which will be penalized during the estimation procedure. We hereby use $\mathcal{H}$ to denote the set that indexes the $H$ variables of interest. Without loss of generality, we order these variables first in $\boldsymbol{X}=[\boldsymbol{X}_\mathcal{H},\boldsymbol{X}_{-\mathcal{H}}]$. Finally, since each horizon $h$ of the LPs is estimated separately, we suppress the dependence on $h$ in our notation.

Local Projection Estimation

We start by estimating (ref) with the lasso to obtain an initial estimate of $\boldsymbol{\beta}$, while leaving the parameters of interest $\boldsymbol{\beta}_{\mathcal{H}}$ unpenalized. The first stage lasso is defined as

equation[equation omitted — 389 chars of source]

where $\boldsymbol{W}$ is an $N\times N$ diagonal matrix with $\boldsymbol{W}_{i,i}=0$ for $i\in \mathcal{H}$ and 1 otherwise. The Frish-Waugh-Lovell-style result of yamada2017 ensures that the penalized and unpenalized estimates can be respectively obtained by

equation*[equation* omitted — 612 chars of source]

where the residual maker $\boldsymbol{M}_{(\boldsymbol{X}_\mathcal{H})}:=(I-\boldsymbol{X}_\mathcal{H}(\boldsymbol{X}_\mathcal{H}^\prime\boldsymbol{X}_\mathcal{H})^{-1}\boldsymbol{X}_\mathcal{H}^\prime)$, and $\hat\boldsymbol{\Sigma}_\mathcal{H}:=\frac{\boldsymbol{X}_\mathcal{H}^\prime\boldsymbol{X}_\mathcal{H}}{T}$. We can therefore use regular lasso estimation on the transformed response $\boldsymbol{M}_{(\boldsymbol{X}_\mathcal{H})}\boldsymbol{y}$ and predictors $\boldsymbol{M}_{(\boldsymbol{X}_\mathcal{H})}\boldsymbol{X}_{-\mathcal{H}}$ to obtain $\hat{\boldsymbol{\beta}}^{(L)}_{-\mathcal{H}}$.

We next “desparsify” this estimate to allow for uniformly valid inference. The desparsified lasso estimator is defined as

equation[equation omitted — 217 chars of source]

where $\hat\boldsymbol{\Theta}$ is an $H\times N$ submatrix of an approximate inverse of $\hat\boldsymbol{\Sigma}:=\frac{\boldsymbol{X}^\prime\boldsymbol{X}}{T}$. We obtain $\hat\boldsymbol{\Theta}$ from `nodewise' regressions, which estimate the linear projections of $\boldsymbol{x}_j$ onto $\boldsymbol{X}_{-j}$, where $\boldsymbol{x}_j$ is the $j$th column of $\boldsymbol{X}$ and $\boldsymbol{X}_{-j}$ is $\boldsymbol{X}$ without the $j$th column. We let

equation*[equation* omitted — 264 chars of source]

and construct $\hat\boldsymbol{\Theta}:=\hat\boldsymbol{\Upsilon}^{-2} \hat\boldsymbol{\Gamma}$, where $\hat\boldsymbol{\Upsilon}^{-2}:=\text{diag}(1/\hat\tau^2_1,\dots 1/\hat\tau^2_H)$, with $\hat\tau_j^2=\left\lVert\hat\boldsymbol{v}_j\right\rVert_2^2/T+\lambda_j\left\lVert\hat\boldsymbol{\gamma}_j\right\rVert_1, \hat\boldsymbol{v}_j=\boldsymbol{x}-\boldsymbol{X}\hat\boldsymbol{\gamma}_j$ and $\hat\boldsymbol{\Gamma}$ stacks the $\hat\boldsymbol{\gamma}_j$'s row-wise:

equation*[equation* omitted — 266 chars of source]

The specific structure of the LPs consisting of the same set of regressors at each horizon $h$ allows for computational and efficiency gains in obtaining the desparsified lasso. The initial lasso in (ref) should be obtained for each horizon, but the population linear projection coefficients $\boldsymbol{\gamma}_j$ in the nodewise regression do not change with the horizon. It is therefore sufficient to compute the nodewise lasso estimator once. This can be done best at horizon zero where most time points are available for estimating the nodewise regression. We therefore partly avoid the loss of efficiency at further horizons where the most recent observation at each new horizon would be lost for estimation otherwise. As an example, for the LP in (ref), we estimate

equation[equation omitted — 182 chars of source]

then store $\hat\boldsymbol{\Theta}$ constructed from $\hat\boldsymbol{\gamma}_{1}=\left[\hat\boldsymbol{\psi}_0^{\prime},\dots, \hat\boldsymbol{\psi}_{K}^{\prime}\right]^\prime$, and the residual vector $\hat\boldsymbol{v}_1$. The former can then be re-used at each horizon in the desparsified lasso estimator in (ref), while the latter can be re-used in the estimation of the long-run covariance matrix which enters the asymptotic distribution of the desparsified lasso, see Section (ref).

Finally, regarding the choice of tuning parameter $\lambda$ in the initial regression and $\lambda_j$ in the nodewise regressions, we follow the approach of adamek2021lasso who propose an iterative plug-in procedure that simulates the quantiles of an appropriate “empirical process” which should be dominated by $\lambda$; see their Section 5.1 for details.

Local Projection Inference

We next discuss the asymptotic properties of $\hat\boldsymbol{\beta}_{\mathcal{H}}$. The assumptions needed for our analysis are listed in detail in Appendix A, and are discussed briefly below. Detailed discussions can be found in adamek2021lasso, on whose approach our theory is built.

We need assumptions on the DGP ((ref)), the sparsity of the parameter ((ref)), regularity conditions for high-dimensional settings ((ref)) and assumptions on the set of parameters on which inference is conducted ((ref)). (ref)(ref) requires that the regressors are contemporaneously uncorrelated with the error term, and that the error terms and the variables in $\boldsymbol{z}_{t}$ have finite unconditional moments up to a certain order. Crucially, (ref)(ref) allows for serial correlation in the error terms, which is a typical feature of LP regressions. (ref)(ref) assumes $\boldsymbol{z}_{t}$ to be near-epoch-dependent (NED). An NED process can be interpreted as a process that is well-approximated by a mixing process, but does not have to be mixing itself. Unlike mixing conditions, which may exclude even very simple autoregressive processes, NED conditions permit very general forms of dependence often encountered in macroeconomic research, see adamek2021lasso for a more detailed discussion.

Comparing our DGP assumptions to the SVMA assumptions of Example (ref), one can verify that the moment condition is satisfied trivially under Gaussianity, and the NED assumption can be shown to hold when the VMA coefficients $\boldsymbol{A}_k$ converge to 0 sufficiently quickly; see Example 17.3 of Davidson02. If one would relax the Gaussianity assumption, a trade-off would occur between the number of moments and the allowed dependence -- the more moments, the more persistent the process may be.

Second, we impose a sparsity assumption on the parameter $\boldsymbol{\beta}$ ((ref)), as typically required for lasso estimators. We derive our results under weak sparsity, thereby recognizing that the true parameters are likely not exactly zero. In particular, we assume that $\left\lVert\boldsymbol{\beta}\right\rVert_{r}^r$ is sufficiently small for some $0\leq r<1$. While for $r=0$ our results boil down to requiring exact sparsity for larger values of $r$, we allow for many non-zero entries in $\boldsymbol{\beta}$ provided they are relatively small. The smaller $r$, the more restrictive this assumption therefore is. We also require weak sparsity in the nodewise regressions, which amounts to assuming that the inverse population covariance matrix $\boldsymbol{\Sigma}^{-1}:=\left(\mathbb{E}\boldsymbol{X}^\prime\boldsymbol{X}/T\right)^{-1}$ is sparse in (induced) $r$-norm. Since (ref) has a similar regressor structure as a VAR, this can be shown to hold under conditions similar to the sparse VAR in Example 6 of adamek2021lasso.

Next, we require that the minimum eigenvalue of $\boldsymbol{\Sigma}$ is bounded away from 0 ((ref)). This assumption ensures that the population covariance matrix of the regressors remains well-behaved with increasing $N$. Again, thanks to the similarity of (ref) to a VAR, lower bounds on the minimum eigenvalue can be derived using the results on page 6 of masini2019regularized. Finally, we also require that the number of unpenalized parameters of interest to be bounded ((ref)). In LPs, this is trivially satisfied, for instance $H=1$ in (ref), and $H=3$ in (ref).\footnote{While the impulse responses are the two parameters of interest, we may also want to leave the dummy parameter unpenalized.}

Under these assumptions, (ref) establishes the asymptotic normality of our desparsified lasso estimator, which allows for valid asymptotic inference.

theoremUnder (ref), for the model in (ref), we have \begin{equation*} \sup_{\underset{z\in\mathds{R}} {\boldsymbol{\beta} \in \boldsymbol{B}_N(r,s_r)}} \left\vert\mathbb{P}\left(\sqrt{T} \frac{\boldsymbol{r}' (\hat\boldsymbol{\beta}_\mathcal{H}-\boldsymbol{\beta}_\mathcal{H})}{\sqrt{\boldsymbol{r}' \hat\boldsymbol{\Upsilon}^{-2} \hat{\boldsymbol{\Omega}} \hat{\boldsymbol{\Upsilon}}^{-2} \boldsymbol{r}}} \leq z\right)-\Phi(z)\right\vert=o_p(1), \end{equation*} where $\Phi(\cdot)$ is the CDF of $N(0,1)$, $\boldsymbol{r}\in\mathds{R}^\mathcal{H}$ is chosen to test hypotheses of the form $\boldsymbol{r}^\prime \boldsymbol{\beta}_H=0$, and $\hat{\boldsymbol{\Omega}}$ is a consistent long-run covariance estimator.

Note that in LPs, the errors $u_{h,t}$ from (ref) are generally autocorrelated due to the leads of $y_t$ on the left-hand side. As in adamek2021lasso, we therefore use an autocorrelation robust Newey-West long-run covariance estimator $\hat\boldsymbol{\Omega}:=\sum_{l=1-Q_T}^{Q_T-1}\left(1-l/Q_T\right)\hat\boldsymbol{\Xi}(l)$, with bandwidth $Q_T<T$, where $\hat\boldsymbol{\Xi}(l):=\frac{1}{T-l}\left.\sum_{t=l+1}^{T}\hat\boldsymbol{q}_t\hat\boldsymbol{q}_{t-l}^{\prime}\right.$, $\hat\boldsymbol{q}_t:=\left(\hat v_{1,t}, \dots, \hat v_{H,t}\right)^\prime\hat u_t$. To obtain our asymptotic results, we generally require that the bandwidth $Q_T$ grows as $T\to\infty$ at a sufficiently slow rate, see (ref) in Appendix A for details. Regarding the choice of $Q_T$ in finite samples, we use the bandwidth estimator of andrews1991heteroskedasticity.

For the LP in (ref), Theorem (ref) permits asymptotic confidence intervals $CI_h(\alpha):=\left[\hat\phi_h\pm z_{\alpha/2}\sqrt{T^{-1}\hat{\omega}_{ h}\hat{\tau}_{h}^{-4}}\right]$, where $\hat\omega_{h}$ and $\hat\tau_{h}^{-2}$ are the scalar versions of $\hat\boldsymbol{\Omega}_h$ and $\hat\boldsymbol{\Upsilon}^{-2}_h$ respectively, and $z_{\alpha/2}:=\Phi^{-1}(1-\alpha/2)$. Note the dependence on $h$ in this result, since the residuals $\hat \boldsymbol{u}=\hat \boldsymbol{u}_{h}$ are different at each horizon, and $\hat\boldsymbol{\Omega}=\hat\boldsymbol{\Omega}_h$ and $\hat\boldsymbol{\Upsilon}^{-2}=\hat\boldsymbol{\Upsilon}_h^{-2}$ as well.

Simulations

We perform a simulation study to analyse the finite sample performance of the desparsified lasso estimator. In (ref) we compare our proposed method with unpenalized parameter of interest to the standard desparsified lasso in a sparse structural VAR. In (ref), we study our proposed method in an empirically calibrated DFM.

Sparse Structural VAR Model

Consider the structural VAR with structural shocks $\boldsymbol{\epsilon}_t$

equation[equation omitted — 319 chars of source]

where the autoregressive parameter matrices $\boldsymbol{A}_k$ are tapered Toeplitz matrices. We analyse the response of $y_t$ to the first shock $\epsilon_{1,t}$ and investigate the coverage and width of 95% point-wise confidence intervals corresponding to the impulse response parameters at horizons 1 to 10. We consider different values for the number of variables $P=\{20, 40, 100\}$ and time series length $T =\{100, 200, 500\}$ and report results for 1000 replications. Details on the DGP and implementation of the HDLP are included in Appendix (ref).

figure[figure omitted — 285 chars of source]

(ref) presents the coverage rates for the proposed desparsified lasso (blue), which leaves the impulse response parameters of interest unpenalized, in comparison to the standard desparsified lasso (red). We only display results for $P=40$, but the results for $P=20 $ and $P=100$ are very similar and included in (ref) in Appendix (ref), alongside the interval widths in (ref).

Our proposed estimator attains good coverage rates close to the nominal level of 95% for all combinations of $P$ and $T$. For both DGPs, coverage rates are usually the lowest around 70-90% for horizon 1, but revert to the nominal level for larger horizons. Good coverage at further horizons is to be expected, since the local projection coefficients will tend to become more sparse in stationary models. Importantly, considerable gains are obtained compared to the standard desparsified lasso. Especially at shorter horizons -- which are usually of most interest in practice -- coverage rates of the latter drop to 25% in the first DGP and even more in the second. We also see in (ref) (Appendix (ref)) that the interval widths are broadly similar or even smaller than the ones for the standard desparsified lasso, meaning we obtain better coverage without losing power.

The poor coverage rates of the standard desparsified lasso occur since the parameters of interest are not selected in the initial lasso regression. A simple alteration of the desparsified lasso that leaves this parameter unpenalized thus brings the coverage rates much closer to the nominal level. Note that the standard desparsified lasso has coverage exceeding our proposed estimator at further horizons; this is because the true impulse response becomes close to zero, where the lasso's bias towards zero is beneficial. This does not run counter to our conclusion, as we believe more uniform coverage over different parameter values is desirable.

Empirically Calibrated Dynamic Factor Model

We now examine the performance of our proposed estimator on an empirically calibrated dynamic factor model, inspired by lazarus2018har and li2021local. Unlike the former works which use the quarterly dataset of stock2016DFM to fit their DFM, we use the monthly FRED-MD database (mccracken2016fred) which we analyse in (ref). Implementation details on the simulation set-up are in Appendix (ref), an overview on the variables in the FRED-MD database in Appendix (ref).

We consider both a “dense” DFM where the factor loadings are obtained using principal components and a factor VAR is estimated by OLS, and a “sparse” DFM where the loadings are obtained using the WF-SOFAR estimator of uematsu2022estimation and the VAR is estimated by the lasso. The dense DFM is inherently difficult for our proposed estimator since the covariance structure of the variables is dense, whereas we assume it to be sparse in (ref), and the DGP is likely to violate the sparsity of the nodewise regression coefficients. However, bounds on the sparsity of the nodewise regressions can be obtained for sparse factor models and VARs, see Example 5 and 6 of adamek2021lasso. We therefore expect better coverage properties of our proposed estimator is this setting. The WF-SOFAR estimator seems particularly appropriate to use since uematsu2022inference find evidence of a sparse factor structure in the FRED-MD data.\footnote{The difference in sparsity between the dense and sparse DFMs is relatively limited. For details, see (ref) in Appendix (ref).}

figure[figure omitted — 170 chars of source]

We study the impulse response of Industrial Production (IP) to the Federal Funds Rate (FFR). Our HDLP method estimates (ref) with $y_t=\text{IP}_t$, $x_t=\text{FFR}_t$, the remaining 120 variables in $\boldsymbol{w}_{s,t}$, $\boldsymbol{w}_{f,t}$ empty, and 3 lags included.\footnote{Each local projection therefore has 488 regressors.} We compare the performance of our proposal to two benchmarks. The first benchmark is a factor-augmented local projection, labelled “FA-LP”. We favour FA-LP by including the true number of six factors, estimated by principal components from the 120 variables, and estimate the FA-LP by OLS. As this method matches the true DGP closely, we expect this benchmark to be a highly competitive standard. Our second benchmark is a 3-variable LP (estimated by OLS) with consumer price index as third variable as in bernanke2005measuring. This benchmark, labelled “LP", allows us to assess the gain in using high-dimensional data. All models are estimated using the R package desla (desla), taking the plug-in constant equal to 0.4 for the HDLP.

In the dense DFM, see (ref) in Appendix C, our proposed HDLP method attains coverage around 70% for the first 5 periods, then steadily improves at further horizons where it reaches nominal coverage around horizon 20 for large sample sizes. The benchmarks perform, however, considerably better as expected. They perform similarly at $T=200$ with 90% coverage, but diverge at larger sample sizes. The FA-LP model has improved coverage at larger samples, reaching nominal at $T=600$, whereas the coverage of the LP deteriorates at larger samples.

In the sparse DFM ((ref)), the performance of the HDLP considerably improves. Its coverage (top panels) is still below the benchmarks for $T=200$, stays around 85% for all horizons, but it matches or exceeds the benchmarks at larger sample sizes. The performance of the FA-LP is similar to that of the dense DFM, and the LP benchmark maintains similar coverage to the FA-LP even at larger sample sizes; the 3-variable LP is likely a better approximation of the sparse DFM than the dense DFM. Importantly though, we see across all sample sizes that the widths (bottom panels) of the FA-LP benchmark are 2-4 times larger than those of the HDLP, while the LP widths are also wider everywhere though to a lesser extent.

One should keep in mind that we set up the simulations advantageously to the FA-LP by using the true number of factors. The superior performance of the FA-LP, particularly for the dense model where our HDLP is expected to suffer, is thus hardly surprising. Yet for the sparse DFM, our HDLP method outperforms both the idealized FA-LP and the 3-variable LP with comparable or superior coverage yet considerably tighter intervals.

To conclude our simulation experiments, we investigate the sensitivity of our method to the choice of long-run covariance estimator. We compare the Newey-West (NW) estimator to the fixed-$b$ Newey-West and Equal-Weighted Cosine (EWC) estimator suggested by lazarus2018har. In Appendix (ref) we find that the first generally performs slightly worse than our default choice of Newey-West (likely due to our adaptive choice of bandwidth), whereas the second performs very well in terms of coverage but has confidence intervals around 4-6 times as wide as the ones from our NW estimator. In settings where our method has good coverage, such as the sparse DFM, EWC thus provides little benefit. We therefore proceed with our default choice of NW estimator in the remainder of the paper, but these alternatives are available in the desla package.

Structural Impulse Responses Estimated by HDLPs

We apply the desparsified lasso for HDLPs to two canonical macroeconomic applications. In (ref), we build on the work by bernanke2005measuring on monetary policy analysis, in (ref) we build on the work by ramey2018government on government spending. All analyses are performed in R using the package \href{https://CRAN.R-project.org/package=desla}{desla}.

Impulse Responses to a Shock in Monetary Policy

We consider the work by bernanke2005measuring, wherein they estimate impulse responses to a monetary policy shock identified by the federal funds rate in a factor-augmented SVAR. We use HDLPs to perform the impulse response analysis and expand the dataset of the original paper to the more recent and larger FRED-MD database McCrackenNg16, thereby considering monthly data from January 1960 to October 2008 for 122 macroeconomic variables. For a full description of the data, see Appendix (ref).

We estimate the HDLP in (ref), where as response variables $y_t$ we take the federal funds rate (FFR), industrial production (IP) and the consumer price index (CPI). As shock variable of interest $x_t$ we consider the federal funds rate (FFR). Furthermore, we use the identification strategy of bernanke2005measuring, which consists of classifying variables as “slow" in $\boldsymbol{w}_{s,t}$ or “fast" in $\boldsymbol{w}_{f,t}$, according to whether these variables can respond within the same month to an unexpected change in the FFR.\footnote{As discussed before, this strategy is equivalent to recursive (partial) identification in an SVAR, as shown in Example 1 of plagborg2021local. In fact, this application can also be seen as a high-dimensional extension of christiano2005nominal.} For variables in FRED-MD that can be directly matched to their counterparts in the DRI/McGraw Hill Basic Economics Database, we use the same slow/fast classification as bernanke2005measuring. For new variables in FRED-MD, we apply the same general classification rule; i.e., variables in the categories Prices, Output & Income, Labour Market, and Consumption are classified as slow, whereas those in the categories Interest & Exchange Rates, Money & Credit, Stock Market, and Housing as fast.\footnote{One exception in the Category Prices is the series “OILPRICEx", which we classify as fast.} We apply the transformation codes of bernanke2005measuring and McCrackenNg16 (for new variables) to remove stochastic trends. Using $K=13$ lags, the resulting HDLP consists of $N=1654$ regressors, while the time series length is $T=572$.\footnote{The exact number of variables in the model depends on which variable is taken as the response, and the effective number of observations depends on the horizon. These numbers are for the model where FFR is the response, at horizon 1.}

figure[figure omitted — 261 chars of source]

We compare the HDLP impulse responses obtained from the desparsified lasso to the ones obtained from a 3-factor FAVAR as used in bernanke2005measuring. Details about the FAVAR estimation are provided in Appendix (ref). (ref) shows the impulse responses of FFR, IP, and CPI to a shock in the FFR of a size such that the FFR has unit response at horizon 0. Since we use the first difference of IP and CPI in the model, we cumulate the impulse response which then corresponds to the response in the original (level) variables.\footnote{Cumulative impulse response functions for the HDLP are obtained by cumulating the dependent variable, i.e. taking $\sum_{\ell=0}^{h}y_{t+\ell}$ on the LHS of (ref). For the FAVAR, we take the cumulative sum of the regular impulse responses from previous horizons. To ensure that these impulse responses are comparable, we scale the FAVAR impulse responses by the response of the FFR at horizon 0. } The comparable figure in bernanke2005measuring is Figure 1.

The impulse responses from the FAVAR (bottom panel (ref)) and HDLP (top panel) are generally similar. The impulse response of the FFR peaks at horizon 1 and then steadily declines to zero, which is in line with Figure 1 in bernanke2005measuring. For the HDLP, the response stays at a higher level for a longer period than for the FAVAR. The response of IP is also in line with bernanke2005measuring, with the largest drop around horizon 20, before eventually returning to zero. Notably, the response obtained with the HDLP is considerably smaller than the one for the FAVAR. Finally for the response of CPI, we see a more pronounced price puzzle with the HDLP. The response is positive and significant for 30 months, peaking at horizon 20. The FAVAR impulse response is largely in line with bernanke2005measuring, with a small positive effect at early horizons, followed by a negative, though mainly insignificant, response after horizon 10.

Impulse Responses to a Shock in Government Spending

ramey2018government estimate impulse responses to a shock in US government spending identified based on military spending news. In this paper, we first augment the authors' main LP specification with more lags as a robustness check, before considering an extended state-dependent HDLP specification with interacting states related to unemployment and recession conditions. We use the quarterly data provided by the authors at \href{https://econweb.ucsd.edu/ vramey/research.html#govt}{econweb.ucsd.edu/$\sim$vramey} covering the period 1889Q1 to 2015Q4.

We estimate the state-dependent HDLP in (ref) for real per capita GDP and government spending as $y_t$, while the shock variable $x_t$ is the military spending news shock. We include taxes as controls in $\boldsymbol{w}_{f,t}$, while no variables are included in $\boldsymbol{w}_{s,t}$.\footnote{While ramey2018government do not include taxes in their main analysis, they mention in footnote 11 that their results were robust to their inclusion. We prefer to include them as additional controls, following the VAR of blanchard2002empirical.} The state dummy variable $I_t$ distinguishes between low and high unemployment states, with $I_t=1$ when unemployment in period $t$ is larger than 6.5%. For full details on the construction of the shock variable and the treatment of the other variables, see section II.B of ramey2018government. We use $K=40$ lags, such that the resulting HDLP consists of $N=464$ regressors, while the time series length is $T=161$.\footnote{As in (ref), these correspond to the model estimated at horizon 1.}

figure[figure omitted — 260 chars of source]

Figure (ref) shows the impulse responses to a shock in government spending. Note that the military spending news variable is scaled by GDP such that the impulse responses can be seen as a reaction to a shock of size 1% of GDP. The comparable figure in ramey2018government is Figure 5. The general shape and magnitude of these impulse responses are similar to the LP model with only 4 lags used in ramey2018government. In our analysis, the peak of the responses in the high unemployment state occurs 2-3 quarters earlier, and experiences a sharper drop after horizon 15. We also find that the impulse response is not significant at horizon 0 in the linear and low-unemployment state.

figure[figure omitted — 226 chars of source]

Next to allowing the inclusion of more lags, the HDLP framework can easily be extended to allow for $S$ states. The LP is then given by

equation*[equation* omitted — 187 chars of source]

In addition to the unemployment state dummy $I_t^{(U)}$, we consider recession dummy $I_t^{(R)}$.\footnote{As in ramey2018government, we follow the NBER classification of recession periods.} Letting them interact results in $S=4$ distinct state dummies.

(ref) displays the impulse response functions with four states. The impulse responses out of recessions resemble the state-dependent model of (ref), both in shape and magnitude, but we see a different pattern in the recession state. During recessions with low unemployment, there is a negative impact response for both government spending and GDP, which becomes insignificant after a few periods. In the high unemployment state, in contrast, we see a large, positive, and persistent effect on both variables.

Conclusion

In this paper, we propose a modified version of the desparsified lasso to estimate local projections in high dimensions. The modification is simple, namely leaving the (small number of) dynamic impulse response parameter(s) of interest unpenalized in the initial lasso regression, yet provides considerable improvements in finite sample performance in terms of better coverage. The modified desparsified lasso estimator still comes with desirable asymptotic normality and uniformly valid asymptotic inference. Finally, we show how this method performs in a simulation using an empirically calibrated dynamic factor model and in canonical macroeconomic applications where we use the high-dimensional local projections framework to estimate structural impulse responses.

Acknowledgements

We thank the co-editor and two referees for their thorough review which substantially improved the quality of the manuscript. The first and second author were financially supported by the Dutch Research Council (NWO) under grant number 452-17-010. Previous versions of this paper were presented at seminars at Aarhus University and ESSEC Business School, the NESG 2022 conference and the 2022 Maastricht Workshop on Dimensionality Reduction and Inference in High-Dimensional Time Series. We gratefully acknowledge comments by participants at these seminars and conferences. We also thank Otilia Boldea and Lenard Lieb for helpful discussions. Remaining errors are our own.