EconBase
← Back to paper

Random Subspace Local Projections

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,084 characters · 18 sections · 70 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Random Subspace Local Projections

abstractWe show how random subspace methods can be adapted to estimating local projections with many controls. Random subspace methods have their roots in the machine learning literature and are implemented by averaging over regressions estimated over different combinations of subsets of these controls. We document three key results: (i) Our approach can successfully recover the impulse response functions across Monte Carlo experiments representative of different macroeconomic settings and identification schemes. (ii) Our results suggest that random subspace methods are more accurate than other dimension reduction methods if the underlying large dataset has a factor structure similar to typical macroeconomic datasets such as FRED-MD. (iii) Our approach leads to differences in the estimated impulse response functions relative to benchmark methods when applied to two widely studied empirical applications.

{\bf Keywords:} { Local Projections, Random Subspace, Impulse Response Functions, Large datasets} \thispagestyle{empty} \setcounter{page}{1}

Introduction

Impulse response functions (IRFs) are a common tool in macroeconomics to study the dynamics of variables in response to shocks. A recent development is to estimate IRFs using local projections (LPs), developed by jorda2005estimation tenreyro2016pushing,ramey18state,swanson2020measuring. LPs are single-equation methods, and can be thought of as the direct forecast counterpart of traditional multiple equation systems, such as a vector autoregressions (VARs), which rely on iterating on a system of equations.

A parallel development in the macroeconometrics literature has been the use of large datasets. The cost of putting together a large dataset for empirical analysis, at least for U.S. macroeconomic data, is now becoming exceedingly trivial, especially with datasets such as FRED-MD and FRED-QD mccracken2020fred. Like any regression model, the choice of control variables to use in an LP forms part of the model specification. With the availability of large datasets, the dimension of this control set has become potentially large.

The contribution of our paper is to introduce random subspace methods for the estimation of IRFs using LPs. The method is simple to implement by three steps: First, we include a random subset of controls in the LP and estimate the IRF. Second, we repeat the first step many times. Finally, we take the average IRF. Through averaging over random subsets of controls, where the subsets are generated independently of the data, the variance of the IRF estimate is reduced while maintaining most of the signal of the controls.

A practical implication of our approach is that it frees up the researcher to focus on the IRFs, which is the key object of interest. Conditional on a base set of controls ensuring identification of the IRF, the random subspace method deals with the issue of appropriately including controls to improve estimation efficiency. In case a large set of controls is required for identification, the averaging across different subspaces mitigates the risk of identification bias, while avoiding the signal of the shock of interest to get swamped by the controls. This is an appealing practical feature because the coefficients on the controls are almost never the object of interest, as long as the appropriate controls are included in order to obtain accurate estimates of the IRFs.

We highlight that random subspace methods apply more naturally within the LP setting despite the now-known result that LPs and VARs recover the same IRF in population plagborg2021local. In a VAR setting, an additional control variable implies an extra equation. In the LP setting, extra variables imply just more controls within the single equation. With the potential control variables numbering into the tens or hundreds, appropriately dealing with extra, possibly extraneous, controls is arguably more manageable than estimating tens or hundreds of extra, possibly irrelevant, equations.

Our key results are as follows. First, we show that random subspace methods can recover the true IRFs under plausible factor structures and usual identification techniques for the structural shock of interest. We base this conclusion of two sets of Monte Carlo experiments, one using a real business cycle model with fiscal foresight introduced by leeper2013fiscal, and one using the dynamic factor model of stock2016dynamic. The proposed method is implemented with both SVAR identification plagborg2021local, and instrumental variable (IV) identification stock2018identification. Random subspace methods help accurately estimate IRFS using a large set of controls that exhibits different factor structures resembling macroeconomic datasets.

Second, the Monte Carlo experiments demonstrate that random subspace LPs are competitive in terms of mean squared error to a wide range of methods, including commonly used VAR methods and LPs with alternative dimension reduction methods. More specifically, the use of random subspaces improves upon the use of factors when the factor structure resembles that of U.S. macroeconomic datasets such as FRED-MD. Random subspaces are also favoured over variable selection methods, possibly in part due to the dense structure of macroeconomic data giannone2021economic, though whether our method outperforms VAR based methods can depend on the specific context.

Third, we document in two widely studied empirical applications (i.e. the effect of the macroeconomy to an identified monetary policy and technology shock) that random subspace LPs produce meaningful differences in the estimated IRFs relative to common empirical strategies. Our results are encouraging as it suggests that the random subspace LP is a simple procedure or robustness check for practitioners using LPs who may be concerned that they may have omitted relevant controls.

We link our work to two strands of the broader literature. First, although random subspace methods have their roots in the machine learning literature, researchers have recently applied them to improve forecast accuracy for economic indicators based on a large number of possible predictors koop2019bayesian,kotchoni2019macroeconomic,boot2019forecasting,pick2022multi. In general, they find strong performance of random subspace methods across different macroeconomic forecasting exercises. Since the datasets used in macroeconomic structural analysis are similar, exploring random subspace methods to estimating IRFs seems natural.

Second, various approaches have been proposed to improve the efficiency of the LP estimator barnichon19smooth,ferreira2023bayesian,lusompa2023local,ho2023averaging, which is known to have a lower bias but higher variance than VARs li2022local. One can view our approach in a similar vein. Given that proliferation of controls leads to inefficiency, researchers naturally economize on the number of controls. Random subspace methods are a form of regularization, in which averaging across subsets of controls targets the bias-variance trade-off between omitting potentially relevant information and including all available data as controls.

The remainder of the paper is organised as follows: Section (ref) provides a detailed discussion of our framework. Section (ref) summarizes how to implement our procedure in a step-by-step manner. Section (ref) uses a Monte Carlo exercise to understand how random subspace methods help in appropriately estimating IRFs. Section (ref) applies our proposed approach to two widely studied empirical applications. Section (ref) provides some concluding remarks.

Method

Consider the standard LP regressions one specifies to estimate the IRF to the variable $y$ from an exogenous one unit impulse on $x_t$:

equation[equation omitted — 116 chars of source]

where $\mu_h, \beta_h$, and ${\Phi}_h$ are projection coefficients, $\xi_{t+h}$ is the projection error, and ${W}_t$ is a vector of controls. We are interested in the response of $y_{t+h}$ with respect to an exogenous one unit impulse to $x_t$, which equals $\beta_{h}$. The problem we investigate in this paper is how one deals with the control set ${W}_t$. While one is almost never interested in the coefficients ${\Phi}_{h}$, the specification of the control set can matter to obtaining accurate estimates of $\beta_{h}$.

We briefly highlight two roles the controls play in the estimation of $\beta_{h}$. First, the set of control variables accounts for relevant determinants which are correlated with $x_t$. In this setting, failure to include the relevant controls results in omitted variable bias, and so biased estimates of the IRF. Second, even if $x_t$ is strictly exogenous, or we possess a strictly exogenous instrument, including additional controls in the projection may reduce the variation in forecasting $y_{t+h}$, which may result in a more precise estimate for $\beta_{h}$. However, in this setting, the effect is less obvious. If the included set of controls is too large, the increase in parameter uncertainty reverses possible efficiency gains, and may then increase the variance of the estimate for $\beta_{h}$.

With large macroeconomic datasets like FRED-MD, empirical work in macroeconomics now has access to over a hundred variables that may serve as potential controls. This leads to the familiar trade-off for practitioners where omitting relevant controls results in bias, but including irrelevant variables leads to an increase in variance. In cases where a variable is known to be irrelevant, such a variable should be omitted, although in most empirical settings it is not obvious whether a variable is irrelevant or not. This is the setting that we propose solving through random subspace local projections (RSLP): we alleviate the issue of specifying ${W}_t$, mindful that a practitioner is often only interested in $\beta_{h}$.

Random Subspace Local Projections (RSLP)

Instead of a generic set of controls ${W}_t$, we make a distinction between the $p_V\times 1$ vector ${V}_{t}$ of variables which are considered essential controls, and the $p_G\times 1$ vector ${G}_t$ of possibly relevant controls. Rewriting (ref) in terms of these essential and possible controls results in

align[align omitted — 111 chars of source]

where ${\Theta}_{h}$ and ${\Psi}_h$ are the projection coefficients of ${V}_t$ and ${G}_t$, respectively. While nothing precludes specifying ${V}_t$ as an empty set, most macroeconomic applications would a priori treat some variables as essential. For example, the usual monetary policy shock application would include at least some measure of real activity, inflation, and the interest rate. The identification strategy for $\beta_h$ may also inform which variables have to be included in ${V}_t$, a matter which we discuss in detail in Section (ref).\footnote{Even in cases where one possesses a strictly exogenous and valid instrument --and no controls are required for the LP estimator to be consistent-- the inclusion of controls known to be relevant in ${V}_t$ may improve the accuracy of the estimate for $\beta_h$.}

Some applications may only provide essential control categories, but leave the precise variable unspecified. For example, even if the model should feature an interest rate variable, there are many such variables in datasets like FRED-MD. Our empirical examples select one of the interest variables to be included in $V_t$ in this case, and include the rest in ${G}_t$. Alternatively, Appendix (ref) extends our approach to using prior knowledge on which categories are essential, without having to select a specific variable from those categories.

The number of possibly relevant controls in ${G}_t$ is potentially large as previously motivated. Instead of estimating $\beta_{h}$ in (ref), a form of dimension reduction is usually applied to ${G}_t$. Consider the linear projection from (ref) with a $k \times p_G$ compression matrix ${R}^{(j)}$ indexed by $j$, where $k\le p_G$:

align[align omitted — 145 chars of source]

where ${\Gamma}_{h}^{(j)}$ is a $k$-dimensional vector of projection coefficients, instead of the $p_G$-dimensional vector of projection coefficients ${\Psi}_h$ in (ref). The construction of the compression matrix can be data-driven. For instance, variable selection methods use data to estimate a selection matrix for ${R}^{(j)}$. Factor-augmented models take ${R}^{(j)}$ as the matrix of the principal component loadings corresponding to the $k$ largest eigenvalues from the sample covariance matrix of ${G}_t$.

The random subspace approach to dimension reduction is to generate the elements of ${R}^{(j)}$ from a probability distribution that is independent of the data. More precisely, $R^{(j)}$ is sampled from a uniform distribution across all combinations of the $p_G$ available predictors of size $k$. It follows that ${R}^{(j)}$ is a random subset selection matrix which randomly selects a subset of $k$ predictors out of the $p_G$ predictors. For instance, if there are 5 possible predictors and we wanted to choose 3, $p_G=5$ and $k=3$, and one possible draw $j$ could be the following:

align*[align* omitted — 120 chars of source]

Conditional on a draw ${R}^{(j)}$, $\beta_h^{(j)}$ in (ref) can be estimated using least squares. The random subspace estimate for $\beta_h^{(j)}$ is constructed by averaging over the least squares estimates $\hat{\beta}_h^{(j)}$ corresponding to different draws ${R}^{(j)}$, with $j=1,\dots,n_R$:

align[align omitted — 100 chars of source]

Through averaging, random subspace methods reduce the variance of the LP estimate, while it maintains most of the signal of the controls. We note that one could combine these subspaces using weights reflecting the performance of each subspace, instead of equal weights. However, estimation of these optimal weights introduces additional variance, and hence a simple average is generally hard to beat, as shown by elliott2013handbook.\footnote{In Appendix (ref), we use a weighting scheme based on the Bayesian information criterion as discussed by elliott2013complete in our empirical examples, which indeed results in little differences relative to equal weighting.}

Implementing the random subspace LPs requires a choice for the subspace dimension $k$. We find that setting $k=50$ is a reasonable choice in our Monte Carlo experiments and empirical applications. While one could use an information criterion to choose the dimension, as we discuss in Appendix (ref), this may not work well in macroeconomic settings with a large number of controls and a small sample size. Sensitivity analyses reported in Appendix (ref) suggest that the optimal subspace dimension is often in the 40-60 range. Hence, we proceed with an empirical strategy that uses $k=50$, and subsequently check for the robustness of the results.

Random subspace methods have shown to improve forecast accuracy for economic indicators based on a large number of possible predictors koop2019bayesian,kotchoni2019macroeconomic,boot2019forecasting,pick2022multi. The relative effectiveness of the random subspace approach may be explained by three features of macroeconomic data. First, macroeconomic variables are often highly correlated. Variables that are irrelevant when all relevant controls are included, are often correlated with omitted relevant controls in a random subset. Hence, the use of random subsets results in a small bias while it reduces the variance. Second, most macroeconomic data is characterized by a low frequency and low signal. As a result, the selection of the controls in ${G}_t$, or the derived factors, may be subject to substantial uncertainty. Third, giannone2021economic, among others, find that often many variables are relevant in explaining macroeconomic quantities of interest. This makes it even more challenging to find an accurate low-dimensional representation of the data by variable selection methods. boot2020subspace discuss random subspace methods and its theoretical justification in the context of macroeconomic forecasting in more detail.

Structural identification within RSLP

The exposition so far discusses how one can use random subspace methods to estimate IRFs in an LP setting. We now connect our approach to recovering the object of interest, namely the IRF from an underlying structural model.

Define ${w}_t$ as an $n \times 1$ vector of macroeconomic variables, in which both the variable of interest $y_t$ and the variable $x_t$ to which we are introducing an exogenous impulse in (ref) are included. The variables in ${w}_t$ are driven by a $m \times 1$ vector of independent and identically distributed (i.i.d.) shocks ${\epsilon}_t$:

equation[equation omitted — 107 chars of source]

with lag polynomial ${\Theta}(L) \equiv \sum_{j=0}^{\infty} {\Theta}_j L^j$, and $n \times m$ coefficient matrices ${\Theta}_j$.

Suppose, without loss of generality, that we are interested in the effect of the first shock $\epsilon_{1t}$, and $y_t$ and $x_t$ are respectively the $i^{th}$ and $j^{th}$ variable in the vector ${w}_t$. After normalizing $\epsilon_{1t}$ to imply a unit increase in $x_t$,\footnote{This normalization is without loss of generality since shocks are unobserved, so the variance and sign of shocks are ultimately normalized. Our choice of normalization simplifies the exposition from needing to introduce a parameter to link $\epsilon_{1t}$ to $x_t$.} we can write $x_t$ as a linear combination of shocks:

align[align omitted — 102 chars of source]

where $\epsilon_{2:m, t}\equiv(\epsilon_{2t},...,\epsilon_{mt})'$ and $f(.)$ is a linear combination of its argument. Leading the $i^{th}$ equation of (ref), which describes the dynamics of $y_t$, by $h$ periods, and subsequently rearranging (ref) to substitute out for $\epsilon_{1t}$, we obtain

equation[equation omitted — 154 chars of source]

where the argument in $f(\cdot)$ now subsumes $y_{t+h}$ being a linear function of past and future shocks and $f(\epsilon_{2:m, t}, \epsilon_{t-1},\epsilon_{t-2},...)$ from (ref). The object of interest is $\theta^{i1}_h$, as it is the IRF from (ref). Based on the form in which (ref) is written it would be analogous to $\beta_h$ in (ref). Due to endogeneity, (ref) cannot be consistently estimated since both $x_t$ and $y_{t+h}$ are a function of {all} past and current shocks. More precisely, in our context, $x_t$ is not only correlated with $\epsilon_{1t}$ but also with other shocks that affect $y_{t+h}$.

Identification through instruments or external variation

The IRF can be estimated using two-stage least squares and an instrument $z_t$ that satisfies the following conditions, termed Condition LP-IV by stock2018identification:

enumerate• {$\mathbb{E}\left[\epsilon_{1t}z_t \right] \neq 0$ (relevance)}, • {$\mathbb{E}\left[\epsilon_{2:m, t}z_t \right] = {0}$ (contemporaneous exogeneity)}, • {$\mathbb{E}\left[\epsilon_{t+j}z_t \right] = {0},\, j \neq 0$ (lead-lag exogeneity)}.

With a strictly exogenous instrument satisfying the LP-IV conditions, the instrument and the shocks are assumed to be uncorrelated contemporaneously and at all lags, and hence the IRF can be estimated consistently. Directly observing the shock is equivalent to observing a $z_t$ which is perfectly correlated with $\epsilon_{1t}$. In this case, Condition LP-IV allows one to directly regress the shock on $y_{t+h}$, similar as in approaches such as romer2004new. In these settings, the inclusion of controls is to the extent that they may lead to efficiency in finite samples, but they are not required for the identification of the IRF.

Condition LP-IV is a strong assumption but can be relaxed by including controls:

eqnarray[eqnarray omitted — 318 chars of source]

where $u_t^{\perp} = u_t - Proj(u_t\mid {W}_t)$ for some variable $u$. This is a more relevant setting empirically, given widely used instruments for monetary policy shocks have been shown to be either contaminated by other shocks miranda2018transmission, or possibly forecastable by other macro variables ramey2016macroeconomic. These suggest violations of strict exogeneity. Controls that account for this information can aid in satisfying conditional exogeneity, and so help constructing a valid instrument $z_t^{\perp}$.\footnote{This argument is analogous to what lloyd_manuel2401 label as the one-step approach, where one uses controls in the first stage to render the instrument conditionally exogenous.} If there are particular controls that one knows to help the instrument to satisfy conditional exogeneity, they should be treated as essential controls. In addition, RSLP may efficiently account for the information in the possibly relevant controls so that the conditions in (ref) and (ref) are satisfied.

We thus extend random subspace methods to the two-stage least squares settings where the first and second stage regressions are

align[align omitted — 299 chars of source]

where ${\Lambda}^{(j)}$ and ${\Upsilon}^{(j)}$ are the projection coefficients from the first stage regression, $\hat{x}_t$ is the fitted value from (ref), and the estimates for $\beta_{h}^{(j)}$ in (ref) are used in (ref) to average across random draws for selection matrices.

Implementing SVAR identification

plagborg2021local show that SVAR or system-based identification (i.e. recursive identification, long-run restrictions, and sign restrictions) can be implemented using LPs. These identification schemes require SVAR invertibility, which allows one to map a VAR in $w_t$ to the form in (ref) by inverting the VAR lag polynomials. Under invertibility there is a linear combination of $w_t$ and its lags which accounts for $f(\epsilon_{2:m, t}, \epsilon_{t-1},\epsilon_{t-2},...)$ in (ref), so that what gets used in place of $x_t$ in (ref) identifies the shock of interest $\epsilon_{1t}$. If not all relevant information is included, identification fails due to not satisfying invertibility. The literature has also long recognized the link between invertibility and the inclusion of all relevant information hansen2019two,fernandez2007abcs,stock2018identification.

Where a standard LP has a VAR representation estimating the same IRF in population, estimating an IRF by RSLP is analogous to averaging over the IRFs from VARs with different sets of variables, but the same identification strategy.\footnote{Note that we will be applying the same sign, long, or short run restriction across different subspaces. Therefore, all that is changing is the information set, or the set of controls, one is using to apply the identification scheme. In this regard, the approach has connections to the approach of ho2023averaging, although they allow the identifying restrictions to differ across the different models that they average over.} In cases where the essential controls in $V_t$ are sufficient for invertibility, RSLP may help reducing the variance in the IRF estimates by regularizing over the information from the rest of the dataset. In cases where the possibly relevant controls in $G_t$ are required to satisfy invertibility, and thus play a role with identification, the averaging over different subspaces of $G_t$ implicitly takes uncertainty about the exact identifying assumptions into account. The averaging over many different information sets is also related to the estimation of IRFs by factor models or large Bayesian VARs, such as respectively those by bernanke2005measuring and banbura2008large. These methods may be viewed as alternative approaches to fulfill the invertibility condition by applying dimension reduction to large information sets.

To illustrate how SVAR identification may be implemented within an RSLP, we consider a recursive identification scheme from the seminal work by christiano1999monetary. They identify a monetary policy shock by restricting the contemporaneous effect of “fast moving” variables on the Federal funds rate to be zero, and the contemporaneous effect of monetary policy shocks on “slow moving” variables to be zero. This suggests that contemporanous values of the “slow moving” variables, such as real activity and price variables, have to be included in the LP. The number of available real activity and price variables in modern macroeconomic datasets is large, and the identification scheme does not provide guidance on which specific variables to select.\footnote{For example, it is known that researchers have used industrial production, unemployment, employment, or real GDP, amongst others, to reflect real activity. While the motivation is to include at least one real activity variable, there is little guidance of which one to include. From this perspective, one could argue that researchers aim to control for the same information even if they disagree on which precise variable reflects that information. } By including a small number of variables in $V_t$, such as industrial production and inflation, and the remaining available “slow moving” variables in $G_t$, one can view RSLP as a procedure to balance the risk of non-invertibility with overfitting.

Error bands with RSLP

We briefly discuss how to construct error bands for the IRF estimated by RSLP. A key challenge is that the random subspace literature, like most machine learning methods, provides little guidance on how to quantify estimation uncertainty.

The standard error of a standard LP can be estimated with a block bootstrap to account for serially correlated errors kilian2011reliable,inoue2023significance.\footnote{We note that if one were certain that the lags of the correct variables are included as essential controls, it may be possible to obviate the need to account for serial correlation montiel2021local and just consider the cross correlation across the regressions.} Since RSLP is constructed by averaging across LP regressions, the bootstrap should not only preserve correlation across residuals in each regression, but also the correlation across the regressions. Therefore, a natural choice would be a block bootstrap, using the same block of residuals across regressions to also preserve the correlation across subspace estimates. Although estimating a single RSLP requires $n_R$ least squares regressions, which would at most take a few seconds in a modern macro context, the bootstrap may amount to substantial computational costs. Given each horizon is bootstrapped separately, an IRF for one variable would require bootstrapping $n_R \times (H+1)$ regressions. The computation time of this procedure can be alleviated by parallelizing the computation across subspaces.

Since the block bootstrap for RSLP is computationally expensive, we also discuss an alternative approach derived from buckland1997model. This approach relies on the assumption that there is perfect correlation across the IRF estimates from different subspaces. Although the correlation across subspaces is high in practice, this assumption is likely too strict, and therefore the bands will be conservative. The error bands of standard LP bands are already known to be wide, although RSLP may alleviate this issue by efficiently incorporating controls that help explain variation in the variable of interest. The implementation details of both the bootstrap and the buckland1997model approach are deferred to Appendix (ref).

Implementing RSLP

We briefly summarize the preceding discussion to outline how RSLP can be implemented:

enumerate• {Make a distinction between essential controls, $V_t$, and the controls that the random subspace method will be applied to, $G_t$. This distinction specifies the number of possible covariates that each random subspace matrix will select from, $p_G$. } • {Choose a subspace dimension, $k$. The subspace dimension dictates how many covariates are included in each subspace regression. In practice, we find a subspace dimension of 50 to work well with U.S. macroeconomic data.} • {Randomly draw from the $p_G$ possible controls, assigning equal probability that any of these $p_G$ controls can be chosen. Create a matrix indexed by $j$, $R^{(j)}$ of dimension $k \times P_G$, which acts as a selector matrix that sets the appropriate element to 1 if a control is chosen, and zero otherwise.} • {For each horizon $h \in \lbrace 0, 1, \hdots, H \rbrace $, estimate ${y}_{t+h}=\mu_h^{(j)}+ \beta_h^{(j)}x_t+{\Theta}_{h}^{(j)} {V}_{t}+ {\Gamma}_{h}^{(j)}{R}^{(j)}{G}_t+\xi_{t+h}^{(j)}$. Depending on the identification strategy, $x_t$ may need to be constructed with a first stage regression. This regression may be similar to (ref) if one used an external instrument, or is specified using system-based identification restrictions such as long-run or sign restrictions.\footnote{See plagborg2021local for a general discussion on how to implement system-based identification in an LP setting. alpanbda2021 discuss sign restrictions, and Appendix (ref) long-run restrictions in the spirit of blanchard1988dynamic in LPs.} Because the first stage regression also requires choosing controls, the $R^{(j)}$'s will also be used in the first stage.} • {Repeat steps 3 and 4 a large number of times. In practice, we find a number of subspace draws $n_R = 1000$ to be sufficiently large that results are stable. The estimated IRF for variable $y$ at horizon $h$ to the identified shock of interest, can be obtained via averaging over the $n_R$ draws: $\beta_h = \frac{1}{n_R}\sum_{j=1}^{n_R} \beta_h^{(j)}$.}

Monte Carlo Experiments

As an illustration of the RSLP procedure, we consider Monte Carlo experiments based on the real business cycle model with fiscal foresight by leeper2013fiscal. The phenomenon of fiscal foresight occurs when at time $t$ agents know the tax rate they will face at time $t+h$. The model includes income taxes, inelastic labor supply, and full capital depreciation. The log-linearized equilibrium condition for capital is given by\footnote{See leeper2013fiscal, page 1118, Equation (4).}

align[align omitted — 139 chars of source]

where $k_t$ and $\hat{\tau}_t$ are respectively the percentage deviations from the steady state capital and tax rate, $\tau$ is the steady state value of the tax rate, $\theta$ and $\alpha$ are parameters satisfying the inequalities $0<\theta<1$ and $0<\alpha<1$, and $u_{at}$ is an i.i.d. technology shock. The tax rule is $\hat{\tau}_{t+h}=u_{\tau t}$, where $u_{\tau t}$ is an i.i.d. tax shock. Allowing $h=2$ (a two-period foresight), (ref) becomes

align[align omitted — 100 chars of source]

where $\kappa=\tau(1-\theta)/(1-\tau)$. The structural moving average representation equals

align[align omitted — 250 chars of source]

To parametrize the above, we follow leeper2013fiscal and forni2014sufficient by setting $\theta=0.2673$, $\tau=0.25$, $\alpha=0.36$, and $u_{\tau t}$ and $u_{at}$ to be from independent standard Normal distributions.

Information sets and identification schemes

The econometrician's objective is to estimate the IRFs of both taxes and capital to a tax shock. Due to fiscal foresight, there is insufficient information to recover the IRFs by just observing capital and the tax rate at time $t$. The Monte Carlo experiments consider different ways of extending the information set of the econometrician.

\paragraph{Strictly exogenous instrument} First, we allow the econometrician to observe a strictly exogenous instrument together with additional series that can aid in recovering the IRFs. The instrument $z_t$ for the tax shock $u_{\tau t}$ is generated as

align[align omitted — 129 chars of source]

where the shocks $\nu_{1,t}$ and $\nu_{2,t}$ are both generated from a $N(0,4)$ independent from the structural model. The instrument identifies the IRF without controls. However, being able to account for variation in $z_t$ due to $\nu_{1t-1}$ and $\nu_{2t-1}$ can lead to efficiency gains. To this end, 100 informational variables $y_{it}^*$ are generated as

align[align omitted — 156 chars of source]

where $b_i$ is a Bernoulli random variable assuming value 1 with probability 0.1.

\paragraph{Conditionally exogenous instrument} Second, we allow the econometrician to observe an instrument for the tax shock which is only exogenous conditional on a set of informational variables. The instrument $z_t$ is generated as

align[align omitted — 131 chars of source]

where $u_{a t-1}$ and $u_{\tau t-1}$ are structural shocks. Failure to control for these lagged shocks will lead to a violation of lead-lag exogeneity. To be able to recover the IRFs with this instrument, we generate 100 informational variables $y_{it}^*$ as

align[align omitted — 158 chars of source]

The settings with strictly and conditionally exogenous instruments will help to understand the extent to which RSLP can achieve either reduction in estimation variance or identification bias.

\paragraph{SVAR identification} Third, we assume that the econometrician knows that the technology shock has no cumulative effect on the tax rate, and has access to the informational series from (ref). As shown by leeper2013fiscal, the model implied by (ref) does not admit an invertible SVAR representation in taxes and capital as the model is missing the two period ahead tax rate (or the current tax shock) in the information set. Since the informational variables contain this information, including these variables admits an invertible VAR representation. It follows from plagborg2021local that subsequently applying SVAR identification with this additional information in an LP can recover the IRF.\footnote{We note that forni2014sufficient already demonstrate that one can use SVAR identification to recover the IRF in this setting by using a FAVAR to account for this information. We impose the SVAR identification restriction in an LP setting by using the procedure presented by plagborg2021local. We leave the implementation details to Appendix (ref).}

\paragraph{Controls with a strong and weakened factor structure} To examine the role of the strength of the factor structure of the information series in recovering the IRFs, we consider two settings that differentiate the factor structure of the informational variables in (ref) and (ref):

eqnarray[eqnarray omitted — 145 chars of source]

The strong case mimics a structure that one would probably encounter in a panel where these individual series have a high level of comovement, such as cross country asset prices and interest rates miranda2020us, or commodity prices west2014factor,alquist2020commodity. The weak case is akin to the sort one would expect to encounter when using U.S. macroeconomic data: Appendix (ref) shows that the first two factors of both the FRED-MD and FRED-QD dataset explain a similar amount of variation as in the weak case.

We run each of the three identification schemes with both the strong and weak cases. This amounts to six Monte Carlo experiments, each simulating 1000 artificial datasets of 200 observations for capital and the tax rate according to (ref). The experiments corresponding to different information sets and identification schemes accompany each of these artificial datasets with: an exogenous instrument and informational variables generated from (ref) and (ref); a conditionally exogenous instrument and a set of informational variables from (ref) and (ref); informational variables from (ref) within SVAR identification.

RSLP can recover the true impulse response functions

To understand if RSLP can recover the IRF from the DGP, we study the expectation of the estimated IRFs. For each artifical dataset, RSLP constructs the IRF following the procedure in Section (ref) with 1000 draws of the selection matrix ${R}^{(j)}$ and ${G}_t$ equal to the first lag of $y^*_t$. To investigate whether RSLP is a valid approach to accounting for the omitted information, we estimate LPs without ${G}_t$, which we label LP with a base set of controls. Both specifications include the two lags of tax rate and capital in ${V}_t$, together with two lags of the instrument when using IV identification. Appendix (ref) elaborates on the specification of the LPs when using SVAR identification.

Figure (ref) presents the IRFs of both capital and tax rate to a tax shock from both the IV and SVAR identification schemes, under both the strong and weak case. We also plot IRFs estimated via LP without the additional controls from the generated informational series, and the true IRFs. The plotted estimated IRFs are taken as the average across the Monte Carlo replications, so deviations from the true IRF represent bias.

figure[figure omitted — 956 chars of source]

Since the tax shock occurs in period 2, the true IRF of the tax rate sees a one off spike in period $t+2$. Capital falls for two periods, then adjusts back in a monotonic fashion. Panel (a) in Figure (ref) shows that with a strictly exogenous instrument the true IRFs are recovered. This result holds even without information on the controls, although the dot-dashed orange line shows some more variation around the true IRF. Now suppose we estimate the IRF without controls in Panel (b) and (c). The dot-dashed orange line shows that we will miss the timing of the tax rate, and as a result, also miss the response of capital. LP with a base set of controls and a conditionally exogenous instrument in Panel (b) is biased given the instrument is invalid due to not being exogenous to the lagged shocks. Panel (c) is an illustration of a fiscal foresight issue, where information on only the tax rate and capital is insufficient to recover the true IRF, which has been previously documented leeper2013fiscal,forni2014sufficient.

Our proposed approach is able to largely recover the true IRF. In particular, RSLP is able to estimate the tax rate spiking in period 2, and recover the dynamics of the response to capital. While our proposed method is close to being unbiased in the strong case, the bias is still minimal under the weak case with either IV or SVAR identification.

Comparisons relative to other methods

This section compares RSLP relative to widely used existing methods for constructing IRFs, and to LPs employing alternative dimension reduction methods. More specifically, we consider FAVARs and factor augmented local projections (FALPs), small and medium Bayesian VARs (BVAR),\footnote{The small BVAR only includes the base set of variables; taxes and capital. The medium BVAR considers 20 variables, with another 18 informational variables in addition to the base set. Extant work suggests that with the standard Minnesota type shrinkage and natural conjugate priors, a 20 variable BVAR often produces similar results relative to large BVARs of over hundred variables banbura2008large,Morley_Wong2020. In our setting, a 20 variable BVAR suffices to make our point.} the Bayesian local projection (BLP) of ferreira2023bayesian, and LPs exploiting sparsity with either a LASSO penalty or the variable selection method OCMT proposed by chudik2018one. Finally, we consider the LP with a base set of controls (Base). Implementation details of these methods are deferred to Appendix (ref).

Table (ref) shows the root mean squared error (RMSE) of the IRFs estimated by each method relative to the RMSE of the IRFs estimated by RSLP.\footnote{Note that there are seven horizons ($h=0,\dots,6$) in the IRFs, across which we average to calculate the RMSE: $\sqrt{\frac{1}{7}\sum_{h=0}^{6}\frac{1}{1000}\sum_{i=1}^{1000}(\hat{\beta}_{hi}-\beta_{h,\text{true}})^2}$, where $\hat{\beta}_{hi}$ is the estimated IRF at horizon $h$ in the $i$-th replication and $\beta_{h,\text{true}}$ is the true IRF at horizon $h$.} Values above one favour RSLP. These values are reported for the six experiments with the real business cycle model with fiscal foresight, which we refer to as the baseline experiments. We also compare the methods in experiments with a dynamic factor model (DFM) in the spirit of stock2016dynamic. Appendix (ref) discusses the DFM experiments in detail. Here we summarize broad observations from our baseline experiments, noting that the DFM experiments have similar results unless otherwise stated.

table[table omitted — 2,950 chars of source]

\paragraph{RSLP improves upon factor models when the factor structure is weakened} We focus our initial comparisons relative to factor models given its established tradition in the macroeconomics literature. We first discuss the FAVAR, since this model is shown to be able to recover the IRF in the baseline DGP with SVAR identification forni2014sufficient.\footnote{The FAVAR augments a bivariate VAR system containing taxes and capital with the first two principal components of the informational variables. The IV identification case uses the instrument as per the approach by mertens2013dynamic and gertler2015monetary. The SVAR identification case identifies the shock by assuming there is no cumulative effect of the technology shock on the tax shock.} Table (ref) shows that in the strong factor case in the conditionally exogenous instrument or SVAR identification schemes, the FAVAR models have a lower or similar RMSE relative to our RSLP approach.\footnote{Note that we focus the comparison of the FAVAR to the RSLP to these two cases as the FAVAR is invertible. The FAVAR is non-invertible in the strictly exogenous instrument case in our baseline DGP as the controls capture noise in the instrument, explaining its underperformance.} However, the conclusions flip when we consider the weak cases, where RSLP outperforms the FAVAR, at times by substantial margins.

The comparison between FAVAR with RSLP may be driven by differences between a VAR and LP, or between using factors or random subspaces. Comparing RSLP with FALP isolates the latter. Although the margin of improvement decreases, RSLP still outperforms FALP in the weak case. This result suggests that the outperformance of the RSLP relative to FAVAR is both driven by the VAR structure and possible issues related to the factors not being able to account for all relevant information. We conclude that with a strong factor structure, such as cross-country data on asset prices and interest rates, factor models do well. RSLP becomes useful with a weaker factor structure, resembling that of a macroeconomic dataset such as the FRED-MD. The DFM experiments also show results in favour of RSLP over FALP. However, since the DGP in those experiments is a closer approximation to a FAVAR, FAVAR is outperforming both.

\paragraph{RSLP does well compared to other local projection methods} RSLP is competitive compared to applying other dimension reduction methods to LPs. Together with FALP, RSLP often performs best in terms of RMSE. Hence, depending on the strength of the underlying factor structure in the data, FALP also appears to be a reasonable method. Variable selection with LASSO does not appear to outperform the FALP or RSLP. OCMT performs well with a conditionally exogenous instrument or SVAR identification, but is less accurate when the instrument is strictly exogenous. Note that, without any guidance from the literature on how to implement variable selection methods within LPs, improvement may be made in selecting the tuning parameters in these methods.

Comparing the base LP with the LPs using dimension reduction in the experiments with a strictly exogenous instrument shows that using a large set of controls often leads to efficiency gains. All the LP methods that use large information sets do similarly or better than an LP with a standard set of controls. Since the controls are not required to identify the shock with a strictly exogenous instrument, any improvements in RMSE are driven by efficiency gains from inclusion of the potentially relevant controls.

In the DFM experiments, both variable selection methods are outperformed by FALP and RSLP. Although BLP is outperformed in the baseline experiments, it performs well in the DFM experiments. The BLP is shrinking towards a BVAR. While a BVAR is a poor approximation in the baseline DGP due to information insufficiency, a BVAR is a reasonable approximation to the DFM DGP, and so correspondingly, the BLP does much better in the latter setting. Therefore, while the BLP is a valuable addition to the empirical toolkit, our exercise also shows that how well the BLP does in practice depends on the underlying BVAR that one chooses to shrink towards.

\paragraph{RSLP may compare favourably relative to VARs, depending on the DGP} Relative to VAR methods other than FAVAR, RSLP performs well across the baseline experiments. Within these DGPs, the small BVAR suffers from information insufficiency. The medium BVAR is technically information sufficient, and hence the loss of accuracy relative to RSLP is driven by two features. First, the VARs include equations for all included variables and are therefore vastly overparametrized. Second, to prevent the overparametrization to lead to substantial estimation uncertainty, the VAR methods apply shrinkage to all coefficients. This shrinkage does not discriminate between essential and potentially relevant controls, which may induce a bias.

Note that the results in Table (ref) are based on mean squared errors averaged across horizons. It is well known that the relative accuracy of VARs improves when the horizons increase. Appendix (ref) shows that this is also the case in our baseline experiments. For instance, the relative performance of RSLP and FALP stays relatively constant across horizons, while the relative accuracy of FAVAR to RSLP improves with larger horizons.

While the baseline experiments show that RSLP can be a serious competitor relative to VAR methods, the DFM experiments show that this is not a general result. In the DFM experiments, the BVAR produces accurate IRFs for most variables considered, and also BLP outperforms the LP methods for most variables. The medium BVAR performs better than RSLP across all variables. These results suggest that VARs are a better approximation to the DFM than to the baseline DGP. The DFM DGP generates the data from factors without distinguishing essential variables, making the application of dimension reduction to all variables appropriate. This may benefit VAR models by reducing both the variance due to potential overparametrization and bias due to the shrinkage priors.

Take-aways

We summarize two key take-aways from our Monte Carlo exercise. First, RSLP is capable of recovering the true IRF, which at least provides prima facie evidence that it is a viable approach for applied work. In particular, RSLP is close to being unbiased across a range of different Monte Carlo experiments. Second, relative to a set of alternative benchmarks RSLP at least performs competitively, suggesting it to be an appropriate addition to the standard toolkit for applied macroeconomists.

Empirical applications

We use RSLP to estimate the dynamic responses to technology and monetary policy shocks. Both applications can be traced to a broad empirical literature, where there are suggestions that baseline specifications with minimal controls are insufficient. Hence, these applications are natural settings to understand whether random subspace is a useful method to incorporate information from a large set of controls.

Technology shock application

Since gali1999technology, a strand of the SVAR literature has investigated the impact of technology shocks on a variety of macroeconomic variables, with a focus on labor market variables francis2005technology,forni2014sufficient,barnichon2010productivity. Keeping in the spirit of gali1999technology, the technology shock is identified as being the only shock that has a long-run impact on labor productivity.\footnote{For settings that we are working with LPs instead of SVARs, we implement the long-run identification restrictions as suggested by plagborg2021local. Implementation requires one to nominate a horizon at which all short-run effects of the shock are expected to dissipate. We set this horizon to 3 years. Appendix (ref) discusses further implementation details.}

Mirroring the baseline bivariate VAR in the broader literature, our baseline set of controls includes four lags of both the growth rate of labor productivity and unemployment in ${V}_t$ as we use quarterly data. The set of possible controls in ${G}_t$ includes 127 macroeconomic time series specified at the quarterly frequency, which includes 117 series from the FRED-QD database mccracken2020fred, 6 total factor productivity series from Fernald's website,\footnote{https://www.johnfernald.net/TFP} and 4 consumer confidence indicators from the Michigan Survey. We consider the first lags for these variables, which gives us a set of 127 possible control variables. Our sample is from 1960Q1 to 2019Q4.

Monetary policy shock application

This application features a baseline set of controls that echoes the external-VAR of gertler2015monetary. As the model is monthly, our baseline set of controls includes 12 lags of the log difference of CPI, log difference of industrial production (IP), the excess bond premium by gilchrist2012credit, and the 1-year government bond rate in ${V}_t$.

Our possible control set ${G}_t$ includes 111 FRED-MD series, and we consider the first lag of these variables. The sample spans January 1990 to June 2012 to match up with the time-span of the instrument. We use the high frequency surprises around announcements by gertler2015monetary as an instrument to identify the effect of monetary policy shocks.

Results

Figure (ref) presents the estimated IRFs from our two applications.\footnote{Note that the IRFs are on the level of labor productivity, CPI and industrial production index. Hence, we follow the usual practice to specify the left-hand side variable in the LP as $y_{t+h}-y_{t-1}$ stock2018identification.} The RSLPs are based on 1000 draws of the selection matrix, a subspace dimension equal to 50,\footnote{Appendix (ref) investigates the robustness of the results with respect to the subspace dimension.} and accompanied by an approximate 90% bootstrap interval constructed as in Appendix (ref). The IRFs are also estimated by the benchmark methods discussed in Section (ref). The implementation details of these methods in both applications are deferred to Appendix (ref).

figure[figure omitted — 1,133 chars of source]

For the technology shock application, we present the response of labor productivity and the unemployment rate to a technology shock which raises labour productivity by $0.25\%$, which is equivalent to a one standard deviation shock in the FAVAR model. Using RSLP, labor productivity increases and the unemployment rate decreases in response to an expansionary technology shock. These results are consistent with the predictions of the real business cycle model and those shown by, for example, christiano2003happens.

For the monetary policy shock application, we present the response of CPI and industrial production to a contractionary monetary policy shock which increases the 1-year government bond rate by a 100 basis points. Using RSLP, both CPI and industrial production fall in response to a contractionary monetary policy shock, consistent with what one expects from a standard monetary policy shock.\footnote{The information effect does not seem to produce puzzles here. We nonetheless note that our approach may be well-suited for dealing with informational effects, as variables representing central bank information can be used as additional controls. We experimented with using the 27 variables representing forecasts and forecast revisions constructed by miranda2018transmission as part of a larger set of possible controls for our monetary policy application, and obtained almost identical results, which suggests that our large information set already accounted for possible informational effects.}

Comparing relative to other methods, we find that while the direction of the response of the variables is largely consistent across all the different methods, our proposed RSLP also leads to noticeable differences relative to the other approaches. These differences can be qualitatively meaningful, as the estimated responses from the different methods can be outside the bootstrapped 90% intervals for RSLP.

More specifically, we find that the relative differences across LP methods vary across the two applications. FALP does not differ much from the base LP in the technology application, but produces IRFs different from all other methods with the monetary policy shock. The IRFs of the LPs with variable selection are similar to the ones from RSLP in the technology application, but do not pick up any information from the controls in the monetary policy shock example. From the comparison with the VAR methods, we find that FAVAR and the medium BVAR are closer to RSLP than BLP and the small BVAR across both applications, indicating that they pick up some signal from the controls. The BLP can be quite similar to the estimated response from the small BVAR, which highlights the finding from our Monte Carlo experiments that the performance of BLP can be driven by the choice of the BVAR one chooses to shrink towards.

Overall, our empirical exercise suggests that RSLP does lead to reasonable empirical results, especially when compared with other methods and the broader literature. Consistent with our Monte Carlo experiments, we are encouraged that our approach can be a useful addition to the applied macroeconomist's toolkit.

Conclusion

We show how one can apply a dimension reduction technique traditionally used in machine learning to LPs in order to estimate IRFs with many controls. This random subspace method is simple to implement and it basically contains 3 steps: Step 1: take a random draw; Step 2: do it many times; and Step 3: average. We have shown that our approach is a plausible addition to the toolkit as it can recover the true IRFs in settings encountered in macroeconomics.

It is worth stressing that while our proposed approach is a plausible empirical strategy, it does not compete to supplant any of the recent innovations in the broader development of LP estimation. To highlight some important developments in the LP literature, barnichon19smooth consider smoothing LP, and ferreira2023bayesian combine LPs with additional prior information. Both approaches can be applied to RSLPs instead of standard LPs with a base set of controls, so providing a more appropriate starting point for their procedures. Since RSLP averages across estimates from standard LP regressions, each regression can in principle be corrected for finite-sample bias as proposed by herbst2021bias, or estimated by generalized least squares as discussed by lusompa2023local.

We also note that while we show that our method can outperform factor models in some empirical plausible settings, we do not argue that one has to choose between either factor or random subspace methods. For instance, there is nothing to stop one from using factors in conjunction with subspace methods. One possibility is the application of subspace methods to a set of estimated factors instead of the original set of controls. Therefore, one should not necessarily view our work as a replacement for existing methods. Instead, we are keen to stress the potential for future work to combine our insights with existing developments in the LP literature in attempts to further improve the properties of these LP estimators.