Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
94,233 characters · 17 sections · 90 citation commands
Local Projections Inference with High-dimensional Covariates without Sparsity
Impulse responses analyze how a shock, such as a change in interest rates, affects various aspects of the economy over time. Knowing how different parts of the economy respond to shocks, such as policy changes or natural shocks, helps policymakers make informed decisions and businesses prepare for potential impacts. In recent years, local projections (LPs), introduced by jorda2005estimation, have gained considerable attention as a flexible alternative to traditional Vector Autoregression (VAR) models. LPs directly estimate the relationship between shocks and future outcomes, offering simplicity and intuitive interpretability.
As the use of LPs has grown, so has the recognition of the need for high-dimensional controls in these models. While methods like LASSO have been proposed to handle high-dimensional covariates, they heavily rely on sparsity assumptions—--the idea that most parameters are exactly zero. Recent work by adamek2022lasso,adamek2023_hdlp proposed using a debiased or desparsified LASSO in time series and local projections with weaker form of sparsity, but their approach still faces challenges in scenarios with dense data generating processes (DGPs). Recent studies by giannone2021illusion and kolesar2023fragility have raised significant concerns about the validity and reliability of these assumptions in economic models, challenging their applicability in many empirical settings.
This paper presents a novel approach to local projections that incorporates high-dimensional covariates without relying on sparsity constraints. The proposed method adapts the Orthogonal Greedy Algorithm with a High-Dimensional Akaike Information Criterion (OGA+HDAIC), originally developed by ing2020, to the context of local projections. In contrast to LASSO-based methods, the proposed approach allows parameters to be nonzero but assumes they decay toward zero as dimensionality grows. This assumption is more flexible and is automatically satisfied by autoregressive processes, making it especially suitable for time series settings. This approach offers a more versatile framework that can handle both sparse and dense scenarios, which is ideal for economic time series analysis.
The Orthogonal Greedy Algorithm proposed in this paper offers several advantages over existing methods, particularly in terms of interpretability. Unlike Principal Component methods (\citealp*{stock2002forecasting}), the proposed method prioritizes variables based on their cross-sectional explanatory power. This feature not only enhances the model's interpretability but also provides valuable insights into the most influential covariates in impulse response dynamics. Such information is crucial for economists and policymakers seeking to understand shock propagation mechanisms in the economy. Recent work by dinh2024random has applied random subspace methods in this domain, sharing the motivation to address high-dimensional controls in local projections. While their approach offers empirical insights, the proposed method provides a statistical inference theory in addition to practical applicability.
From a theoretical perspective, this paper contributes to the literature by extending the OGA+HDAIC method of ing2020 to the context of local projections. My coauthors and I previously adapted this method to a cross-sectional double machine learning framework (cha2023inference). In the current paper, I further develop this approach for the time series context by incorporating explicit assumptions on the dependence structure, using the near epoch dependence (NED) assumption as employed by adamek2022lasso, adamek2023_hdlp. A key advantage of using the NED assumption is that it inherits the mixingale properties that I employ in deriving the error bounds.\footnote{I have benefited enormously from the comprehensive foundation of dependence concepts and properties in the exquisite textbook davidson1994stochastic.} I leverage the triplex inequality from jiang2009uniform in conjunction with the results in ing2020 to derive the inference of the proposed method. In deriving the inference, I incorporate a double selection method (bch2013res), which effectively addresses potential overfitting issues in high-dimensional dynamics involving lagged covariates and outcomes. This framework enables the derivation of error bounds and asymptotic Gaussianity, providing both theoretical guarantees and practical implementation strategies.
Building on the solid statistical background, the proposed method's ability to handle high-dimensional controls has important implications for identification, addressing several key challenges. From a causal identification perspective, angrist2018_lpstringtheory proposed interpreting local projections as causal parameters, with the key identifying assumption being unconfoundedness. The inclusion of large sets of covariates, which the proposed method efficiently manages, enhances the plausibility of this assumption (damour2021overlap; rosenbaum2002overt). By incorporating a comprehensive set of potential confounders, we can more confidently assume unconfoundedness, strengthening the causal interpretation of the estimates.
In terms of identifying structural impulse responses, the proposed method addresses the invertibility condition by essentially solving an omitted variable bias problem (STOCKwatsonchap8, p. 450). When relying on the LP-VAR equivalence results from plagborg2021local, we must consider the bias term arising from finite lag approximations. This bias decreases as the number of included lags increases, further justifying the use of high-dimensional controls.
In a Difference-in-Differences (DiD) framework, dube2023local extend local projections to estimate DiD parameters, with the parallel trends assumption the key assumption for identification. Building on this identification scheme, my framework introduces high-dimensional controls, which greatly enhance the credibility of the conditional parallel trends assumption. As highlighted by heckman1997matching,heckman1998characterizing, factors like demographic differences or regional economic conditions can lead to non-parallel trends, introducing bias into traditional DiD estimates. By incorporating a rich set of covariates, high-dimensional methods address these potential biases, making the conditional parallel trends assumption more plausible.
By addressing these identification challenges through the inclusion of high-dimensional controls, the proposed method enhances their interpretability and potential for causal inference. In the following Section (ref), I will demonstrate how these high-dimensional methods are applied to various identification schemes, including structural impulse response estimation and DiD approaches, providing a comprehensive framework for modern empirical analysis.
To demonstrate the effectiveness of this approach, I conduct illustrative simulation studies comparing the proposed method with conventional LP and LASSO-based approaches. For the simulation design, I adapt from MontielMikkel2021_ecta_lpinference, where they proposed lag augmentation---adding one more lag as controls than the true autoregressive model suggests, to achieve robust inference with persistent data a longer horizons. I adopt lag augmentation for all the methods to see if such robust inference holds in finite samples. Simulation results suggest that the proposed method performs consistently well, especially in dense and more persistent scenarios. This is particularly relevant for macroeconomic applications where data often exhibit high persistence and complex interdependencies.
In empirical applications, I apply the proposed method to the followning two studies. First, I revisit the analysis of acemoglu2019democracy on the causal effects of democracy on economic growth. The proposed method demonstrates significant efficiency gains over the conventional local projections approach. It also shows robustness to high-dimensional controls. This application highlights the method's practical use in empirical analysis, where high-dimensional confounders play a critical role.
Next, I reexamine the study by bhandari2024survey, which explores how subjective beliefs, particularly pessimism, influence macroeconomic aggregates like inflation and unemployment. The proposed method shows robustness even as model complexity increases, particularly when incorporating a moderate number of variables with extended lags. This robustness across varying model specifications is crucial in empirical analysis, where the true underlying model structure is often unknown. Furthermore, I introduce a novel state-dependent analysis that distinguishes between good and bad economic conditions. The results reveal that in good states, inflation follows a pattern consistent with the rational behavior of firms. In contrast, in bad states, inflation rises due to firms acting on subjective, pessimistic expectations of future costs. These state-dependent findings provide new insights into the role of belief shocks in shaping macroeconomic outcomes across different economic environments.
The remainder of this paper is organized as follows. Section (ref) presents the notations and implementation details of the proposed method with two main applications for the identification schemes. Section (ref) illustrates its finite sample performance through simulation studies. Section (ref) outlines the theoretical framework, including definitions, assumptions, and main results. Section (ref) presents two empirical applications: one examining the effect of democratization on GDP growth, and the other investigating the impact of pessimism on macroeconomic aggregates. Finally, Section (ref) includes concluding remarks.
I will first introduce a general estimation framework and then provide further details on the estimation procedures in the latter part of this section. Consider the following $h-$step ahead local projection regression model
where $y_{t+h}$ is the $h$-step ahead response variable, $x_t$ the innovation, $\boldsymbol{r}_{t}$ are contemporaneous controls, and $\boldsymbol{z}_{t}$ are lagged controls. The index $t=1,\dots,\bar{T}$ is an index for the observations, $h=1,\dots,H_{\max}$ is an index for horizons, and $\ell=1,\dots,L$ is the number of lags included in the model. This is a general representation of local projections in the literature, as introduced in plagborg2021local. Following the spirit of jorda2005estimation, I fix $h$ and focus on each horizon of interest.
We are interested in the response of $y_{t+h}$ with respect to a shock in $x_t$, and our parameter of interest $\beta_h$, is defined as
The parameter can be interpreted as impulse responses or treatment effects according to different identification assumptions. Below I introduce two applications which include the identification schemes.
As was originally proposed by jorda2005estimation, local projections are tools to estimate impulse responses with correctly specified controls. plagborg2021local present a detailed review on how the LPs can be used to estimate the structural impulse responses, using their equaivalence results between the VAR and LPs. I will briefly introduce this approach following Section 3.1. of plagborg2021local.
Denote the data as $w_t = (\textbf{r}_t', x_t, y_t, \textbf{z}_{t}')'$ with the dimension of $\textbf{r}_t$ as $n_r$. Consider the following Structural Vector Moving Average (SVMA) as the DGP
and assume normality for the structural shocks: $\epsilon_t \overset{i.i.d.}{\sim} N(0,I_{n_\epsilon})$. Under some regularity conditions, the $(i,j)$th element of $\Theta_{i,j,h}$ is the impulse response of variable $i$-th element of $w_t$ to the shock of $j$-th element of $\epsilon_t$ at horizon $h$. Consider the case where we are interested in the response of $y_t$ with respect to a shock in $\epsilon_{1,t}$. The parameter of interest is then $\theta_h := \Theta_{n_r+2,1,h}$ for $h=0,1,\dots$.
Assuming that the structural shock can be recovered as a function of both current and past data, denoted as $\epsilon_{1,t} \in \text{span}(\{w_\tau\}_{\tau\le t})$, the structural shock can be identified as a linear combination of the Wold forecast errors using an SVAR identification scheme. This can be expressed as $\tilde{\epsilon}_{1,t} = b'e_t$, where $b$ is obtained as a function of reduced-form VAR parameters depending on identification schemes. The LP approach involves using reduced-form LP parameters instead of VAR counterparts to generate structural impulse responses. Using this method, the population estimand would be equivalent. For a detailed explanation with different identification schemes, please refer to plagborg2021local.
In cases where the model is non-invertible, the straightforward approach would be to use instruments. I cannot efficiently deal with such cases since this paper does not provide a theory for the instrumental variable LP. However, as mentioned in kanzig2021macroeconomic, non-invertibility is essentially the omitted variable bias problem. In that sense, I can argue that high-dimensional controls still help because the model spans more information.
The above results leverage the VAR-LP equivalence with infinite lags. In practical applications, we rely on a finite lag approximation. The use of finite lags introduces approximation errors, which must be carefully considered in empirical applications. A detailed discussion on the implications of using finite lags is provided in Appendix (ref).
In the empirical analysis in Section (ref), I apply the impulse response analysis scheme outlined above to study the effects of subjective beliefs on some key economic aggregates. I consider several model specifications with an increasing number of lags, revisiting the issues of invertibility and finite lag approximations. This allows me to assess the robustness of the results to potential non-invertibility and to explore the trade-offs associated with different lag lengths in capturing the dynamic effects of belief shocks.
A recent paper by dube2023local proposes a local projections approach to DiD event studies. Following their identification schemes, the parameter of interest $\beta_h$ can be interpreted as a DiD estimator. In the following, I will briefly introduce their identification schemes with application to high-dimensional controls.
Consider a potential outcomes framework (rubin1974estimating) in a panel setting with a binary treatment. Denote $y_{it}(0)$ and $y_{it}(p)$ as the potential outcomes with respect to the control and treatment at time $p \neq \infty$ status. Units are divided into groups, \( g \in \{0, 1, \ldots, G \} \), where each group $g$ shares the same treatment timing, $p_g$. Denote the group \( g = 0 \) as the never treated group with $p_0 = \infty$. Treatment is absorbing, meaning that once a unit receives treatment, it remains treated in all subsequent periods.
The group-specific average treatment effect on the treated (ATT) at horizon \( h \) for group \( g \), which starts treatment at time \( p \), is given by:
The DiD approach hinges on two main assumptions: parallel trends and no anticipation. The no anticipation assumption implies that for any time \( t \) before the treatment time \( p \), the expected difference in potential outcomes is (\( E[y_{it}(p) - y_{it}(0)] = 0 \)). The parallel trend assumption states that $E[y_{i,t}(0) - y_{i,1}(0)|p_i = p] = E[y_{i,t}(0) - y_{i,1}(0)],$ for all $t\ge 2$ and for all $p\in\{1,\dots,T,\infty\}$. Also assume a simple structure to the never treated outcome to be
Now consider the following estimating equation.
where $\delta_t^h$ are time specific controls and $e_{it}^h$ the error terms. The LP-DiD parameter then identifies
From this expression, we can clearly see the two bias terms. The contribution of the referenced paper is to restrict the sample such that both bias terms become zero.
The two key terms contributing to the biases $\Delta D_{i,t+j}$ for $j\le h$ and $\Delta D_{i, t-j}$ for $1\ge j$. These terms go to zero if $\Delta D_{i,t-j} = 0$ for $-h<j< \infty$, which simplifies to $D_{t+h} = 0$ under the assumption of absorbing treatment. This is where the sample restriction plays a critical role.
By restricting to the sample that are either
the LP-DiD parameter identifies
where it provides a convex combination of all group-specific effects $\tau_g(h)$ by removing previously treated observations and observations treated between $t + 1$ and $t + h$ from the control group.
Now, consider the scenario where the researcher believes the parallel trend holds only after controling for a vector of covariates, $\bold{w}_{it}$. Including controls becomes necessary for the identification purpose, resulting in the following conditional parallel trend assumption:
The conditional parallel trends assumption becomes crucial when allowing for variations in outcome dynamics across different groups, provided that these differences can be fully accounted for by the covariates. This makes the assumption far more realistic than the conventional parallel trends assumption, especially in settings where untreated outcomes are not expected to evolve similarly across treatment groups. (heckman1997matching; abadie2005semiparametric; callaway2021difference).
Adopting my estimation framework introduces a new layer to this argument: by accounting for high-dimensional controls, the conditional parallel trends assumption becomes even more realistic and robust. By leveraging these controls, we ensure that any remaining variation in trends across groups is largely orthogonal to the treatment, further strengthening the identification strategy. This approach makes the conditional parallel trends assumption not only more plausible but also more applicable to modern datasets, where a wide range of observable factors can be accounted for in high-dimensional frameworks.
Then the estimating equation becomes
where $\bold{w}_{i,t}$ can include lagged terms of the dependent variable and other controls. By restricting the sample to be newly treated group and the clean control group as before, the $\text{LP-DiD}^{\bold{w}}$ parameter identifies
This framework enables us to focus on the portion of outcome evolution that is not explained by the covariates, isolating the treatment effect more effectively. I will revisit this identification framework and its implications with the empirical application in Section (ref).
For notation simplicity, stack the covariates except for $x_t$ into $\boldsymbol{w}_t = (\boldsymbol{r}_{t},\boldsymbol{z}_{t-1},\dots,\boldsymbol{z}_{t-L})$ and write
where $\beta_{-h} = (\delta_{h,0}',\delta_{h,1}',\dots,\delta_{h,L}')'$. To address the bias that arises from excluding higher-order lag controls, the dimension of $\boldsymbol{w}_t$ expands as the lag order, $L$, increases. Turning to the causal identification scheme of angrist2018_lpstringtheory, it is intuitive to add a large set of covariates as controls to account for unobserved counfounding variables, which also contributes to increasing the dimension of $\boldsymbol{w}_t$.
A straightforward approach to (ref) with high dimensionality would be to apply a regularization method such as the LASSO. However, it is now well known that these methods come with an associated cost known as regularization bias. One way to address this challenge is a debiasing strategy using node-wise regression, regressing each covariate on the remaining covariates, as illustrated by van2014asymptotically and zhang2014confidence. Using the estimates from the node-wise regressions, they construct a remedy for the bias term and control for it. An alternative approach is a double selection method that establishes an orthogonality condition in a similar spirit to bch2013res. This method is analogous to the debiasing process of the LASSO, as noted in semenova2023_hdpanel and chernozhukov2021_lassotimespace. In the context of (ref), the main idea of both approaches is to incorporate another set of regressions of the covariates to control for the regularization bias, either by debiasing or by using the covariates selected in both regression models.
In this paper, I consider the double selection method with a model selection method OGA+HDAIC by ing2020, where it consists of two steps. OGA first orders the covariates $\boldsymbol{w}_t$ by their explanatory power, and then we select the number of covariates that minimizes the high-dimensional AIC criteria. The selections come from the following regression models,
Although not necessary in (ref), I used the subscript $h$ for consistency across both equations. The concept involves applying OGA+HDAIC on both equations (ref) and (ref). To remove the regularization bias, I use the union of two selected covariates as $\boldsymbol{w}_t$ to finally run the local projection regression in (ref). For a clear presentation of the OGA+HDAIC procedure, I fix some notations below.
First, consider the OGA part of $y_{t+h}$ on $\boldsymbol{w}_t$. The ordering of the covariates requires the following definition, which indicates the explanatory power of the covariate $\boldsymbol{w}_i$.
As it looks, it works as a scaled version of a single regressor regression of $\boldsymbol{y}_h$ on each regressor $\boldsymbol{w}_i$. Since it is scaled by the square root of the $L_2$ norm of each regressor, the scale of $\boldsymbol{w}_i$ does not affect the magnitude of $\mu_i(J)$. To choose the one with the most explanatory power, we start with the setting $J=\emptyset$ and choose the one with the largest value of $\left|\mu_i(\emptyset)\right|$. Let this covariate index as
and define the first chosen set as $\widehat{J}_1 := \{\widehat{j}_1\}$. To select the second order covariate, we use the residuals from the first step, using $\mu_i(\widehat{J}_1)$ to measure explanatory power. We choose the second order covariate from the remaining covariates, $\widehat{j}_2 := \arg\max_{i\in [p]\setminus \widehat{J}_1} \left|\mu_i(\widehat{J}_1)\right|$, then update the chosen set of covariates, $\widehat{J}_2 = \widehat{J}_1 \cup \{\widehat{j}_2\}$. Generalizing to the $m-$th order covariate, suppose we have the previously chosen set of covariates, $\widehat{J}_{m-1}$. Compute the following coefficient for all $i\in[p]\setminus \widehat{J}_{m-1}$, where
and select the one with the largest absolute value of $\mu_i(\widehat{J}_{m-1})$,
and update the chosen set of covariates, $\widehat{J}_m = \widehat{J}_{m-1} \cup \{\widehat{j}_m\}$. Repeating this procedure orders the covariates in descending order of their explanatory power, conditioning on the previously chosen set of covariates. Now that the covariates are ordered, the remaining task is to select the threshold for the number of covariates. The information criterion we use is HDAIC\footnote{Note that it is called HD because it takes into account the penalization on $\log p /T$. This notion, unlike the traditional AIC vs. BIC framework, takes into account the size of the entire feature space $p$ relative to the available data $T$. This holistic approach to model complexity resonates in the high-dimensional setting, where the effect of the penalty is amplified as $\log p/T$ diverges with increasing dimensionality.}, deinfed as
where $\widehat{\sigma}_J^2 = \boldsymbol{y}_h'(I-P_J)\boldsymbol{y}_h /T$. The number of covariates to be included in the model is the one which minimizes the information criteria,
Denote the chosen set of covariates as $\widehat{J}^{[1]} := \widehat{J}_{\widehat{m}}$. Next, repeat the OGA+HDAIC for (ref) and obtain the set of covariates $\widehat{J}^{[2]}$. Finally, denote the union set as $\widetilde{J} = \widehat{J}^{[1]}\cup \widehat{J}^{[2]}$ and run the local projection regression (ref) with the selected covariates:
The least squares estimator for $\beta_h$ in this final model is our proposed estimator. For the variance estimator, define $\psi_{t} = v_{t,h} e_{t,h}$ and $\tau^2 = E[\sum_{t=1}^T v_{t,h}^2/T]$. The variance estimator is then defined as
where $\widehat{\Omega}$ is the Newey-West estimator with a bandwidth parameter $K$, which is assumed to be increasing with respect to increasing sample size. We borrow arguments from andrews1991heteroskedasticity for the choice of $K$. $\widehat{\psi}_t$ is the sample analogue of $\psi_t$, where $\widehat{v}_{t,h}$ and $\widehat{e}_{t,h}$ are the residuals from (ref) and (ref).
The proposed algorithm is detailed in Algorithm (ref) after a remark on tuning parameters.
With the above algorithm, one can construct an $100(1-\alpha)\%$ confidence interval for $h-$step ahead impulse response estimator of the following form
where $z_{\alpha} = \Phi^{-1}(1-\alpha/2)$ is the $(1-\alpha/2)$ quantile of the standard normal distribution.
This section highlights the advantages of the proposed method over the commonly used LASSO approach, especially when dealing with unknown sparsity assumptions and the degree of persistence. I present a comparative analysis of the proposed estimator, the standard local projection estimator, and the debiased LASSO\footnote{For the debiased lasso estimation, I use the R package desla provided by adamek2023_hdlp.}, as described in adamek2023_hdlp. For the DGP, I adapt from MontielMikkel2021_ecta_lpinference. Consider the following DGP
where $\boldsymbol{y}_t\in \mathbb{R}^{n}$. We are interested in estimating the reduced-form impulse response of $y_{2,t}$ with respect to the innovation $u_{1,t}$. I evaluate the finite sample properties with $95\%$ coverage probabilities and the median widths of the confidence intervals over $1000$ iterations.
For parameter settings, I set $\tau$ to be $0.3$, $n = 10$, and I consider the sample size of $T=300$. For estimation settings, I adopt the lag-augmentation and set the number of lags included in the model as $L=21$, while the lags included in the DGP is $12$. To evaluate performance over longer horizons, I estimate the reduced-form impulse responses for horizons from $1$ to $60$. The simulations explore different levels of persistence and sparsity by varying $\rho$ and $B_\ell$ for $\ell= 1,\dots,12$. Two levels of persistence are considered: $\rho=0.5$ and $\rho=0.95$. For sparsity, I define the values of $B_\ell$ using a vector $a \in \mathbb{R}^{n-1}$ of different magnitudes. I defined the even elements of each row ${B\ell}_{(j,)}$ as polynomials of $a_j$ with alternating signs. Different values of $a$ are considered to generate different levels of sparsity, where I set $a^{s}= (0.4, -0.36, ...,,-0.094, 0.05)'$ for a sparse scenario and $a^{d}= (0.8, -0.73, ...,,-0.28, 0.2)'$ for a dense scenario.
The resulting coefficients of the corresponding local projection equations are depicted in Figure (ref). The coefficients are ordered in descending order of their absolute magnitude. The left panel displays the regression coefficients from the local projection equation using the sparse vector \(a_s\), while the right panel shows the coefficients using the dense vector \(a_d\). The top panels correspond to data with lower persistence (\(\rho=0.5\)), and the bottom panels correspond to data with higher persistence (\(\rho=0.95\)). Each line represents the regression coefficients from the \(h\)-step ahead local projection equations, with horizons spanning from 3 to 59. Note that \(h\) ranges from 1 to 60, and I have selected 9 points within this range for illustration. In the left panel, the coefficients are sparse, with fewer than 10 coefficients being nonzero. In contrast, the right panel shows a dense set of coefficients; although the absolute magnitudes decay, a larger number of coefficients remain nonzero. The more persistent the data, the more amplified the largest ordered coefficients become.
Figures (ref) and (ref) show the simulation results for different levels of persistence. In the left panel of Figure (ref), the model is sparse and the data are less persistent, making it easier for all estimation methods to perform well. Most methods achieve satisfactory coverage probabilities, except for debiased LASSO, which performs slightly worse. Notably, while the conventional LP achieves $95\%$ coverage rates, it does so at the cost of significantly larger confidence interval widths compared to the high-dimensional methods. Additionally, the conventional LP method shows slight underperformance for the last $5$ horizons, despite its large standard errors. This can be attributed to the sample size of $T=300$ and the fact that the number of covariates included in the model is about $70\%$ of the sample size. As a result, the standard erorr gets larger when fitting all covariates, and this issue worsens with increasing horizons due to the reduced effective sample size.
We can observe a similar pattern for the dense scenario in the right panel, except for a sharp failure in the initial horizons using the debiased LASSO method. The proposed method achieves the coverage rates with small standard errors across the different specifications. The huge gain in efficiency comes from the model selection, and it still achieves efficiency in the longer horizons. This is true for both high-dimensional methods.
Figure (ref) presents the results for the higher persistence scenario. There is a huge failure of the debiased LASSO for either sparse or dense scenarios, especially when it gets to longer horizons. The reason for this failure can be inferred from the median width plot, where LASSO rules out too many regressors. The conventional LP performs similarly to the less persistent case, where it achieves the coverage rates but with larger standard errors. Overall, the proposed method maintains the coverage rates at a much lower cost compared to the conventional LP.
The simulation restuls illustrates that the proposed method maintains desirable coverage rates with more stable and narrower confidence intervals across different specifications. While the conventional LP also achieves the coverage probabilities, it suffers from larger standard errors, especially as the horizon increases. As the debiased LASSO consistently gives the smallest standard errors throughout all scenarios, there might be an appropriate DGP where it outperforms the proposed method at least in terms of efficiency, but less likely in dense scenarios, as the performance of debiased LASSO tends to decline with denser DGPs. Given the uncertainties regarding the underlying sparsity and persistence of the data, the proposed method provides a robust and reliable alternative that performs well across various conditions.
In this subsection, I introduce the definitions used throughout the theory. Most of the definitions and explanations are taken from davidson2021stochastic. I will use these definitions in the context of the main text in the following subsection.
Consider a probability space $(\Omega,\mathcal{F},P)$. With time series data, time flows in one direction: The past is known while the future is unknown. It is hence important to condition on previous information set in the context of time series analysis. The accumulation of information is represented by an increasing sequence of sub $\sigma-$fields, $\{\mathcal{F}_t\}_{-\infty}^\infty$, where $\mathcal{F}_s \subseteq \mathcal{F}_t$ for $s\le t$. With the uncertainty of the future, the natural next step is to take expectations. If $X_t$ is $\mathcal{F}_t-$measurable for each $t$, $\{X_t,\mathcal{F}_t\}_{-\infty}^\infty$ is called an adapted sequence, and $E[X_t|\mathcal{F}_{t-1}]$ is defined. Also, if $E[X_{t+s}|\mathcal{F}_t] = X_t$ a.s. for all $s\ge 0$ under adaptation, the sequence is identified by its history alone, and $\{X_t\}_{-\infty}^\infty$ is called a causal stochastic sequence. I start by introducing the most widely used dependence concept, the martingale difference (m.d.) sequences.
As is clear from the definition, m.d. assumes one-step-ahead unpredictability. It follows that it is uncorrelated with any measurable function of its lagged values, and thus behaves like independent sequences in classical limit results. As shown in davidson2021stochastic, many limit theorems that hold under independence also hold under the m.d. assumption with few additional assumptions about the marginal distributions, making it a preferred dependence assumption for econometricians. To extend our understanding beyond one-step-ahead unpredictability, the next definition introduces the concept of asymptotic unpredictability.
A mixingale generalizes a m.d. as a special case where $\rho_m =0$ for all $m>0$. The equations show a diminishing effect of past information on the present, while suggesting eventual complete knowledge of the present in the distant future. Note that if the sequence is adapted, then $E[X_t|F_{-\infty}^{t+m}]=X_t$ for all $m\ge0$ and (ref) is satisfied. Just as m.d.s behave like independent sequences in limit theorems, mixingales behave like mixing sequences, which implies asymptotic independence with the following definitions.
The $\alpha-$mixing coefficient $\alpha_m$ measures the dependence between the sequences separated by a lag of $m$. If a process $X_t$ is $\alpha-$mixing, the dependence diminishes as the lag increases. Mixing conditions often play a crucial role in further establishing asymptotic properties. Although the mixingale condition offers several advantageous properties, an important limitation arises: even if a function of mixingales is independent, it loses its mixing property with an infinite number of lags (davidson2021stochastic, Chapter 18). The following definition introduces a mapping from a mixing process to a random sequence. This transformation enables the sequence to inherit some desired mixingale properties.
The NED condition is the main dependence assumption I use on the data, following adamek2022lasso and adamek2023_hdlp. In the rest of the subsection, I introduce some useful lemmas using the aforementioned definitions.
The following lemma gives a concentration bound without a specific assumption on the dependence structure, which will be used both in Theorems (ref) and (ref).
This inequality is called the triplex inequality because it has three components. The first element is a Bernstein-type bound, the second deals with dependence, and the last component is on the tail. This bound is a central theory used in the proofs, where I use the mixingale property inherited by the NED assumption to simplify the dependence and the tail bounds. The following lemma provides the simplified bounds by imposing the mixingale assumptions on $\{X_t\}_{-\infty}^{\infty}$.
The definitions of the variables and the parameters come from the baseline models (ref) -- (ref). I start by introducing the assumptions on the data generating processes.
Assumption (ref) (ref) assumes that the error terms in (ref) and (ref) are not correlated with the contemporaneous regressors. Note that I impose the adaptation assumption to use asymptotic martingale difference property the Bernstein blocks must have (davidson1994stochastic, page 387). According to davidson2021stochastic, it is a widely employed assumption in time series econometrics, assuming that future shock information cannot help in predicting $X_t$ given its history (davidson2021stochastic, page 383). Assumption (ref) (ref) ensures the processes have bounded $2\bar{q}-$th moments. Assumption (ref) (ref) is on the dependence structure. While $(x_t, \boldsymbol{w}_t, u_{t,h}, e_{t,h})$ themselves are not mixing processes, they depend almost entirely on `near epoch' of $\{\Upsilon_t\}$ (davidson2021stochastic, page 368), which is $\alpha-$mixing. This allows $(x_t, \boldsymbol{w}_t, u_t, e_{t,h})$ to inherit some mixingale properties, as will be stated in the following lemma.
This lemma shows that some transformations of NEDs and their demeaned processes are mixingales. It thus allows us to use Lemma (ref). Next, I impose some assumptions on the mixingale constants to simplify some bounds used in the proof of the main theorems.
This assumption is on the mixingale constants, where I define the constants to align with the constants that inherit the mixingale properties. The sequence $\rho_m\to 0$ is hence the mixingale sequence for the transformed variables in Lemma (ref). Assumption (ref) gives the assumption of how fast $\rho_m$ should decay. The parameter $\kappa$ governs the strength of the dependence, with larger $\kappa$ values indicating a weaker dependence. Note that this is a technical assumption I need to prove Thoerem (ref), where it simplifies the dependence bound in the triplex inequality (ref). The following is the assumptions on the degree of sparseness of the underlying coefficients, $\lambda_h$ and $\gamma_h$.
This assumption is the main difference between this method and LASSO, in that it doesn't limit the number of nonzero coefficients, but rather restricts the magnitudes of the coefficients. While we do not require that the parameters to have a natural order, these conditions can be expressed more simply by rearranging the parameters in descending order of their magnitude $|\xi_j|$. Denote the rearrangement as $|\xi_{(1)}|\ge |\xi_{(2)}|\ge \dots\ge |\xi_{(p)}|$. Then, Assumption (ref) (ref) implies
where from (ref) we call Assumption (ref) (ref) as a polynomial decay case: if (ref) holds for some $\delta > 1$, then (ref) holds for the same $\delta$, as shown in Lemma A.2 in ing2020. Apart from (ref), it can be shown that Assumption (ref) (ref) also implies
which is a frequently adopted assumption in the high-dimensional literature, as in wang2014adaptive. Assume that (ref) holds for some $\delta$. Applying Hölder's inequality,
and we can see that by setting $C_\delta=C^{\delta/(2\delta-1)}$, (ref) holds. The parameter $\delta$ governs the degree of sparseness, with larger values indicating a faster decay of the coefficients. Thus, we can see that Assumption (ref) covers a wider class of sparsity conditions where the condition trivially implies the exact sparsity case.
Assumption (ref) are some additional assumptions on the covariates and growth rates of the dimensionality $dim(\boldsymbol{w}_t)=p$. Assumption (ref) (ref) limits strong correlations between covariates, ensuring that the OLS coefficients of one covariate on any set of other covariates remain finite. Assumption (ref) (ref) is on minimum eigenvalue of the Gram matrix, with a remark that this condition is on the population level, not on the sample matrix as in the restricted eigenvalue conditions in bickel2009simultaneous. The third condition addresses the relationship between the growth of the covariate dimension $p$ and sample size $T$. It specifies that $\log p$ should not grow too quickly compared to $T$. This condition is crucial for bounding the triplex inequality in the proof of Theorem (ref). Note that this assumption is stronger than the usual growth rates in the high-dimensional literature with i.i.d. data, for example $\log p=o(T^{1/3})$ in bch2013res. While bounding either the Bernstein bound or the tail bound would require $\log p = o(T^{(\delta-1)/(2\delta -1)})$, bounding all the terms simultaneously requires a stricter condition $(\log p)^3 = o(T^{(\delta-1)/(2\delta -1)})$. More detailed explanations are given in the proofs of Theorems (ref) and (ref). While this may sound restrictive, note that this assumption is on the growth rate of $\log p$, where $p$ can grow at much faster rates.
In this subsection I establish the inference for the parameter of interest with each horizon $h$. I use a matrix representation as the baseline instead of (ref) -- (ref) for simplicity. Let $\boldsymbol{y}_h$, $\boldsymbol{x}_h$, $\boldsymbol{u}_h$, and $\boldsymbol{v}_h$ be the $T \times 1$ vector and $W_h$ be the $T \times p$ matrix, where $p = dim(\boldsymbol{w}_t)$. Then the models can be written in a compact matrix form:
Note that the convergence rates are not the fastest rates proposed in the main theorem in ing2020: the error bounds are $(\log p/T)^{1-1/2\delta}$ for the polynomial decay case. The main results require assumptions (A1) and (A2) in ing2020, and with Assumption (ref) and Lemma (ref) those assumptions are not satisfied. I hence use the bounds from weakened assumptions in equations (2.33) and (2.34) in ing2020, resulting in the slower converge rates. As noted in equation (2.35) of ing2020, these relaxed assumptions require $\log p = o(T^{1/3}),$ which is implied by Assumption (ref) (ref). While Theorem (ref) itself is useful in deriving error bounds for the high-dimensional nuisance parameters, our parameter of interest is ${\beta}_h$ in equation (ref). The following theorem derives asymptotic distribution of $\widehat{\beta}_h$ using the error bounds derived in Theorem (ref).
By Theorem (ref), the estimator of interest achieves the square root $T$ convergence rate regardless of the slower convergence rate of the nuisance parameters presented in Theorem (ref). The following theorem establishes the validity of the proposed variance estimator.
With Theorems (ref) and (ref), we can now construct $100(1-\alpha)\%$ confidence intervals for each horizon. Given a confidence level $\alpha$, an asymptotic $100(1-\alpha)\%$ confidence interval $\widehat{\mathcal{I}}_{\alpha,h}$ is given by
where $z_{\alpha} = \Phi^{-1}(1-\alpha/2)$ is the $(1-\alpha/2)$ quantile of the standard normal distribution and $\hat{\sigma}_h$ is defined in Theorem (ref).
I illustrate the performance of the proposed method through an empirical investigation of the response of GDP growth to democratization, revisiting the LP framework in acemoglu2019democracy. They develop a binary index of democracy by gathering information from several different sources of data. The data covers $184$ countries from $1960$ to $2010$. In part of their analysis, they adopt the LP specificataion to see how the effect of democratization evolves over time. I adopt the baseline LP specification in acemoglu2019democracy
where $y_{c,t}$ denotes the log GDP per capita in country $c$ at time $t$, $D_{ct}$ is the binary democracy variable, and $\delta_t^h$ the time dummy. The baseline model includes the lagged terms of the GDP to address the selection into democracy, as the descriptive data suggests the dip in GDP preceding democratizations. The identifying assumption is that conditional on the lags of GDP, countries that democratize do not follow a different GDP trend relative to other nondemocracies. This specification can be viewed as a version of the LP-DiD identification scheme by dube2023local, where the sample is restricted to
I ran (ref) with the conventional LP and with my method, with $p=4$ lags of $y_{c,t}$ included as controls. The results are depicted in Figure (ref). While the point estimates are very similar, the proposed method has efficiency gain overall, especially for the logner horizon estimates. As expected, the standard errors increase with the horizon due to the decreasing effective sample size. However, even after accounting for long-run variance, the proposed method results in smaller standard errors, attributable to the model selection process. The empirical findings align with the original results, indicating significant GDP growth of 20 percent higher than the controls over a 20-year period.
However, as noted in their paper, even after controlling for fixed effects and GDP dynamics, there are possible sources of bias coming from unobservables related to future GDP and the change in democracy. To investigate these factors, the authors show the robustness of the results with different sets of controls. The sets of controls include potential trends related to differences in the level of GDP in the initial period, dummies for the democracy of Soviet and Soviet satellite countries, lags of trade and financial flows, lags of demographic structure, and the full set of region$\times$initial regime$\times$year dummies, and more in the appendix. To see the robustness against high-dimensional controls, I have combined these controls into a union set. This brings the number of covariates included in the model to $276$ with a sample size of $774$.
The results are displayed in Figure (ref). As expected for the conventional LP, the standard errors are larger due to the inclusion of all regressors, and the estimates are generally insignificant. Even with the proposed method, most estimates remain insignificant, although it manages to maintain moderate standard errors. However, the proposed method identifies significant GDP growth at longer horizons, starting from year $25$ onward. The percentage increase reaches $17.33$ percent in year $30$, which is consistent with the original findings using baseline controls. This result suggests that the proposed method provides reliable estimates even when considering a large number of potential confounders, which is essential for the causal interpretation of the parameter.
In this subsection, I revisit the empirical study by bhandari2024survey, which explores the impact of subjective beliefs on macroeconomic aggregates. Their research develops a theory on how the bias of subjective beliefs from rational beliefs affects key economic indicators, formalizing these departures using model-consistent notions of pessimism and optimism.
The theoretical framework posits that fluctuations in beliefs, particularly increases in pessimism, have significant effects on the macroeconomy. This pessimism is hypothesized to be contractionary and to increase belief biases in both inflation and unemployment forecasts. The mechanism operates through various channels: pessimistic households lower current demand due to consumption smoothing, firms adjust their pricing and hiring strategies based on expectations of future economic conditions, and labor market frictions amplify these effects. This shared pessimistic outlook creates a positive correlation between the biases in inflation and unemployment forecasts.
Figure (ref) replicates the dynamic responses originally presented in Figure 10 of bhandari2024survey. It compares the dynamic responses under two scenarios: one where all agents, including firms, hold subjective beliefs, and another where firms adopt rational expectations while households retain subjective beliefs. The solid line represents the case where all agents follow subjective beliefs, while the dashed line reflects the responses when firms adopt rational expectations. This comparison highlights important differences, where rational firms exhibit more muted fluctuations compared to the scenario where all agents are subjective. Specifically, rational firms keep inflation lower on impact, as they perceive an increase in pessimism as contractionary but do not anticipate higher future marginal costs, leading to smaller price adjustments. This figure provides a baseline understanding of the impulse responses under different belief structures, which will serve as the foundation for further analysis.
To test this theory empirically, the authors utilize data from multiple sources. The key variable is the belief bias, which they term the "belief wedge"---defined as the difference between subjective beliefs and rational forecasts. Subjective beliefs are measured using the University of Michigan Surveys of Consumers, while rational predictions are constructed using both VAR predictions and forecasts from the Survey of Professional Forecasters (SPF). The study focuses on quarterly data from 1982Q1 to 2019Q4.
Since their model predicts a one-factor structure of the belief wedges, they define the belief shock as the first principal component derived from the standardized inflation belief wedge and unemployment belief wedge. To address a possible misspecification of VAR forecasts and hence the belief wedges, the authors construct a belief shock using the principal component of the belief wedges between the Michigan and SPF forecasts. The question here is how positive deviation to this belief shock (pessimism) affects the economy.
The general estimating equation can be written as:
where $\theta_t$ is the belief shock, the first PC of the inflation and unemployment belief wedges, $\bold{w}_t$ includes additional controls, and $u_{t,h}$ the error term.
Figures (ref)--(ref) present the responses of the inflation, unemployment, and the belief wedges to a positive innovation in the belief shock across different model specifications. Each set of responses is estimated using the conventional LP and the proposed method. To provide context, the model-implied impulse response functions are included, where all agents are assumed to hold subjective beliefs. This provides a theoretical benchmark against which the empirical results can be evaluated.
I begin by presenting the results for the baseline model, where no additional controls are included and the lag order is set to $4$. As shown in Figure (ref), the empirical results generally align with the model's predictions. Both LPs and the model show initial increases in inflation. While the model does not generate the hump-shaped responses by both methods, the authors say the magnitude is comparable to what the model implies. These outcomes align with increased pessimism, resulting in higher unemployment and inflation wedges. This alignment provides initial support for the robustness of the original study's conclusions.
To ensure robustness, alternative model specifications are considered, including adding variables used to create the VAR forecasts as controls and increasing their lags. These controls, $\bold{w}_t$ in (ref), include key economic indicators such as GDP, inflation, unemployment, among others.
Invoking the recursive identification scheme, this approach assumes that these economic factors influence the belief shock variable, while the belief shock does not, in turn, affect these factors. This assumption aligns intuitively, as both subjective beliefs and rational forecasts are derived from the current state of economic indicators. The picture becomes more complex when additional controls are introduced. When I added the 9 variables used for the VAR prediction with 4 lags (Figure (ref)), the standard errors for the conventional LP increased substantially. This increase in uncertainty makes it challenging to interpret some results coming from LP, particularly for inflation, where no significant effects are observed.
Further increasing the lags to 8 and 10 (Figures (ref) and (ref)), the conventional LP produces increasingly erratic results with inconsistent patterns. In contrast, the proposed method maintains consistent results with narrower confidence intervals across different specifications. Notably, the proposed method closely mimics the model-implied impulse responses, with the exception of unemployment. While neither methods do not comply with the model-implied responses for unemployment, the proposed method consistently find an approximately 0.5 percent increase in unemployment in magnitude across all model specifications.
These results reveal important insights about the robustness of different estimation methods across various model specifications. As the number of included variables and lags increases, the conventional LP method shows increasing variability in its results, producing wiggly estimates with wider standard errors. In contrast, the proposed method demonstrates remarkable robustness, maintaining consistent patterns and narrower confidence intervals across different specifications.
This robustness becomes particularly evident when the model includes more variables and higher lag orders. Importantly, the proposed method performs consistently in both simple and more complex settings, suggesting there's no disadvantage to using it when the underlying model structure is uncertain---a common scenario in empirical analysis.
For comparison, I also ran debiased LASSO\footnote{For the debiased lasso estimation, I use the R package desla provided by adamek2023_hdlp.} from adamek2023_hdlp in Figure (ref). The results with other model specifications are in Appendix (ref), where debiased LASSO shows similar patterns across different model specifications. In particular, the method yields highly persistent responses for inflation and inflation wedges. The unemployment response, while significant, tends to be overestimated, and the unemployment wedge exhibits a persistent negative bias over longer horizons.
These results are challenging to interpret both intuitively and within the theoretical perspective proposed by bhandari2024survey. One possible explanation is that these models are relatively low-dimensional, which is not an ideal setting for LASSO. Alternatively, as suggested by the simulations, the high persistence in the data might be driving the estimates away from expected values.
Another findings of their paper is the cyclical patterns in the belief wedges. While their empirical analysis only included the baseline model for impulse responses, I extend the analysis by exploring different empirical specifications to account for the cyclicality. To examine how the impulse resonses differ across differrent economic conditions, I use NBER recession indicators to separate belief shocks into those occurring in bad states ($\theta_t^b$) during recessions and good states ($\theta_t^g$) outside of recession periods. If this specification shows no difference in responses, we can conclude that economic states do not significantly alter the effects of pessimism shocks. However, if we observe differences, we can infer important implications from these results. The model is specified as follows:
where the key interest lies in comparing the estimates of $\beta_h^{g}$ and $\beta_h^{b}$. I begin with a baseline model without controls, then add controls used in VAR predictions, with both 4 and 8 lags in subsequent models.
In terms of results, I focus on inflation as it shows noticeable differences between good and bad states. The results for other dependent variables are presented in Appendix (ref), where they generally behave similarly across states: increase in both wedges, and hump-shaped patterns for the unemployment. However, inflation exhibits distinctive patterns depending on the economic state. Figure (ref) shows the responses of inflation in good (left panel) and bad (right panel) states. For the baseline model, there is no clear pattern in good states, as no significant effects are observed. In bad states, we observe an initial positive response followed by negative responses. The behavior in bad states, especially for early horizons, resembles the subjective agent model predictions from bhandari2024survey, as shown in Figure (ref).
As in the previous discussion, adding controls make intuitive sense in that key economic indicators affect in the formation of both the subjective and rational beliefs. Importantly, focusing on the models with VAR controls, the proposed model yields robust estimates regardless of lag length. And this is where we see more obvious patterns. Unlike in the baseline model were we couldn't find any significant effect for the good state, we now see significant negative responses in the beginning horizons, in both models regardless of the lag length. Also, we observe significant positive responses in bad states, while the magnitude is smaller than what the model predicts. For convenience, I also included the point estimates and the standard errors in Table (ref). For the bad states in subtable (b), we do see positive point estimates compared to negative ones in the good states coutnerparts.
These findings suggest that the economy behaves more rationally in good states, even in the presence of a pessimism shock, while firms act more like subjective agents in bad states. In good states, as in the rational firms model of bhandari2024survey, firms recognize a pessimism shock but do not associate it with higher future inflation, leading to only a slight and temporary dip in inflation before recovery. This behavior reflects a confidence in the economy's strong fundamentals, which prevents firms from overreacting to the shock.
In bad states, however, pessimism raises the subjective probability of negative outcomes like lower productivity and tighter monetary policy, as noted by the subjective agents model of bhandari2024survey. This leads firms to expect higher future marginal costs, thus reducing their incentive to lower prices even during a contraction. The “muted" inflation response in bad states, which actually involves an increase in inflation, is due to firms' defensive pricing behavior as they brace for worsening conditions. This mirrors shiller2003efficient's concept of negative feedback loops where bad economic states reinforce pessimism, leading to irrational decisions and inflationary pressures as firms anticipate continued declines in productivity and demand.
In summary, the empirical analysis largely supports the findings of bhandari2024survey, showing that positive shocks to belief wedges lead to increases in both inflation and unemployment, consistent with the hypothesized effects of heightened pessimism. The proposed method demonstrates robustness across different model specifications, producing stable results with added controls and varying lag lengths. A key contribution of this analysis is the exploration of belief shock responses in different economic states. In good states, inflation shows a temporary dip followed by recovery, reflecting rational firm behavior. However, in bad states, inflation increases, driven by firms’ pessimistic expectations of higher future marginal costs, as predicted by the subjective agent model. This state-dependent inflation response highlights the role of economic conditions in shaping the effects of pessimism shocks. While other variables, such as unemployment, exhibit consistent dynamics across states, the differentiated inflation response underscores the importance of considering state-dependent behaviors in empirical analyses. These results open avenues for future research to refine theoretical models to account for these distinct patterns across economic conditions.
Local projection has emerged as the preferred alternative to VARs in impulse response analysis due to its robustness to model misspecification and ease of implementation (MontielMikkel2021_ecta_lpinference; olea2024double). As the use of local projections has gained traction, recent literature has addressed the challenges of incorporating high-dimensional covariates, with a particular focus on methods like LASSO. However, the reliance on strong sparsity assumptions in these methods limits their applicability in many empirical settings, especially when dense DGPs are present.
This paper aims to address the gap in the literature by introducing a high-dimensional local projection approach that caters to both sparse and dense settings and takes into account the uncertainty of sparseness in the DGP. Building on the OGA with HDAIC method proposed by ing2020, the proposed method relaxes the need for strict sparsity by allowing parameters to decay toward zero as dimensionality increases, making it especially suitable for economic time series with autoregressive properties.
The proposed framework demonstrates significant advantages in both DiD estimation and impulse response analysis. In addition, the proposed approach has the advantage of interpretability through the use of OGA, which orders covariates based on their explanatory power. Through simulations, I show strong performance in handling persistent data and longer horizons estimations.
By utilizing the OGA with HDAIC and the NED assumptions, the local projection estimator achieves $\sqrt{T}$ asymptotic normality with HAC standard errors. The theoretical foundation of this method is based on the error bounds of ing2020, the triplex inequality of jiang2009uniform, and double selection arguments of bch2013res. These advances strengthen the framework's capacity to handle complex time series data while maintaining reliable inference.
In conclusion, this method advances econometric modeling by offering a practical, robust solution for high-dimensional datasets. Its applicability to both sparse and dense scenarios makes it a practical tool for researchers working on a range of empirical applications, from macroeconomic analysis to event studies, providing a framework that ensures consistency and reliability across a broad range of empirical applications.