EconBase
← Back to paper

Design-Based Inference for Time-Series GMM

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

117,365 characters · 15 sections · 42 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Design-Based Inference for Time-Series GMM

abstractThis paper studies inference for time-series GMM when uncertainty comes from shock assignment within a realized historical episode. Rather than treating the data as one random draw from a population of hypothetical economies, the framework conditions on the historical environment and considers alternative realizations of shocks and instruments. For locally correctly specified GMM estimators, the centered moment has design long-run variance $\Omega_R$, which determines the sandwich covariance for the finite-history estimand. Conventional HAC estimators instead converge to $\Omega_R^+=\Omega_R+\Omega_\mu$, where $\Omega_\mu\succeq0$ is the long-run variance of the centered mean-moment path. HAC inference is therefore conservative for scalar functions of the finite-history estimand. Projection adjustment using predetermined covariates can reduce this HAC variance limit in Loewner order and, under an additional long-run orthogonality condition, yields a tighter conservative bound on the corresponding asymptotic covariance. Monte Carlo evidence shows when the distinction is quantitatively important. In a monetary-policy application, standard-error reductions from rich macro covariates provide a diagnostic for economically meaningful predictable variation in the mean-moment path.

\noindentKeywords: Design-based inference; Time-series GMM; HAC standard errors; Local projections; Finite-population asymptotics; Regression adjustment.

\noindentJEL codes: C12, C22, C32, C36.

Introduction

Suppose a researcher uses a local projection or SVAR to study a past historical episode, such as monetary-policy shocks over a particular time period. Often, researchers are not directly interested in how estimates of dynamic treatment effects---such as impulse responses---would vary across hypothetical economies drawn from the same stationary process. Instead, the question of interest may be the following: holding the historical environment fixed, how different would the estimated impulse response have been under alternative realizations of the economic shocks? This question often corresponds more closely to the object of interest in applied work, but it requires a different notion of variance from the one that is most commonly used. Standard HAC standard errors treat the observed moment path as a draw from an unconditional time-series process. By contrast, the design approach below conditions on the realized finite history and treats only the assignment shocks, and possibly instruments, as random. The estimand is therefore the minimizer of the sample-period population GMM criterion based on average design mean moments over the realized dates, with variance given by the design variance under re-randomization of the shocks.

The distinction is the time-series analogue of the finite-population variance calculation in randomized experiments. In cross-sectional settings, design-based inference asks about treatment effects in the sample, with uncertainty arising only from treatment assignment. Analogously, this framework targets dynamic treatment effects for the observed time period only, with uncertainty generated only by alternative realizations of shocks or instruments. In a simple randomized experiment, studied by Neyman1923_thesis, the variance of a difference-in-means estimator contains an unidentified negative term proportional to treatment-effect heterogeneity, so the usual plug-in formula is conservative. Recent finite-population results extend this logic to regression, M-estimation, and overidentified GMM AbadieAtheyImbensWooldridge2020,Xu2021FinitePopulationMEstimators,KakehiMatsushitaOtsu2026FinitePopulationGMM. The time-series question is how the same unidentified heterogeneity term changes when a shock at one date can affect outcomes and moments at other dates.

This paper answers that question for a class of time-series GMM estimators. The conditioning environment $\mathcal E_T$ collects the outcome maps, state-transition rules, deterministic sampling windows, and covariate paths that the design holds fixed. Conditional on $\mathcal E_T$, only the assignment shocks are re-randomized. The estimand is denoted $\theta_T^\star$, which may be, for example, the average $h$-period ahead impulse response over the observed time period. While standard methods such as LPs and SVARs can deliver valid point estimation of $\theta_T^\star$ under the design-based framework, statistical inference requires more care. Specifically, write the moment used in estimation as $g_t(W_t,\theta)=\mu_{T,t}(\theta) + e_{T,t}(W_t,\theta)$. Here $\mu_{T,t}(\theta)$ is the design-conditional mean of the date-$t$ moments, so the sequence $(\mu_{T,t}(\theta_T^\star))_{t\le T}$ is the mean path, recording predictable date-by-date movement in moments over the fixed historical episode. The key decomposition is that the usual asymptotic variance $\Omega_R^+$ satisfies \[ \Omega_R^+ = \Omega_R + \Omega_\mu, \qquad \Omega_\mu\succeq 0. \] Here $\Omega_R$ is the asymptotic design-based variance of the centered moments. Conventional HAC procedures estimate $\Omega_R^+$ because they are applied to the observed moment sequence and therefore include both design innovation variation and fixed movement in the mean path. The component $\Omega_\mu$ is the long-run variance obtained by applying the same HAC long-run-variance calculation to the mean path after subtracting its sample-period average. This is the dependent-data analogue of the finite-population heterogeneity correction and characterizes when $\Omega_R^+$ and $\Omega_R$ are not equal.

The first contribution of the paper is to show that conventional HAC estimators and HAC multiplier bootstraps estimate a conservative variance limit for scalar design-based inference. If the target is instead an unconditional population of hypothetical economies, as in some forecasting exercises, the usual asymptotic variance remains the relevant object. The results below apply to the finite-history question, where the realized historical environment is held fixed and only the assignment shocks are re-randomized.

The second contribution is a projection-adjusted, or regression-adjusted, refinement based on predetermined covariates. Projecting the moments on such covariates and applying the same HAC long-run-variance calculation to the residualized moments gives an adjusted variance limit $\Omega_R^+(r)$ satisfying $\Omega_R^+(r)\preceq\Omega_R^+$. This positive-semidefinite, or Loewner-order, reduction result does not by itself assert that the adjusted matrix is conservative for the design variance. The stronger interpretation $\Omega_R\preceq\Omega_R^+(r)\preceq\Omega_R^+$ requires long-run orthogonality, meaning zero HAC long-run cross-covariance, between the centered innovation and the part of the date-specific mean path left after projecting on the adjustment covariates. Equality with $\Omega_R$ holds when the residualized centered mean path has zero long-run variance; a simple sufficient case is that the centered mean path lies exactly in the linear span of the adjustment covariates. Without those restrictions, the adjusted matrix is a diagnostic and a feasible reduction of the HAC variance limit rather than an estimator of $\Omega_R$.

The paper studies estimators defined by unconditional moment restrictions, including just-identified LPs, stable VAR normal equations, proxy or event-study IV moments, and locally correctly specified overidentified GMM\@. The main text records primitive sufficient routes for fixed-dimensional LPs and stable VAR normal equations, while Appendix (ref) gives the corresponding details. The Monte Carlo section evaluates the variance limits directly under fixed conditioning environments. The simulations show that conventional HAC can be materially conservative when movement in the date-specific mean path is large and that aligned predetermined adjustment can remove a substantial part of the gap. They also show that covariate adjustment meaningfully improves standard errors only when the adjustment span tracks the predictable component of that mean path. The empirical application to monthly U.S. monetary-policy shocks follows the same logic. It asks whether a richer set of predetermined macro covariates explains part of the HAC moment variation in standard LP and state-dependent specifications. The exercise diagnoses the variance-limit distinction in a familiar monetary-shock setting rather than supplying new structural evidence about monetary transmission.

The paper proceeds as follows. Section (ref) introduces the design-based time-series perspective, Section (ref) relates the argument to design-based econometrics and macro time-series inference, and Section (ref) defines the finite-history GMM estimand and gives the first-order theory. Sections (ref)--(ref) then establish the HAC and bootstrap results, give the projection-adjusted refinement, and report the simulations and monetary-shock application. The appendices contain additional supporting theory, proofs, additional simulations and application diagnostics, and examples showing how other macro estimators fit the moment-level setup.

Motivation

This section makes the conditioning convention concrete before the GMM notation is introduced. It begins with the direct potential-outcome framework of BojinovShephard2019_TimeSeriesExperiments and RambachanShephard2019_NonparametricDynamicCausalModel, under which assignment shocks are the source of randomness in the economy. The examples in this section use i.i.d.\ shocks for clarity; the GMM theory below allows weak dependence and states the needed long-run covariance and CLT conditions at the centered-moment level. To keep the observed-period notation one-sided, write $w_{1:t}=(w_1,\ldots,w_t)$ for the realized shock history up to time $t\ge 1$.

Suppose that $(W_t)_{t \in \mathbb{Z}}$ has common law $\mathcal{F}_W$ on $\mathcal{W}\subseteq\mathbb{R}^{d_w}$ and that, for each observed date $t\ge 1$, there is a vector-valued potential-outcome map $Y_t:\mathcal{W}^t\to \mathbb{R}^{d_y}$ with $w_{1:t}\mapsto Y_t(w_{1:t})$. Conditional on these maps, the assignment shocks are the source of design randomness, and the maps are nonanticipating in the sense that the date-$t$ outcome depends only on shocks realized up to date $t$. This construction motivates the moment-level conditioning convention in Assumption (ref). Fix an observed date $s\ge 1$ and a horizon $h\ge 0$. For any $w,w'\in\mathcal{W}$ and any finite off-date shock values $w_{-s}^{(h)}:=(w_1,\ldots,w_{s-1},w_{s+1},\ldots,w_{s+h})$, the $h$-step causal effect of replacing $W_s=w'$ by $w$ (holding those off-date shocks fixed) is \[ \tau_{s,h}(w,w';w_{-s}^{(h)}) := Y_{s+h}\!\big(w_{1:s-1},\,w,\,w_{s+1:s+h}\big) - Y_{s+h}\!\big(w_{1:s-1},\,w',\,w_{s+1:s+h}\big). \] In the notation of RambachanShephard2019_NonparametricDynamicCausalModel, the usual $h$-period impulse response at date $s$ is $\tau_{s,h}(1,0;W_{-s}^{(h)})$. The observed data are $(Y_t(W_{1:t}), W_t)_{t=1}^T$. This convention is the time-series analogue of the Neyman--Rubin finite-population setup, with shocks playing the role of randomized treatments.

The shock-conditioning convention is already implicit in many time-series models. A stationary SVAR written in its MA\((\infty)\) representation, \(Y_t=\sum_{s=0}^\infty A_sW_{t-s}\), treats the structural shocks as the stochastic input and maps each possible shock history into an outcome path. Conditional on the coefficients and any initial conditions, a shock sequence \(w=(w_1,\ldots,w_T)\) therefore determines a potential outcome path \(Y_t(w)\), while the observed path is the one selected by the realized shocks \(W_t\). The design-based convention takes this potential-outcome map as fixed over the historical episode and re-randomizes only the assignment shocks. The difference from conventional stationary asymptotics is not that SVARs make outcomes random while the design approach does not; outcomes are random in both views. The difference is what is averaged over: conventional asymptotics average over the law generating the economy, whereas the design calculation conditions on the realized finite-history outcome functions and averages only over alternative shock draws.

In order for there to be a difference between the usual and the design-based notions of asymptotic variance, there needs to be heterogeneity in dynamic causal effects. In the cross-sectional Neyman formula, the gap between design-based and conventional variance limits is driven by treatment-effect heterogeneity. In time series, the analogous object is heterogeneity in impulse responses across dates. For example, in the constant-coefficient SVAR(1), \(Y_t=AY_{t-1}+BW_t\), the date-\(t\), horizon-\(h\) response is the same at every date, \(A^hB\). If instead the impact matrix depends on a state variable, \(B=B(Z_t)\), with \(Z_t\) predetermined relative to \(W_t\), then the response is \(A^hB(Z_t)\) and varies with $Z_t$. Holding that fixed, a natural estimand of interest would be the sample average \(\theta_T^\star := (T-h)^{-1}\sum_{t=1}^{T-h}A^hB(Z_t)\). This variation need not imply nonstationarity: the observables and effects may remain strictly stationary, for example if \((Z_t)_{t\in\mathbb Z}\) is strictly stationary. Nor must the heterogeneity itself be identified from a single realized history, as when \(B(z)\) is orthogonal for every \(z\). Design-based and conventional standard errors differ only when the realized economy is treated as fixed and dynamic effects vary over that history. Empirical state dependence is documented, for example, by AuerbachGorodnichenko2012, TenreyroThwaites2016, and RameyZubairy2018, but its relevance must be assessed in each application. The main results do not require the researcher to specify the form of the heterogeneity; the possibility of predictable dynamic-effect variation is enough to make the variance distinction relevant.

The following simple LP with i.i.d.\ Bernoulli shocks gives a Neyman-style setting in which, for each date, the two potential outcomes are treated as fixed. Fix $h$ and set $T_h:=T-h$; that is, the pairs $(Y^{(1)}_{t+h},Y^{(0)}_{t+h})_{t\le T_h}$ are treated as fixed constants. All expectations and variances in this paragraph and Lemma (ref) are conditional on that fixed array. The calculation therefore isolates the same-date, or lag-zero, finite-population variance component. If a full dynamic re-randomization lets overlapping histories move together, serial covariance terms can appear; those terms are handled by the later GMM long-run variance.

After this conditioning, write \[ Y^{(1)}_{t+h}:=Y_{t+h}(W_{1:t-1},1,W_{t+1:t+h}), \qquad Y^{(0)}_{t+h}:=Y_{t+h}(W_{1:t-1},0,W_{t+1:t+h}), \] with finite-history averages $\bar Y_h^{(w)}:=T_h^{-1}\sum_{t=1}^{T_h}Y^{(w)}_{t+h}$ for $w\in\{0,1\}$. The observed outcome in the horizon-$h$ LP is $Y_{t+h}^{\mathrm{obs}}=W_tY^{(1)}_{t+h}+(1-W_t)Y^{(0)}_{t+h}$. Let $\tau_{t,h}:=Y^{(1)}_{t+h}-Y^{(0)}_{t+h}$ and $\bar\tau_h:=\bar Y_h^{(1)}-\bar Y_h^{(0)}$. The first-order linear term for the LP slope is \[ \phi_{t,h}:= \frac{W_t}{p}\big(Y^{(1)}_{t+h}-\bar Y_h^{(1)}\big) -\frac{1-W_t}{1-p}\big(Y^{(0)}_{t+h}-\bar Y_h^{(0)}\big). \]

lemmaConditional on the fixed array described above, suppose $W_t \sim \operatorname{Bernoulli}(p)$ i.i.d.\ with $p\in(0,1)$ fixed, and suppose $T_h^{-1}\sum_{t=1}^{T_h}(|Y^{(1)}_{t+h}|^4+|Y^{(0)}_{t+h}|^4)=O(1)$. Then, as $T_h\to\infty$, the LP slope $\widehat\tau_h$ obtained from the intercept regression $Y_{t+h}^{\mathrm{obs}}=\alpha_h+\tau_h W_t+u_{t,h}$ admits the design expansion \[ \sqrt{T_h}(\widehat\tau_h-\bar\tau_h) = T_h^{-1/2}\sum_{t=1}^{T_h}\phi_{t,h}+r_{T,h}, \qquad \mathbb E[r_{T,h}^2]=o(1). \] Consequently, \[ \mathrm{Var}(\widehat\tau_h) =\frac{1}{T_h}\Gamma_h(0)+o(T_h^{-1}), \] where \[ \Gamma_h(0) :=\frac{1}{T_h}\sum_{t=1}^{T_h}\left[ \frac{\big(Y^{(1)}_{t+h}-\bar Y_h^{(1)}\big)^2}{p} + \frac{\big(Y^{(0)}_{t+h}-\bar Y_h^{(0)}\big)^2}{1-p} - \big(\tau_{t,h}-\bar\tau_h\big)^2 \right]. \] Equivalently, $\Gamma_h(0)=S^2_{1,h}/p+S^2_{0,h}/(1-p)-S^2_{\tau,h}$, where $S^2_{1,h}$, $S^2_{0,h}$, and $S^2_{\tau,h}$ are the finite-history variances, computed with denominator $T_h$, of $Y^{(1)}_{t+h}$, $Y^{(0)}_{t+h}$, and $\tau_{t,h}$, respectively.

Lemma (ref) is a simple special case that connects the time-series setup to the finite-population result of Neyman1923_thesis. Once the array of outcomes obtained by setting $W_t=1$ and $W_t=0$ at each date is fixed, the same-date, or lag-zero, term is the Neyman finite-array component. The negative term $-S^2_{\tau,h}$ is the treatment-effect-heterogeneity adjustment: it is not identified from one assignment because $Y^{(1)}_{t+h}$ and $Y^{(0)}_{t+h}$ are not jointly observed at the same date.

The rest of the paper generalizes this kind of result to the GMM setting. The full theory allows centered moment or influence contributions to be serially dependent. For a scalar centered contribution $U_{t,h}$, write $\Gamma_h^U(0):=T_h^{-1}\sum_{t=1}^{T_h}\mathrm{Var}(U_{t,h})$ and $\Gamma_h^U(\ell):=(T_h-\ell)^{-1}\sum_{t=\ell+1}^{T_h}\mathrm{Cov}(U_{t,h},U_{t-\ell,h})$ for $1\le \ell<T_h$. Then \[ \mathrm{Var}\!\left(T_h^{-1/2}\sum_{t=1}^{T_h}U_{t,h}\right) = \Gamma_h^U(0)+2\sum_{\ell=1}^{T_h-1}\left(1-\frac{\ell}{T_h}\right)\Gamma_h^U(\ell). \] For vector-valued contributions, the positive-lag term is $\Gamma_h^U(\ell)+\Gamma_h^U(\ell)^\top$ rather than $2\Gamma_h^U(\ell)$. The nonzero lag terms are the time-series component. They vanish in the fixed-array Bernoulli calculation above because $\phi_{t,h}$ is a function only of $W_t$ and fixed constants, but they need not vanish for the serially dependent centered moment arrays used below. In the GMM notation, these lag terms enter the design long-run variance $\Omega_R$ of the centered moment, while movement in the date-specific mean path creates the conservative-HAC gap $\Omega_\mu$.

Related literature

The first connection is to design-based econometrics. The finite-population analysis of Neyman1923_thesis conditions on the potential outcomes and treats only the assignment as random. In that setting, the usual conservative variance formula reflects the fact that treatment-effect heterogeneity is not identified from one assignment. AbadieAtheyImbensWooldridge2020, Xu2021FinitePopulationMEstimators, and KakehiMatsushitaOtsu2026FinitePopulationGMM extend related finite-population reasoning to regression, M-estimation, and GMM. Relatedly, Sancibrian2025RepeatedFinitePopulations studies estimation uncertainty in repeated finite populations with latent attributes and noisy repeated measurements; that setting is model-based rather than assignment-based, but it shares the distinction between uncertainty about a realized finite population and uncertainty about a superpopulation. The present paper keeps the same conditioning logic but replaces cross-sectional independent assignment with a dependent time-series moment. The resulting correction is not a single finite-population variance of treatment effects; it is the long-run variance $\Omega_\mu$ of the centered mean path, the object estimated by a growing-bandwidth HAC calculation.

The second connection is to regression adjustment in randomized experiments. Freedman2008RegressionAdjustment emphasizes that regression adjustment is not automatically justified by random assignment, while Lin2013AgnosticRegressionAdjustment gives conditions under which adjustment improves precision. The adjustment studied here has the same model-free character. Predetermined covariates are used to remove predictable components of the mean path. The result is a positive-semidefinite reduction in the Loewner order for the HAC variance limit, while the interpretation as a tighter conservative variance bound requires additional orthogonality between the residualized innovation and the span of adjustment covariates.

A third connection is to time-series experiments and dynamic causal effects. BojinovShephard2019_TimeSeriesExperiments and RambachanShephard2019_NonparametricDynamicCausalModel develop potential-outcome frameworks in which treatment or assignment histories generate dynamic effects, and BojinovRambachanShephard2021PanelExperiments extends this logic to panels. For switchback designs, BojinovSimchiLeviZhao2023Switchback study optimal design and randomization-based inference. LiangRecht2025NEqualsOne study randomization inference for a single time-series unit using links to system identification, and LinDing2025TimeSeriesExperiments study regression-based and design-based inference in time-series experiments. The present paper differs in its estimand and estimators: it studies macroeconomic GMM moments, including smooth transformations, overidentified systems, HAC estimators, dependent multiplier bootstraps calibrated to the HAC long-run variance, and LP or VAR-style impulse-response applications.

Finally, this paper reinterprets standard macroeconometric inference. Classical GMM and extremum theory use long-run covariance matrices to approximate sampling uncertainty Hansen1982GMM,NeweyMcFadden1994_LargeSampleEstimation, and HAC estimators estimate those matrices under weak dependence NeweyWest1987,Andrews1991HAC. In conventional macro applications, the long-run variance averages over both innovation variation and predictable variation in the mean moment under the maintained stochastic process. In the finite-history interpretation developed here, the environment and the date-specific mean path are conditioned on, so the same HAC calculation estimates $\Omega_R^+=\Omega_R+\Omega_\mu$. This estimand/variance-limit distinction is complementary to work on robust LP and VAR inference, which studies how to approximate the sampling law of impulse-response estimators under alternative persistence, lag-order, and horizon asymptotics. The appendix connects the finite-history decomposition to the LP and VAR impulse-response literature, including the general LP survey of JordaTaylor2025, robust LP inference in MontielOleaPlagborgMoller2021 and InoueJordaKuersteiner2024, and LP--VAR equivalence and misspecification comparisons in PlagborgMollerWolf2021 and MontielOleaPlagborgMollerQianWolf2024.

Setup and main results

Estimand and interpretation

As mentioned above, the estimand is both finite-history and design-based. For each sample length, the design conditions on a fixed environment $\mathcal{E}_T$, which records the potential-outcome maps, state-transition rules, deterministic windows, and covariate paths held fixed by the design.

assumptionFor each sample length $T$, the design specifies a conditioning environment $\mathcal{E}_T$. This environment contains the potential-outcome maps, state-transition rules, deterministic sampling windows, and any covariate or state paths explicitly held fixed by the design. Conditional on $\mathcal{E}_T$, the moment array used below is generated by the assignment shocks through maps from shock histories to moment contributions. The assignment shocks are the only source of design randomness, and all expectations, probabilities, and variances are taken with respect to this conditional design distribution.

Assumption (ref) is a conditioning convention, not a law of large numbers or a central limit theorem. It fixes the probability model before the stochastic regularity conditions are imposed. The later mean moments, covariance limits, and variance comparisons are all evaluated under this conditional design distribution, and the asymptotic arguments consider sequences of such conditioning environments. Thus the choice of $\mathcal{E}_T$ is part of the estimand: changing what is held fixed changes the counterfactual experiment and can change the variance target.

This convention also separates two ideas that are often conflated. A variable may be predetermined, in the sense that it is dated before the shock being adjusted for, without being fixed under the design. It is fixed only if its realized path is included in $\mathcal{E}_T$; otherwise it remains part of the shock-history map and may change when earlier assignment shocks are re-randomized. Hence the same object should not be treated as fixed when defining the estimand and stochastic when computing the variance, unless an explicit two-layer model is introduced. Two versions of this choice recur below. In a fixed-state design, the relevant pre-shock covariate, state, or regime path is included in $\mathcal{E}_T$, so conditional means given $\mathcal{I}_{T,t}$ are fixed date-specific quantities. In a dynamic-state design, those variables are themselves generated by earlier assignment shocks and therefore move under re-randomization; the date-specific mean moment then averages over the moving pre-shock state under the design distribution. The analysis uses whichever convention is built into $\mathcal{E}_T$ and keeps that convention fixed within a variance calculation.

Generated shocks and first-stage instruments raise a separate issue. If a high-frequency shock, proxy, or instrument is estimated rather than directly assigned, its first-stage error is neither just a fixed state nor just a moving predetermined variable. Treating that error as part of the design randomness requires augmenting the moment vector, as in Appendix (ref). The main decomposition in Proposition (ref) uses the fixed-environment convention. However, the projection-adjustment result can also be read under a dynamic-state convention as a reduction in the conservative HAC variance limit, but its interpretation as a tighter conservative bound then requires the additional orthogonality conditions in Appendix (ref).

To cover the estimators used in practice, the analysis uses a GMM framework that includes just-identified systems and locally correctly specified overidentified systems. For each date $t$, let $g_t(W_t,\theta)\in\mathbb{R}^k$ denote the moment at parameter value $\theta\in\Theta\subset\mathbb{R}^p$, where $\Theta$ is compact. The argument $W_t$ is shorthand for the relevant shock history under the design distribution, not necessarily only the current shock. Thus an LP moment may contain $y_{t+h}$, lagged controls, or generated shocks, and a VAR moment may contain lagged outcomes. After conditioning on $\mathcal{E}_T$, these objects are components of the map from the shock path to the moment. For brevity, write $g_t(\theta):=g_t(W_t,\theta)$. The main text treats the macro sample as fully observed, so $N=T$; the incomplete-observation case is deferred to the appendix.

Throughout Sections (ref)--(ref), probabilities, expectations, and variances are conditional on $\mathcal{E}_T$. Write $\mathbb{E}_T[\cdot]$ for this design expectation. The date-specific mean moment is $\mu_{T,t}(\theta):=\mathbb{E}_T[g_t(W_t,\theta)]$, and the sequence $(\mu_{T,t}(\theta_T^\star))_{t\le T}$ is the mean path at the estimand. The centered innovation is $e_{T,t}(\theta):=g_t(W_t,\theta)-\mu_{T,t}(\theta)$, and the centered mean path at the estimand is $\tilde\mu_{T,t}:=\mu_{T,t}(\theta_T^\star)-T^{-1}\sum_{s=1}^T\mu_{T,s}(\theta_T^\star)$. Conditional-on-$\mathcal{I}_{T,t}$ expressions in the LP and VAR examples below are primitive sufficient conditions for these fixed date-specific mean moments, not an additional source of randomness in the main decomposition. When the $T$-subscript is not essential, write $\mathcal{I}_t$, $\mu_t$, and $\tilde\mu_t$.

With this convention, the sample moment averages a single design draw of the mapping from shock histories to moments, while the population object averages that same fixed mapping under the conditional design distribution. Define the sample moment and GMM estimator by \[ g_N(\theta):=\frac{1}{T}\sum_{t=1}^T g_t(W_t,\theta), \qquad \widehat\theta_N\in\arg\min_{\theta\in\Theta} g_N(\theta)^\top \widehat A_N g_N(\theta), \] where $\widehat A_N$ is a positive definite estimator of a fixed weighting matrix $A\succ0$. Define the sample-period mean moment by \[ \bar m_T(\theta):=\frac{1}{T}\sum_{t=1}^T \mu_{T,t}(\theta) =\frac{1}{T}\sum_{t=1}^T \mathbb{E}_T[g_t(W_t,\theta)], \] and let \[ Q_T(\theta):=\bar m_T(\theta)^\top A\bar m_T(\theta), \qquad \theta_T^\star\in\arg\min_{\theta\in\Theta} Q_T(\theta). \] The objects of interest are $\theta_T^\star$ and smooth functions $h(\theta_T^\star)$, such as impulse responses or forecast error variance decompositions.

The theory does not require that $\theta_T^\star$ converges to some $\theta^\star$ as $T \rightarrow \infty$; however, this feature can be generated by supposing that $Q_T$ converges uniformly to a deterministic limit $Q$ with unique minimizer $\theta^\star$. When the mean path drifts with $T$, however, $\theta_T^\star$ remains the relevant econometric object for inference, and $\theta^\star$ is best viewed as a convenient asymptotic approximation rather than the estimand.

The potential-outcome examples of RambachanShephard2019_NonparametricDynamicCausalModel often use i.i.d.\ assignments. However, the GMM theory below is stated directly for the centered moment innovations and lag products because those are the objects entering the CLT and HAC estimator. A strictly stationary weakly dependent shock process is a common sufficient route (see, e.g., Doukhan1994,Rio2000), but the raw observed moment may still have a nonconstant date-specific mean. Applications often describe exogeneity through an information set. Let $\mathcal F_t:=\sigma(W_s:s\le t)$ be the filtration generated by the shocks, and let $Z_t$ be a vector of predetermined covariates, such as deterministic trends, regime indicators, lags, and pre-announced variables. Throughout, $\mathcal{I}_{T,t}$ denotes the information set with respect to which the assignment mechanism is exogenous. Section (ref) discusses how suitable choices of $Z_t$ can be used to tighten conservative variance bounds. By convention the contemporaneous shock $W_t$ is not included in $Z_t$. A variable is predetermined when it is dated before the shock whose moment is being adjusted; this timing condition does not by itself mean that the variable is fixed under every possible design distribution.

The rest of the paper uses this convention to compare three moment-level covariance objects. The centered innovation component has design long-run variance $\Omega_R$; the conventional HAC limit is $\Omega_R^+$; and the difference is the nonnegative mean-path component $\Omega_\mu$, so that $\Omega_R=\Omega_R^+-\Omega_\mu$. Once these are established, the relevant asymptotic variance of GMM estimators follows by the usual theory.

Examples and mean-path drift

Two canonical examples below can be used to show where a mean path drift can come from: LPs and VARs.

For a given horizon $h$, let the assignment shock be $x_t$ and let $c_t$ denote predetermined controls, such as lags or trends. Define $\psi_t := (1, x_t, c_t^\top)^\top$ and $\theta_h := (\alpha_h,\beta_h,\gamma_h^\top)^\top$. The LP moment condition is $g^{\mathrm{LP}}_{t,h}(\theta_h)=\psi_t(y_{t+h}-\psi_t^\top\theta_h)$, and so stacking over horizons gives $g^{\mathrm{LP}}_t(\theta)=\big(g^{\mathrm{LP}}_{t,h}(\theta_h)\big)_{h\in\mathcal H}$. This one-shock, one-outcome LP is the running example for the design long-run variance in Theorem (ref), the HAC conservative-variance result in Section (ref), and the projection-adjustment construction in Section (ref). If $y_{t+h}$ is vector-valued, let $r_{t,h}(\theta_h):=Y_{t+h}-M_h\psi_t$ with coefficient matrix $M_h$, and write $g^{\mathrm{LP}}_{t,h}(\theta_h)=(I_n\otimes\psi_t)r_{t,h}(\theta_h)$, so that $\operatorname{vec}(M_h)$ is identified by orthogonality between regressors and residuals.

A second example is a stable VAR($p$) with observed shock vector $W_t$, $Y_t=\Phi_0+\Phi_1Y_{t-1}+\cdots+\Phi_pY_{t-p}+BW_t$, where $Y_t\in\mathbb R^n$, $\Phi_i\in\mathbb R^{n\times n}$, and $B\in\mathbb R^{n\times m}$. Let \[ X_t^\top := (\,1,\, Y_{t-1}^\top,\, \ldots,\, Y_{t-p}^\top,\, W_t^\top\,), \qquad u_t(\phi) := Y_t - \Phi_0 - \Phi_1 Y_{t-1} - \cdots - \Phi_p Y_{t-p} - B W_t, \] with parameter vector $\phi:=\operatorname{vec}(\Phi_0,\Phi_1,\ldots,\Phi_p,B)$. The VAR moment function is $g^{\mathrm{VAR}}_t(\phi)=\operatorname{vec}\!\big(X_t u_t(\phi)^\top\big)=(I_n\otimes X_t)u_t(\phi)$.

In both LPs and VARs, heterogeneous dynamic causal effects imply that $\mu_{T,t}(\theta_T^\star):=\mathbb{E}_T[g_t(W_t,\theta_T^\star)]$ may not be 0 for all $t$. Instead, orthogonality holds on average at $T$, but not necessarily at dates $t < T$. The resulting mean-path drift is the time-series analogue of the treatment-effect-heterogeneity term in Neyman's finite-population variance formula. Remark (ref) records the corresponding conditional mean moments. These formulas are not needed for the abstract GMM theorem, but they explain the economic source of $\Omega_\mu$ in the running examples.

remarkConsider the cases below. \begin{enumerate}[label=(\roman*)] • Suppose, in the local-projection setting, that the shock is centered relative to the assignment information set, $\mathbb{E}[x_t\mid\mathcal{I}_t]=0$, that $\tau_{t,h}$ is $\mathcal{I}_t$-measurable, and that the potential outcome satisfies \[ \mathbb E\!\left[x_t\{y_{t+h}(0)-\mathbb E[y_{t+h}(0)\mid\mathcal{I}_t]\}\mid\mathcal{I}_t\right]=0. \] At the population best linear projection coefficient $\theta_h^\star=(\alpha_h^\star,\beta_h^\star,\gamma_h^{\star\top})^\top$, the $x_t$-row of the conditional mean moment satisfies \[ e_x^\top\mathbb{E}\big[g^{\mathrm{LP}}_{t,h}(\theta_h^\star)\,\big|\,\mathcal{I}_t\big] = \mathrm{Var}(x_t\mid\mathcal{I}_t)\,\big(\tau_{t,h}-\beta_h^\star\big), \] where $e_x$ selects the shock row. If, in addition, the intercept/control block is correctly specified date by date in the sense that $\mathbb{E}[y_{t+h}(0)\mid\mathcal{I}_t]=\alpha_h^\star+\gamma_h^{\star\top}c_t$, then \[ \mathbb{E}\big[g^{\mathrm{LP}}_{t,h}(\theta_h^\star)\,\big|\,\mathcal{I}_t\big] =\begin{bmatrix} 0\\[0.2em] \mathrm{Var}(x_t\mid\mathcal{I}_t)\,\big(\tau_{t,h}-\beta_h^\star\big)\\[0.2em] \mathbf 0 \end{bmatrix}. \] In that case, the conditional mean is constant across $t$ (and, because its sample-period average is zero, equal to zero) whenever $\mathrm{Var}(x_t\mid\mathcal{I}_t)\,\big(\tau_{t,h}-\beta_h^\star\big)\equiv 0$. • Suppose, in the VAR setting, that the true law is $Y_t=\Phi_0(Z_t)+\sum_{i=1}^p \Phi_i(Z_t)Y_{t-i}+B(Z_t)W_t+\varepsilon_t$, with $Z_t$ measurable with respect to the pre-shock information set $\mathcal{I}_t$, $\mathbb{E}[W_t\mid\mathcal{I}_t]=0$, and $\mathbb{E}[\varepsilon_t\mid\mathcal{I}_t,W_t]=0$. At the sample-period estimand $\phi^\star$, define $\Delta_t:=(\Phi_0(Z_t)-\Phi_0)+\sum_{i=1}^p(\Phi_i(Z_t)-\Phi_i)Y_{t-i}$, $\Sigma_{W,t}:=\mathbb{E}[W_t W_t^\top\mid\mathcal{I}_t]$, and $Z^{\mathrm{pred}}_t:=\big(1,\,Y_{t-1}^\top,\,\ldots,\,Y_{t-p}^\top\big)^\top$. If the realized path of $Z_t$ is included in $\mathcal E_T$, it is fixed under the design; otherwise it moves with past assignment shocks under re-randomization. Then \[ \mathbb{E}\big[g^{\mathrm{VAR}}_t(\phi^\star)\,\big|\,\mathcal{I}_t\big] =\operatorname{vec}\!\begin{pmatrix} Z^{\mathrm{pred}}_t\,\Delta_t^\top \\[0.3em] \Sigma_{W,t}\,(B(Z_t)-B)^\top \end{pmatrix}. \] Thus the conditional mean moment vanishes whenever $\Delta_t\equiv 0$ and $\Sigma_{W,t}(B(Z_t)-B)^\top\equiv 0$; in particular this holds when $\Phi_0(Z_t)=\Phi_0$, $\Phi_i(Z_t)=\Phi_i$ for all $i$, and $B(Z_t)=B$. \end{enumerate}

Since DSGE models can often be rewritten via their state-space form into a VAR model Giacomini2013_DSGE_VAR_Relationship, Remark (ref)(ii) also characterizes when GMM-estimated DSGE models exhibit mean drift. Many other macro procedures fit into this setting; representative examples include LP Jorda2005, external-instrument (proxy) SVARs and high-frequency/event-study IV StockWatson2012_BPEA,MertensRavn2013,GertlerKaradi2015, heteroskedasticity-based identification Rigobon2003,LanneLutkepohl2008, FAVAR BernankeBoivinEliasz2005, and minimum-distance / indirect inference that matches model-implied autocovariances or IRFs GourierouxMonfortRenault1993,HallInoueNasonRossi2012. All of these admit per-period moment vectors $g_t(W_t,\theta)$ and a sample-period estimand $\theta_T^\star$. Appendix (ref) records how these estimators fit into the framework.

These examples show that the notation is not special to LPs or VARs. The next subsection therefore states the moment-level assumptions behind the asymptotic results.

Assumptions and asymptotic theory

The estimator assumptions are stated at the moment level. This keeps the theorem general enough for LPs, VARs, proxy procedures, and minimum-distance estimators, while the surrounding prose records primitive sufficient routes for the canonical cases.

assumptionLet $e_{T,t}:=e_{T,t}(\theta_T^\star)$. Conditional on $\mathcal{E}_T$, the centered innovation array satisfies: \begin{enumerate}[label=(\roman*)] • For each fixed $\ell\ge0$, the limit \[ \Gamma_e(\ell):=\lim_{T\to\infty}T^{-1}\sum_{t=\ell+1}^T \mathbb{E}_T[e_{T,t}e_{T,t-\ell}^\top] \] exists, with $\Gamma_e(-\ell):=\Gamma_e(\ell)^\top$ for $\ell>0$, and $\sum_{\ell\in\mathbb{Z}}\|\Gamma_e(\ell)\|<\infty$. • The normalized innovation sum satisfies \[ T^{-1/2}\sum_{t=1}^T e_{T,t}\Rightarrow \mathcal{N}(0,\Omega_R), \qquad \Omega_R:=\sum_{\ell\in\mathbb{Z}}\Gamma_e(\ell). \] • For some $\delta>0$, $\sup_{T,t}\mathbb{E}_T\|e_{T,t}\|^{2+\delta}<\infty$. \end{enumerate}

Assumption (ref) is a condition on the innovation component of the moment, not on the raw moment, the observable series, or the date-specific mean path. It allows the raw moment process to have a changing conditional mean through $\mu_{T,t}$. Strict stationarity of $(e_{T,t})$ is therefore not imposed as part of the baseline theory; it is only one sufficient route.

The primitive sufficient conditions are standard: the assumption follows, for example, if the triangular array $(e_{T,t})$ is uniformly strongly mixing, has the displayed uniform $2+\delta$ moment bound, satisfies a Lindeberg condition, and has fixed-lag Ces\`aro covariance limits with an absolutely summable covariance envelope. In finite-lag LPs with i.i.d.\ or finite-memory assignment shocks, the centered moment is $q_h$-dependent. In stable VARs with short-memory innovations and absolutely summable linear filters, the centered moment is a short-memory linear process. In both cases the assumption reduces to the usual moment and weak-dependence requirements for a time-series CLT. Thus, when a stationary or mixing primitive is invoked below, it is a sufficient condition for the assignment-shock process or for the centered innovation array conditional on $\mathcal E_T$. It is not an assumption that the raw moment $g_t(W_t,\theta_T^\star)$, the observed series, or the fixed mean path $\mu_{T,t}$ is stationary.

The main text specializes to the case where the entire time series of interest is observed. By contrast, the appendix records the modifications under random or deterministic sub-sampling; in other words, the cases where a researcher may be interested in some $T$-length time series, but only observe $N<T$ of those time periods. The asymptotic theory also requires smoothness of the moment map and point identification of the sample-period object. In the canonical LP and VAR cases these requirements reduce to familiar moment, nonsingularity, and stability conditions; the appendix gives the corresponding verification routes.

Let $B_{g,T,t}$, $b_{g,T,t}$, $B_{g,1,T,t}$, and $b_{g,1,T,t}$ denote nonnegative envelope and Lipschitz variables for the moment and its Jacobian.

assumptionFor some $\delta>0$, for each $t$, $g_t(W_t,\theta)$ is measurable and continuously differentiable in $\theta$. For all $\theta,\theta'\in\Theta$, $\|g_t(W_t,\theta)\|\le B_{g,T,t}$, $\|g_t(W_t,\theta)-g_t(W_t,\theta')\|\le b_{g,T,t}\|\theta-\theta'\|$, $\|\nabla_\theta g_t(W_t,\theta)\|\le B_{g,1,T,t}$, and $\|\nabla_\theta g_t(W_t,\theta)-\nabla_\theta g_t(W_t,\theta')\|\le b_{g,1,T,t}\|\theta-\theta'\|$. The envelope moments are uniformly bounded: $\sup_{T,t}\mathbb{E}_T[B_{g,T,t}^{2+\delta}+B_{g,1,T,t}^{2}+b_{g,T,t}^{2}+b_{g,1,T,t}]<\infty$. Finally, $G_T:=T^{-1}\sum_{t=1}^T \mathbb{E}_T[\nabla_\theta g_t(W_t,\theta_T^\star)]\to G$, where $G$ has full column rank.

Assumption (ref) is the standard empirical-process/GMM regularity condition. The envelope and Lipschitz bounds allow the proofs to pass from pointwise convergence to uniform convergence over $\Theta$, control plug-in errors when $\widehat\theta_N$ replaces $\theta_T^\star$, and linearize the sample moment map through the Jacobian. In the appendix proofs, the assumption is used mainly for the uniform laws of large numbers for moments and Jacobians and for the plug-in steps in HAC and bootstrap consistency. To keep the appendix notation light, write $B_g(W_t)$, $b_g(W_t)$, $B_{g,1}(W_t)$, and $b_{g,1}(W_t)$ for these triangular-array envelopes when no confusion arises.

The main consistency proof also needs a uniform law of large numbers over $\Theta$ for the moment function and Jacobian. This condition is stated separately because the innovation short-memory condition at $\theta_T^\star$ does not by itself deliver generic-$\theta$ uniform convergence, and the separation avoids making point identification depend on a stronger process-level stationarity condition.

assumptionThe moment function and Jacobian satisfy \[ \sup_{\theta\in\Theta}\left\|\frac{1}{T}\sum_{t=1}^T\Big(g_t(W_t,\theta)-\mathbb{E}_T[g_t(W_t,\theta)]\Big)\right\|\xrightarrow{p}0, \] \[ \sup_{\theta\in\Theta}\left\|\frac{1}{T}\sum_{t=1}^T\Big(\nabla_\theta g_t(W_t,\theta)-\mathbb{E}_T[\nabla_\theta g_t(W_t,\theta)]\Big)\right\|\xrightarrow{p}0. \]

Assumption (ref) makes explicit the generic-$\theta$ regularity needed for consistency and linearization. It is conceptually distinct from the innovation CLT imposed at $\theta_T^\star$. A primitive route is compactness of $\Theta$, the Lipschitz envelopes in Assumption (ref), and a uniform law of large numbers for the underlying finite-dimensional moment function and Jacobian arrays. For the linear LP and VAR moments used below, this reduces to a law of large numbers for finitely many products of shocks, outcomes, instruments, and controls. The main result further requires local correct specification of the moments.

assumptionThe sample-period estimand satisfies $\sqrt T\|\bar m_T(\theta_T^\star)\|=o(1)$.

Assumption (ref) is automatic in just-identified linear LPs and VARs because the sample-period normal equations imply $\bar m_T(\theta_T^\star)=0$. In overidentified GMM it rules out a first-order misspecification component in the main theorem. Appendix (ref) records the corresponding expansion for the fixed-weight criterion with $\widehat A_N\equiv A$ when this condition is not imposed; the first-order moment then includes Jacobian fluctuations multiplied by the nonzero pseudo-true mean. If the weight matrix is estimated in that locally misspecified case, the weight-estimation expansion is an additional first-order ingredient rather than part of the main theorem. Note further that this average-moment restriction does not require $\mu_{T,t}(\theta_T^\star)=0$ date by date and therefore does not rule out mean-path drift. The centered path $\tilde\mu_{T,t}=\mu_{T,t}(\theta_T^\star)-\bar m_T(\theta_T^\star)$ can be nonzero even when $\bar m_T(\theta_T^\star)=o(T^{-1/2})$.

Finally, standard NeweyMcFadden1994_LargeSampleEstimation-style conditions on the GMM objective function and weight matrix are required to justify estimation. These are the same conditions that would be required in the non-design setting.

assumptionFix a positive definite weight matrix $A\succ0$. Define $Q_T(\theta):=\bar m_T(\theta)^\top A\bar m_T(\theta)$, and assume that for each $T$, $Q_T(\theta)$ is minimized at some $\theta_T^\star\in\Theta$. A feasible weight matrix satisfies $\widehat A_N\to_p A$.

Assumption (ref) is the existence-and-weighting part of the GMM setup. It fixes the sample-period population criterion for each sample length and requires the feasible weight matrix to converge to a fixed nonsingular limit. In the proofs, this condition permits comparison of the sample and population criteria and stabilizes the linear GMM expansion. Consistency for a moving estimand also requires $Q_T$ to keep separating $\theta_T^\star$ from the rest of the parameter space as $T$ changes, so the next assumption states that condition explicitly.

assumptionFor every $\varepsilon>0$, $\liminf_{T\to\infty}\inf_{\|\theta-\theta_T^\star\|\ge\varepsilon}[Q_T(\theta)-Q_T(\theta_T^\star)]>0$.

Assumption (ref) is the moving-estimand identification condition. Because the estimand is $\theta_T^\star$ rather than a fixed $\theta^\star$, the paper needs a statement saying that the criterion continues to separate $\theta_T^\star$ from other parameter values as $T$ changes. Its role is narrow but essential: it is the condition that turns uniform convergence of the sample criterion into consistency of $\widehat\theta_N$ for the sample-period estimand.

For the linearization in part (ii), the theorem uses the usual interior local solution. In the linear GMM applications considered below, this is the standard closed-form estimator on the event that the sample Gram matrix is nonsingular, so the condition is substantive mainly when boundary solutions or singular sample criteria are plausible.

assumptionThere exists $\eta>0$ such that $\{\theta\in\mathbb R^p:\|\theta-\theta_T^\star\|\le \eta\}\subseteq\Theta$ for all sufficiently large $T$. With probability approaching one, $\widehat\theta_N$ lies in this neighborhood and is an interior local minimizer of $J_N(\theta):=g_N(\theta)^\top \widehat A_N g_N(\theta)$.

Assumption (ref) is only needed for the local first-order expansion in Theorem (ref)(ii). The deterministic ball condition makes $\theta_T^\star$ an interior point of $\Theta$ for all large $T$, so differentiability of the population criterion yields the population first-order condition at $\theta_T^\star$. The high-probability sample-local-minimizer condition separately ensures that the sample first-order condition is available and that the line segment used in the mean-value expansion remains inside $\Theta$. The assumption is not needed for the basic consistency conclusion in Theorem (ref).

The first main result is that, under the conditions above, the GMM estimator is consistent and asymptotically normal, with asymptotic variance induced by the design variance of the moments.

theoremFor the estimator with a fully observed sample in this section, $N=T$. Under Assumptions (ref), (ref), (ref), (ref), and (ref), any measurable selection $\widehat\theta_N\in\arg\min_{\theta\in\Theta} g_N(\theta)^\top \widehat A_N g_N(\theta)$ satisfies $\widehat\theta_N-\theta_T^\star\to_p 0$. If Assumptions (ref), (ref), and (ref) also hold, then \begin{equation*} \sqrt{N}\,(\widehat\theta_N-\theta_T^\star)\Rightarrow \mathcal{N}(0,\Sigma), \end{equation*} where \[ \Sigma:=(G^\top A G)^{-1}G^\top A\,\Omega_R\,A G\,(G^\top A G)^{-1}. \] For any continuously differentiable $h:\Theta\to\mathbb{R}^q$, let $\mathcal J_T:=\nabla_\theta h(\theta_T^\star)$. If $\mathcal J_T\to\mathcal J$, then $\sqrt{N}\,\big(h(\widehat\theta_N)-h(\theta_T^\star)\big)\Rightarrow \mathcal{N}(0,\mathcal J\Sigma\mathcal J^\top)$.

When $\Omega_R$ is nonsingular, the inverse-design-covariance weighting is $A^\star=\Omega_R^{-1}$ and the corresponding asymptotic covariance is $(G^\top\Omega_R^{-1}G)^{-1}$. This theorem is a finite-history limit theory for the design-based estimand $\theta_T^\star$. It does not reinterpret $\theta_T^\star$ as a population parameter. As with all the theory in this paper, the randomness in the limiting distribution is the assignment-shock randomness conditional on $\mathcal{E}_T$; the mean path itself is held fixed throughout the theorem.

The local correct-specification restriction has a narrow role: it keeps the main influence function in the familiar GMM form based on the centered moment innovation $e_{T,t}$. In an overidentified system with first-order pseudo-true mean $\bar m_T^\star:=\bar m_T(\theta_T^\star)$, the relevant moment is augmented by Jacobian noise. With $J_{T,t}:=\nabla_\theta g_t(W_t,\theta_T^\star)$ and $G_T:=T^{-1}\sum_t\mathbb E_T[J_{T,t}]$, the first-order term is \[ \zeta_{T,t}:=G_T^\top A e_{T,t}+\big(J_{T,t}-G_T\big)^\top A\bar m_T^\star . \] Appendix (ref) records the corresponding expansion for the fixed-weight criterion with $\widehat A_N\equiv A$. Thus the clean HAC decomposition for the raw moment is a locally correctly specified result, which becomes more complicated under local misspecification.

For the running LP example, the theorem is obtained under the familiar requirements that the assignment shock and predetermined controls generate a finite-memory or short-memory moment process, that the relevant products have finite moments, and that the LP Gram matrix remains nonsingular. For the stable VAR example, it is obtained when the innovation process is short-memory, the autoregressive filters are absolutely summable, and $\theta_T^\star$ is locally identified. For HAC, these first-order conditions are not enough: the lag-product arrays, meaning arrays formed from products of moments at different lags, must also obey the covariance-envelope restrictions used in Proposition (ref).

The following corollary records one sufficient route for the running LP example. Fix a finite horizon set $\mathcal H$ and finite-dimensional regressors $\psi_t=(1,x_t,c_t^\top)^\top$. For each $h\in\mathcal H$, write $u_{T,t,h}^\star:=y_{t+h}-\psi_t^\top\theta_{T,h}^\star$ and $g^{\mathrm{LP}}_{T,t,h}:=\psi_tu_{T,t,h}^\star$. Suppose Assumption (ref) holds. Conditional on $\mathcal E_T$, impose three sufficient restrictions. First, for some fixed integer $q<\infty$, the vector collecting $\psi_t$, $(u_{T,t,h}^\star)_{h\in\mathcal H}$, and the products entering the LP moment function and Jacobian is measurable with respect to the design shocks in the finite window $t-q,\ldots,t+\max\mathcal H+q$. Second, after enlarging $q$ if necessary, the underlying design-shock array is $q$-dependent, and the primitive products entering the moment, Jacobian, and moment-product arrays have uniformly bounded $4+\delta$ moments for some $\delta>0$. Third, the fixed-lag Ces\`aro limits of the primitive second moments exist, $T^{-1}\sum_{t=1}^T\mathbb E_T[\psi_t\psi_t^\top]\to Q_\psi\succ0$, and the sample-period LP normal equations are locally correctly specified.

corollaryUnder the LP finite-window, moment, Ces\`aro-limit, nonsingularity, and local correct specification conditions in the preceding paragraph, the LP moment satisfies Assumptions (ref), (ref), (ref), and (ref); the same finite-window product-array restrictions verify the stochastic part of Assumption (ref) for the kernel and bandwidth under consideration. If Assumptions (ref), (ref), and (ref) also hold for the stacked LP criterion, then Theorem (ref) applies to the stacked LP estimator. If the centered LP mean path also satisfies Assumption (ref) and the kernel and bandwidth satisfy Assumption (ref), then Proposition (ref) and Theorem (ref) apply to the same LP moment.

The next corollary gives the analogous sufficient route for the stable VAR normal equations used as a second running example. Fix $p$, $n$, and $m$, and consider the VAR moments $g^{\mathrm{VAR}}_t(\phi)=\operatorname{vec}(X_tu_t(\phi)^\top)$ with $X_t^\top=(1,Y_{t-1}^\top,\ldots,Y_{t-p}^\top,W_t^\top)$. Suppose Assumption (ref) holds and, conditional on $\mathcal E_T$, the VAR moment function, its Jacobian, and the moment-product arrays satisfy the fully observed sample sufficient conditions in Assumption (ref) of Appendix (ref) after replacing the generic moment by $g_t^{\mathrm{VAR}}$. Suppose also that $T^{-1}\sum_{t=1}^T\mathbb E_T[X_tX_t^\top]\to Q_X\succ0$ and that the sample-period VAR normal equations are locally correctly specified.

corollaryUnder the VAR moment-product, nonsingularity, and local correct specification conditions in the preceding paragraph, the VAR moment satisfies Assumptions (ref), (ref), (ref), and (ref); the product-array restrictions verify Assumption (ref) for the kernel and bandwidth covered by Assumption (ref). If Assumptions (ref), (ref), and (ref) also hold for the VAR criterion, then Theorem (ref) applies to the VAR coefficient estimator. If the centered mean path satisfies Assumption (ref) and the kernel and bandwidth satisfy Assumption (ref), then Proposition (ref) and Theorem (ref) apply to the same VAR moment.

The two primitive corollaries are stated for fixed-dimensional LPs and VARs with a fixed horizon or lag order. The sufficient conditions in Appendix (ref) are stronger than necessary, especially the moment-product and covariance-envelope conditions used for HAC, but they make the product-array content of the high-level assumptions explicit. Cases with many horizons, high-dimensional controls, time-varying parameter dimension, data-selected adjustment sets, or bandwidth rules outside the stated lag-window conditions require separate empirical-process and HAC arguments.

Mean-path decomposition

The next assumption is imposed on the centered mean component of the moment to which the HAC calculation is applied. It requires that this fixed path have a stable short-run, or HAC, long-run covariance.

assumptionFor each fixed $\ell\ge0$, the limit \[ \Gamma_{g,\mu}(\ell):=\lim_{T\to\infty}T^{-1}\sum_{t=\ell+1}^T \tilde \mu_{T,t}\tilde \mu_{T,t-\ell}^\top \] exists, with $\Gamma_{g,\mu}(-\ell):=\Gamma_{g,\mu}(\ell)^\top$ for $\ell>0$, and $\sum_{\ell\in\mathbb{Z}}\|\Gamma_{g,\mu}(\ell)\|<\infty$.

Assumption (ref) should not be read as ruling out deterministic changes in the historical environment. Under the design-based interpretation, the realized paths of states, regimes, calendar variables, sampling windows, and potential dynamic effects are fixed once the conditioning environment is fixed. In that sense, the mean path $(\tilde\mu_{T,t})_{t\le T}$ is itself a deterministic object. The assumption answers a narrower question: after the components treated as fixed by the design have been accounted for in the moment, is the remaining sample-centered mean path the kind of object to which a growing-bandwidth HAC long-run-variance calculation can be applied?

The restriction is therefore a restriction on unresolved low-frequency variation in the moment, not a requirement that impulse responses or shock variances be time invariant. It allows predictable movement in dynamic causal effects to contribute to the mean-path term as long as the centered mean path has stable fixed-lag Ces\`aro autocovariances and those limiting autocovariances are summable in lag. This is the time-series analogue of allowing treatment-effect heterogeneity in the finite-population cross-sectional setting: the heterogeneity may be present in the fixed history, but the variance calculation still needs a regular limiting object. For example, variation driven by a realized short-memory state path can satisfy the assumption even though, conditional on the realized history, that path is fixed.

What the assumption does exclude from the raw HAC calculation is a persistent low-frequency component that has not otherwise been modeled or partialled out. A permanent deterministic break, a monotone trend, a fixed seasonal pattern, or a low-frequency drift of fixed magnitude can have fixed-lag Ces\`aro autocovariances that remain nonnegligible over arbitrarily many lags. A growing-bandwidth HAC estimator would then be accumulating a deterministic low-frequency component rather than converging to a finite long-run covariance. That failure does not contradict the motivation of the paper; it says that such components should not be forced into the same short-memory HAC approximation.

The condition is imposed on the moment after any residualization that is part of the design. If the fixed mean path can be decomposed as \[ \tilde\mu_{T,t}=D_{T,t}\lambda_T+\nu_{T,t}, \] where $D_{T,t}$ is a fixed low-dimensional deterministic block, such as trends, calendar terms, sampling-window indicators, or regime indicators, then raw HAC applied to the unreduced moment need not have a finite long-run limit when $D_{T,t}\lambda_T$ is persistent. Including $D_{T,t}$ in the residualization step removes this deterministic component before the lag-window limit is taken. The relevant mean-path condition is then the corresponding condition on the residual path $\nu_{T,t}$, or on the residualized moment more generally, not on the unreduced low-frequency component. This is why the implementation first residualizes deterministic terms associated with the fixed design and only then uses stochastic predetermined covariates for projection adjustment.

In the special case of time-invariant impulse responses and time-invariant shock variances, $\tilde\mu_{T,t}\equiv0$, so Assumption (ref) is automatic. More generally, the assumption permits deterministic time variation that is either explicitly controlled by the fixed design or leaves behind a residual centered mean path with finite HAC long-run covariance.

Write

equation[equation omitted — 90 chars of source]

The sum in (ref) is a HAC long-run covariance: fixed-lag Ces\`aro limits are taken first, and the resulting lag sequence is then summed. It is not the finite-$T$ sum of all sample autocovariances of the centered path, which would vanish because $\sum_t\tilde\mu_{T,t}=0$. Thus $\Omega_\mu$ measures the short-memory long-run covariance of the residual predictable mean path that remains in the moment being fed to HAC.

The following elementary decomposition is the core object used by the HAC results. For each fixed $\ell\ge0$, define

align[align omitted — 372 chars of source]

and let $\Omega_R^+:=\sum_{\ell\in\mathbb{Z}}\Gamma_{g,\mathrm{hac}}(\ell)$ whenever the sum exists.

propositionUnder Assumptions (ref), (ref), and (ref), the fixed-lag limits in (ref) exist, $\sum_{\ell\in\mathbb Z}\|\Gamma_{g,\mathrm{hac}}(\ell)\|<\infty$, and \begin{equation} \qquad \Omega_R=\Omega_R^+-\Omega_\mu, \end{equation} where $\Omega_\mu\succeq0$.

This decomposition is one of the main results and highlights the main differences between the usual and design-based variances in the time series setting. Proposition (ref) compares two covariance calculations applied to the same moment. The design CLT uses the centered innovation $e_{T,t}$, while HAC applied to the centered observed moment also contains the fixed path $\tilde\mu_{T,t}$. This is the formal reason that the HAC limit contains the design variance plus a nonnegative mean-path component rather than an unrestricted collection of covariance terms. If one instead lets the conditioning object that defines the mean path be random under the design distribution, the same clean decomposition requires additional long-run orthogonality between the innovation and the predictable path. The main theory avoids that ambiguity by conditioning on $\mathcal{E}_T$.

In (ref), the key term in the asymptotic variance is the design-based long-run covariance $\Omega_R$. Proposition (ref) shows that $\Omega_R$ equals the conventional HAC limit minus the nonnegative mean-path term $\Omega_\mu$. The formula reduces to the usual HAC long-run variance when $\tilde \mu_t\equiv 0$, and in the i.i.d.\ no-overlap setting of KakehiMatsushitaOtsu2026FinitePopulationGMM it collapses to $\Omega_R= \Gamma_{g,\mathrm{hac}}(0) - \Gamma_{g,\mu}(0)$. In general, $\Omega_R$ equals the HAC-type component minus the mean-path term $\Omega_\mu$, which is nonnegative but unidentified. This is why conventional standard errors are conservative and why projection or regression adjustment can sharpen the bound under the conditions below. Appendix (ref) records the corresponding formulas under incomplete observation. The applied reading is simple: HAC counts predictable changes in the mean moment as time-series uncertainty, while the design variance does not when those changes are part of the fixed historical environment.

For LPs and VARs, Remark (ref) gives a direct interpretation of $\Omega_\mu$. Under the fixed-state convention, the displayed conditional expectations are the fixed date-specific means $\mu_{T,t}(\theta_T^\star)$. Under the dynamic-state convention, the same formulas describe predictable components before averaging over the moving pre-shock state; the fixed mean path in Assumption (ref) is then the corresponding $\mathbb E_T$ expectation. In the LP running example, these predictable components capture changes in the conditional impulse response $\tau_{t,h}$ or in the shock variance $\mathrm{Var}(x_t\mid\mathcal{I}_t)$ over $t$. Thus $\Omega_R$ captures the short-memory part of moment variation, while $\Omega_\mu=\Omega_R^+-\Omega_R$ captures variation in these predictable components.

remarkWrite $\Omega_\mu:=\sum_{\ell\in\mathbb{Z}}\Gamma_{g,\mu}(\ell)=\mathrm{LRV}(\tilde\mu_t)$, where $\tilde\mu_t:=\mu_{T,t}(\theta_T^\star)-\bar m_T(\theta_T^\star)$ and the displayed conditional expectations below give sufficient expressions for $\mu_{T,t}(\theta_T^\star)$. Then, under the additional LP condition stated in Remark (ref)(i), if $d_{t,h}^{LP}:=\mathrm{Var}(x_t\mid\mathcal{I}_t)(\tau_{t,h}-\beta_h^\star)$ and $\bar d_{T,h}^{LP}:=T^{-1}\sum_{t=1}^T d_{t,h}^{LP}$, then $\Omega_{\mu}^{LP}=\mathrm{LRV}(d_{t,h}^{LP}-\bar d_{T,h}^{LP})\,e_xe_x^\top$, where $e_x$ selects the $x$-row of the LP moment. For the VAR case, if \[ d_t^{VAR}:= \begin{bmatrix} \operatorname{vec}\!\big(Z_t^{\mathrm{pred}}\Delta_t^\top\big)\\[0.2em] \operatorname{vec}\!\big(\Sigma_{W,t}(B(Z_t)-B)^\top\big) \end{bmatrix}, \qquad \bar d_T^{VAR}:=\frac1T\sum_{t=1}^T d_t^{VAR}, \] then $\Omega^{VAR}_\mu=\mathrm{LRV}(d_t^{VAR}-\bar d_T^{VAR})$. In both cases $\Omega_\mu\succeq 0$ (see Lemma (ref)).

The appendix proof of Remark (ref) derives the displayed conditional mean paths. Combining those paths with the definition of $\Omega_\mu$ gives the formulas in Remark (ref). For LPs, under the additional mean condition, the extra term is a rank-one block proportional to the long-run covariance of the centered scalar drift in the $x_t$-row. For VARs, the drift has two blocks: time variation in the deterministic and autoregressive coefficients through $Z_t^{\mathrm{pred}}\Delta_t^\top$ and state dependence in the impact matrix through $\Sigma_{W,t}(B(Z_t)-B)^\top$. By Lemma (ref), positive semidefiniteness follows, so $\Omega_\mu$ measures time variation in predictable dynamic effects, the time-series analogue of heterogeneity in the cross-sectional design-based literature.

Assumptions (ref) and (ref) separate the innovation part of the moment from the predictable mean-path drift induced by time-varying dynamic causal effects. For LPs with finite-lag controls and for stable VARs with short-memory innovations, these assumptions reduce to standard finite-moment and weak-dependence conditions, while allowing rich state dependence in impulse responses so long as the centered drift has finite long-run variance. Appendix (ref) records the incomplete-observation extension. Appendix (ref) gives some basic results on the bias-variance tradeoff in the choice of LPs versus VARs in the design-based setting. The results are essentially the same as in the standard case.

Calculating standard errors

The HAC variance

Under the finite-history design interpretation, HAC standard errors play the same role as Eicker--Huber--White standard errors in the cross-sectional finite-population setting: under the conditions below, they estimate a conservative variance limit for scalar design inference. They are computed as

equation[equation omitted — 237 chars of source]

where $K$ is a bounded, symmetric kernel function and $L=L_T$ is a bandwidth parameter. The empirical lag covariances are, for $\ell\ge 0$, \[ \widehat{\Gamma}_{\mathrm{hac}}(\ell) :=\frac{1}{T}\sum_{t=\ell+1}^{T} \big(g_t(W_t,\widehat\theta_N)-g_N(\widehat\theta_N)\big)\big(g_{t-\ell}(W_{t-\ell},\widehat\theta_N)-g_N(\widehat\theta_N)\big)^{\!\top}. \]

Because Assumption (ref) imposes $\sqrt{T}\,\|\bar m_T(\theta_T^\star)\|=o(1)$, the raw and centered fixed-lag limits coincide, and recentering by either $g_N(\widehat\theta_N)$ or the infeasible $\bar m_T(\theta_T^\star)$ changes the weighted HAC sum by at most $O_p(L_T\|\bar m_T(\theta_T^\star)\|^2)+o_p(1)=o_p(1)$ under Assumption (ref). The centered formula is the natural finite-sample analogue of the centered long-run variance limit $\Omega_R^{+}$; the conventional uncentered formula has the same limit under the stated local correct specification condition. Write $\widehat{\Omega}_R^{+}$ as shorthand for $\widehat{\Omega}_R^{+}(L_T,K)$. The paper places mild assumptions on the kernel in order to rely on standard HAC theory.

assumption$K$ is bounded and symmetric, with $K(0)=1$, $K(x)=0$ for $|x|>1$, and continuity at $0$. The bandwidth parameter $L_T$ satisfies $L_T\to\infty$ and $L_T/\sqrt T\to 0$.

Assumption (ref) is the kernel and bandwidth condition for feasible HAC estimation. The rate $L_T/\sqrt T\to0$ is slightly stronger than the minimal infeasible lag-window requirement $L_T/T\to0$ for the same centered path; it is used only to make the feasible plug-in replacement $g_t(W_t,\widehat\theta_N)$ negligible uniformly across the $O(L_T)$ lags in the HAC sum. The assumption is otherwise the usual lag-window condition for consistent long-run variance estimation with a diverging bandwidth. It is only an implementation assumption for HAC and the multiplier bootstrap; it is not part of the core design-based consistency/CLT result for the GMM estimator itself. Thus admissible bandwidths share the same first-order limit, although finite-sample conservativeness can depend materially on the bandwidth. Let $s_{T,t}:=g_t(W_t,\theta_T^\star)-\bar m_T(\theta_T^\star)$. For $\ell\ge0$, write $\widehat\Gamma_s(\ell):=T^{-1}\sum_{t=\ell+1}^{T}s_{T,t}s_{T,t-\ell}^{\top}$, with the transpose convention for negative lags.

assumptionFor the kernel and bandwidth sequence used in the HAC estimator, the fixed-lag limits $\Gamma_{g,\mathrm{hac}}(\ell)$ in (ref) exist and satisfy $\sum_{\ell\in\mathbb{Z}}\|\Gamma_{g,\mathrm{hac}}(\ell)\|<\infty$. For every fixed $\ell$, $\widehat\Gamma_s(\ell)\xrightarrow{p}\Gamma_{g,\mathrm{hac}}(\ell)$. For that same kernel and bandwidth sequence, \[ \left\|\sum_{1\le |\ell|\le L_T}K(|\ell|/L_T) \big(\widehat\Gamma_s(\ell)-\Gamma_{g,\mathrm{hac}}(\ell)\big)\right\|\xrightarrow{p}0. \] If $\widehat\Gamma_{\hat s}(\ell)$ denotes the same lag covariance with $\widehat s_t:=g_t(W_t,\widehat\theta_N)-g_N(\widehat\theta_N)$ in place of $s_{T,t}$, then for each fixed $\ell$, $\|\widehat\Gamma_{\hat s}(\ell)-\widehat\Gamma_s(\ell)\|\to_p0$, and \[ \left\|\sum_{1\le |\ell|\le L_T}K(|\ell|/L_T) \big[ \widehat\Gamma_{\hat s}(\ell)-\widehat\Gamma_s(\ell) \big]\right\|\to_p0 . \]

Assumption (ref) is the lag-window law of large numbers needed for HAC consistency. The fixed-lag clause controls the zero and finitely many nonzero lags, while the growing-window clause controls the additional lags admitted as $L_T$ diverges. The assumption is weaker than stationarity of the raw moment and stronger than the innovation CLT alone because HAC uses products over a growing set of lags.

Appendix (ref) gives a formal sufficient route. It imposes a uniform covariance-sum bound on the centered moment-product arrays over the retained lags; together with $L_T/\sqrt T\to0$, this bound gives the growing-lag law by Chebyshev's inequality. The feasible plug-in step then follows from the $T^{-1/2}$ GMM rate and the same bandwidth restriction. The $4+\delta$ moment-envelope moments used there are sufficient for products and are not part of the abstract moment-level theorem.

The lag-window class covered by Assumption (ref) includes Bartlett, Parzen, and compactly supported flat-top kernels. Quadratic-spectral kernels, prewhitening, and fixed-$b$ asymptotics require separate arguments. If the centered moment has an exact finite dependence window, a fixed flat-top kernel whose $K=1$ region covers that window recovers the same conservative variance limit; this special case is stated as the fixed-bandwidth alternative in the theorem. The theorem therefore gives consistency for $\Omega_R^+$, the conservative HAC variance limit, rather than for $\Omega_R$ itself.

theoremSuppose the sample is fully observed. Under Assumptions (ref), (ref), (ref), (ref), (ref), (ref), (ref), (ref), and (ref), if the kernel and bandwidth satisfy Assumption (ref), then the feasible centered HAC estimator satisfies $\widehat{\Omega}_R^{+}(L_T,K) \xrightarrow{p} \Omega_R^{+}$. The same conclusion holds with fixed $L$ if the fixed-lag and plug-in clauses of Assumption (ref) hold for that fixed lag set, there exists $q<\infty$ such that the finite-$T$ centered lag covariance of the moment is zero for all $|\ell|>q$, eventually in $T$, and the flat-top kernel satisfies $K(\ell/L)=1$ for every integer $|\ell|\le q$.

This theorem identifies the probability limit of the feasible HAC estimator as the conservative moment matrix $\Omega_R^+$, not as the design long-run variance $\Omega_R$ of the centered moment. The distinction matters exactly when $\Omega_\mu\neq 0$: HAC treats the fixed mean path as part of the long-run moment variation, whereas the design moment variance treats only deviations from that path as random.

corollaryLet \[ \Sigma^+:=(G^\top A G)^{-1}G^\top A\,\Omega_R^+\,A G\,(G^\top A G)^{-1}. \] Under the conditions of Theorems (ref) and (ref), $\Sigma^+-\Sigma\succeq0$. Hence, for any fixed vector $a$, $a^\top\Sigma a\le a^\top\Sigma^+a$.

Furthermore, the following result holds.

corollarySuppose the conditions of Theorems (ref) and (ref) hold. Let $\tau_T:=a^\top h(\theta_T^\star)$ and $\widehat\tau:=a^\top h(\widehat\theta_N)$ for a continuously differentiable functional $h:\Theta\to\mathbb R^q$ and a fixed vector $a\in\mathbb R^q$. Suppose $\mathcal J_T:=\nabla_\theta h(\theta_T^\star)\to\mathcal J$, set $d:=\mathcal J^\top a$, define $v:=d^\top\Sigma d$ and $v^+:=d^\top\Sigma^+d$, and let $\widehat v^+\to_p v^+$. If $v^+>0$, then the Wald interval \[ \widehat\tau \pm z_{1-\alpha/2}\sqrt{\widehat v^+/T} \] has asymptotic design coverage at least $1-\alpha$. If $v>0$, the limiting coverage is \[ 2\Phi\!\left(z_{1-\alpha/2}\sqrt{v^+/v}\right)-1\ge 1-\alpha . \]

Theorem (ref) and Corollaries (ref)--(ref) show that HAC standard errors deliver asymptotically conservative scalar design inference for smooth functions of the finite-history estimand under the fully observed, locally correctly specified conditions. In a constant-coefficient VAR with martingale-difference innovations and no mean-path drift, the relevant score autocovariances may vanish beyond lag zero, so the HAC expression collapses to the usual VAR covariance formula. The HAC notation is used here as a common long-run-variance representation that also covers state dependence, predictable mean-path drift, and other macro GMM moments with nontrivial serial dependence.

As mentioned above, $\Omega_\mu$ is not identified from a single realized history without additional structure. Section (ref) shows that projecting the moment on predetermined covariates can remove predictable components and reduce the conservative HAC variance limit; its interpretation as a tighter conservative variance bound requires the stronger conditions discussed there and in the appendix.

The multiplier bootstrap

Macro bootstrap procedures often approximate impulse-response distributions by resampling moment-contribution-like objects GoncalvesKilian2004,KilianLuetkepohl2017. The design-based analogue here is a dependent Gaussian multiplier applied to centered GMM moments, and it estimates the same conservative HAC covariance matrix as in Theorem (ref). Let the per-date GMM moments stack into $g_t(\theta)\in\mathbb{R}^k$, with sample mean $g_N(\theta):=T^{-1}\sum_{t=1}^T g_t(\theta)$ as above and Jacobian $\widehat G_N(\theta):=\nabla_\theta g_N(\theta)$.

The following two examples help to make the Jacobian in the bootstrap update explicit. For LPs at each horizon $h\in\mathcal H$, let $\psi_t:=(1,x_t,c_t^\top)^\top$ and $\theta_h:=(\alpha_h,\beta_h,\gamma_h^\top)^\top$. If $g^{\mathrm{LP}}_{t,h}(\theta_h)=\psi_t(y_{t+h}-\psi_t^\top\theta_h)$ and $g_t(\theta)=\big(g^{\mathrm{LP}}_{t,h}(\theta_h)\big)_{h\in\mathcal H}$, then $\widehat G_N(\theta)$ is block diagonal with $h$-blocks $-T^{-1}\sum_t\psi_t\psi_t^\top$. When the conditional Gram average converges, the corresponding limit is $G_{\mathrm{LP}}=-\lim_T T^{-1}\sum_t\mathbb E_T[\psi_t\psi_t^\top]$. Moreover, for a VAR$(p)$, let $X_t^\top=(1,y_{t-1}^\top,\ldots,y_{t-p}^\top,x_t^\top)$, $u_t(\varphi):=y_t-\Phi_0-\Phi_1y_{t-1}-\cdots-\Phi_py_{t-p}-Bx_t$, and $\varphi:=\operatorname{vec}(\Phi_0,\Phi_1,\ldots,\Phi_p,B)$. The moment is $g^{\mathrm{VAR}}_t(\varphi)=\operatorname{vec}\!\big(X_t u_t(\varphi)^\top\big)=(I_n\!\otimes X_t)u_t(\varphi)$, so $\widehat G_N(\varphi)=-T^{-1}\sum_t(I_n\!\otimes X_t)\mathcal R_t$, where $\mathcal R_t$ is the linear map with $u_t(\varphi)=y_t-\mathcal R_t\varphi$.

For the multiplier bootstrap, choose a kernel/bandwidth pair such that the Toeplitz matrix $\big(K(|t-s|/L)\big)_{1\le t,s\le T}$ is positive semidefinite. Draw a mean-zero Gaussian vector $\xi=(\xi_1,\ldots,\xi_T)$ with covariance $\mathrm{Cov}^*(\xi_t,\xi_s)=K(|t-s|/L)$. Form the bootstrap moment average $B_N^*:=T^{-1/2}\sum_{t=1}^T \xi_t\big(g_t(\widehat\theta_N)-g_N(\widehat\theta_N)\big)$. Write $\widehat G_N:=\widehat G_N(\widehat\theta_N)$. The bootstrap update is \[ \sqrt T\!\left(\widehat\theta_N^*-\widehat\theta_N\right) = -\Big(\widehat G_N^\top \widehat A_N \widehat G_N\Big)^{-1}\widehat G_N^\top \widehat A_N\,B_N^*, \] whose conditional covariance is the plug-in GMM covariance induced by the HAC moment matrix with kernel $K$ and bandwidth $L$. For the Gaussian multiplier bootstrap, the finite-sample covariance matrix of the multipliers must be positive semidefinite. The paper therefore restricts attention to kernels and bandwidths for which the Toeplitz matrix $\big(K(|t-s|/L_T)\big)_{1\le t,s\le T}$ is positive semidefinite.

theoremSuppose the conditions of Theorem (ref) hold, Assumption (ref) holds, and either the kernel and bandwidth satisfy Assumption (ref) or the fixed flat-top alternative in Theorem (ref) applies. Suppose also, where $L_T$ denotes the chosen bandwidth, possibly constant under the fixed flat-top alternative, that the multiplier Toeplitz matrix $\big(K(|t-s|/L_T)\big)_{1\le t,s\le T}$ is positive semidefinite for each $T$. Conditionally on the data, $\sqrt{T}(\widehat\theta_N^\ast-\widehat\theta_N)\Rightarrow^\ast \mathcal{N}(0,\Sigma^+)$ in probability, where \[ \Sigma^+\;:=\;(G^\top A G)^{-1}G^\top A\,\Omega_R^+\,A G\,(G^\top A G)^{-1}. \]

Thus the bootstrap reproduces the conservative Gaussian law with covariance $\Sigma^+$; it reproduces the exact design law only in the no-mean-path case $\Omega_\mu=0$. Appendix (ref) records how the moment covariance changes under incomplete observation.

Projection and regression adjustment

This section studies a feasible way to reduce the conservative HAC variance limit $\Omega_R^+$. The basic result is a positive-semidefinite, or Loewner-order, reduction of the HAC limit. A stronger interpretation as a tighter conservative bound for $\Omega_R$ requires long-run orthogonality, meaning zero HAC long-run cross-covariance, between the centered innovation and the part of the date-specific mean path left after projecting on the adjustment covariates. Equality with the design variance obtains when the residualized centered mean path has zero long-run variance; a simple sufficient case is that the centered mean path lies exactly in the linear span of the adjustment covariates.

If the date-specific design mean $\mu_t=\mathbb{E}_T[g_t(W_t,\theta_T^\star)]$ were known, one could form the infeasible residual moment $r_t^\star:=g_t(W_t,\theta_T^\star)-\mu_t$ and apply HAC to $(r_t^\star)_{t\le T}$. The resulting long-run matrix would estimate $\Omega_R$ because $\mathbb{E}_T[r_t^\star]=0$ date by date. The difficulty is that $\mu_t$ is not identified from one realized history.

The feasible procedure uses predetermined adjustment variables to approximate the date-specific design mean. The econometric idea is to subtract variation that is predictable before the shock before applying HAC. This can reduce the reported HAC variance limit when those variables track the date-specific mean path, but it does not by itself prove that the adjusted matrix estimates the design variance. Here a variable is predetermined when it is dated before the shock whose moment is being adjusted. After centering or residualizing variables treated as fixed by the design, denote the retained adjustment vector by $z_t$. When such variables are themselves generated by past shocks, the basic positive-semidefinite reduction result applies to the stacked process $(z_t,s_t)$, while the stronger conservative variance interpretation requires the long-run orthogonality condition stated in Proposition (ref).

The population long-run projection objects are defined as follows. For vector processes $(a_t)$ and $(b_t)$ with absolutely summable two-sided cross-covariances, write $m_{a,T,t}:=\mathbb{E}_T[a_t]$ and $\tilde m_{a,T,t}:=m_{a,T,t}-T^{-1}\sum_{s=1}^T m_{a,T,s}$, with analogous definitions for $b_t$. For $\ell\ge0$, set $\Gamma_{ab,\mathrm{hac}}(\ell):=\lim_{T\to\infty}T^{-1}\sum_{t=\ell+1}^T \mathbb{E}_T[a_t b_{t-\ell}^\top]$ and $\Gamma_{ab,\mu}(\ell):=\lim_{T\to\infty}T^{-1}\sum_{t=\ell+1}^T \tilde m_{a,T,t}\tilde m_{b,T,t-\ell}^\top$. For negative lags use $\Gamma_{ab,\mathrm{hac}}(-\ell):=\Gamma_{ba,\mathrm{hac}}(\ell)^\top$ and $\Gamma_{ab,\mu}(-\ell):=\Gamma_{ba,\mu}(\ell)^\top$ for $\ell>0$. Define $S_{ab}^{+}:=\sum_{\ell\in\mathbb{Z}}\Gamma_{ab,\mathrm{hac}}(\ell)$ and $S_{ab}^{R}:=S_{ab}^{+}-\sum_{\ell\in\mathbb{Z}}\Gamma_{ab,\mu}(\ell)$. Because Assumption (ref) imposes $\bar m_T(\theta_T^\star)\to0$, the raw and centered fixed-lag limits coincide for the moment, so $S_{gg}^{+}=\Omega_R^{+}$ and $S_{gg}^{R}=\Omega_R$ when $a_t=b_t=g_t(W_t,\theta_T^\star)$.

In implementation, intercepts, deterministic trends, and regime dummies are first treated as variables fixed by the design and residualized out. The stochastic predetermined adjustment variables are centered or residualized with respect to those deterministic terms before forming long-run projection matrices. Start from a vector of raw $\mathcal{I}_t$-measurable adjustment variables, and let $z_t\in\mathbb{R}^{m}$ denote the resulting retained covariate vector after dropping directions that are linearly redundant in long-run second moment. Write $s_t:=g_t(W_t,\theta_T^\star)$ for the moment to be adjusted. Here $s_t$ denotes the moment being residualized, not an assignment shock. If $z_t$ is generated by past shocks, the basic positive-semidefinite reduction result applies to the stacked process $(z_t,s_t)$; the stronger interpretation of the adjusted matrix as a conservative refinement of $\Omega_R$ requires the long-run orthogonality condition stated in Proposition (ref).

For empirical analogues, define $\widehat\Gamma_{ab,\mathrm{hac}}(\ell):=T^{-1}\sum_{t=\ell+1}^{T} a_t b_{t-\ell}^{\!\top}$ for $\ell\ge0$, set $\widehat\Gamma_{ab,\mathrm{hac}}(-\ell):=\widehat\Gamma_{ba,\mathrm{hac}}(\ell)^\top$ for $\ell>0$, and let $\hat S_{ab}^{+}:=\sum_{|\ell|\le L} K(|\ell|/L)\,\widehat\Gamma_{ab,\mathrm{hac}}(\ell)$. Because $z_t$ has been centered or residualized on variables treated as fixed by the design, and because local correct specification makes the sample-period mean of the moment asymptotically negligible, these uncentered cross-HAC formulas have the same first-order limits as their centered analogues.

assumptionThe stacked process $(z_t^\top,s_t^\top)^\top$ satisfies the same fixed-lag and lag-window HAC regularity as in Assumption (ref) componentwise, and $S_{zz}^{+}:=\sum_{\ell}\Gamma_{zz,\mathrm{hac}}(\ell)\succ 0$.

Assumption (ref) bundles the nondegeneracy of the retained covariate span with the stacked HAC regularity needed to estimate $S_{zz}^+$, $S_{zs}^+$, and $S_{sz}^+$ consistently. The positive-definiteness condition is therefore a normalization of the chosen covariate span, not a substantive restriction that every raw candidate covariate be noncollinear. This assumption is used only in Proposition (ref) and the corresponding regression-adjusted covariance formulas.

The feasible projection coefficient is $\hat B:=(\hat S_{zz}^{+})^{-1}\hat S_{zs}^{+}$. The residualized moments are $\hat r_t:=g_t(W_t,\widehat\theta_N)-\hat B^\top z_t$. The reported adjusted matrix is $\widehat\Omega_{R}^{+}(\hat r):=\hat S_{\hat r\hat r}^{+}$. The corresponding GMM covariance matrix is \[ \widehat\Sigma_{\mathrm{RA}}(A) := (\widehat G_N^\top A \widehat G_N)^{-1} \widehat G_N^\top A\,\widehat\Omega_{R}^{+}(\hat r)\,A \widehat G_N (\widehat G_N^\top A \widehat G_N)^{-1}. \]

The following result shows that the projection-adjusted standard errors converge to their limits under the design-based asymptotics. It also states the exact sense in which asymptotic variance is reduced.

propositionSuppose that Assumptions (ref), (ref), (ref), (ref), (ref), (ref), (ref), (ref),(ref), (ref), and (ref) hold. If $L_T\to\infty$ and $L_T/\sqrt T\to0$, then $\widehat\Omega_{R}^{+}(\hat r)\to_p\Omega_R^+(r):=S_{rr}^+$, where $r_t=s_t-B^{\circ\top}z_t$ and $B^\circ:=(S_{zz}^+)^{-1}S_{zs}^+$. Moreover: \begin{enumerate}[label=(\roman*)] • The adjusted HAC variance limit satisfies $\Omega_R^+(r)=S_{ss}^{+}-S_{sz}^{+}(S_{zz}^{+})^{-1}S_{zs}^{+}\preceq S_{ss}^{+}=\Omega_R^+$. Equality holds if and only if $S_{zs}^+=0$. • For any fixed weighting matrix $A\succ0$, the corresponding GMM asymptotic covariance satisfies $\Sigma_{\mathrm{RA}}(A):=(G^\top A G)^{-1}G^\top A\,\Omega_R^+(r)\,A G\,(G^\top A G)^{-1}\preceq\Sigma^+(A)$, where $\Sigma^+(A)$ is the same sandwich with $\Omega_R^+$ in place of $\Omega_R^+(r)$. • Whenever $\Omega_R^+(r)$ and $\Omega_R^+$ are positive definite, the corresponding inverse-variance-weight covariances satisfy $\Sigma_{\mathrm{RA}}^\star:=(G^\top \Omega_R^+(r)^{-1}G)^{-1}\preceq (G^\top (\Omega_R^+)^{-1}G)^{-1}=:\Sigma_{\mathrm{raw}}^\star$. \end{enumerate}

The next proposition states the stronger conservative variance interpretation available when the residualized adjustment component is long-run orthogonal to the centered innovation component of the moment. In the fully observed setup, date-specific centering of $e_{T,t}$ already gives $S_{e\mu}^+=0$. The additional substantive restriction is the long-run orthogonality $S_{ez}^+=0$. This condition holds automatically for deterministic residualization variables fixed under the design, such as intercepts and deterministic trends, and is substantive for stochastic lagged macro variables generated by past shocks.

propositionSuppose the conditions of Proposition (ref) hold with a fully observed sample. Let $B^\circ:=(S_{zz}^{+})^{-1}S_{zs}^{+}$, $r_t:=s_t-B^{\circ\top}z_t$, and $d_t:=\tilde\mu_{T,t}-B^{\circ\top}z_t$. Suppose also that the long-run cross matrix $S_{ez}^{+}$ exists and equals zero, and that $S_{dd}^{+}$ exists. Then \[ \Omega_R^+(r)=\Omega_R+S_{dd}^{+} =\Omega_R+S_{\mu\mu}^{+}-S_{\mu z}^{+}(S_{zz}^{+})^{-1}S_{z\mu}^{+}, \] and hence $\Omega_R\preceq\Omega_R^+(r)\preceq\Omega_R^+$. Equality with $\Omega_R$ holds if and only if $S_{dd}^{+}=0$; in particular, equality holds when there is a fixed matrix $B_0$ with $\tilde\mu_{T,t}=B_0^\top z_t$ along the sequence.

Proposition (ref) concerns positive-semidefinite reduction of the conservative HAC variance limit $\Omega_R^+$. Proposition (ref) gives the additional conditions under which the adjusted object is also conservative for $\Omega_R$. In the fully observed fixed-environment case, the cross term $S_{e\mu}^{+}$ vanishes by date-specific centering of $e_{T,t}$; the substantive long-run orthogonality condition in the proposition is $S_{ez}^{+}=0$.

The distinction matters because “predetermined” is a timing statement, not a guarantee of design validity. A lagged macro variable is observed before the current shock, but if it was generated by earlier shocks it moves under designs that redraw those earlier shocks. In that case Proposition (ref) still gives a reduction of the HAC limit; Proposition (ref) supplies the extra orthogonality condition needed to interpret the adjusted matrix as a tighter conservative bound for $\Omega_R$. The result is fixed-dimensional; high-dimensional or data-selected adjustment sets require incorporating the selection step into the moment or proving that it is asymptotically negligible. In finite samples either standard error can be larger, so the result does not justify adding covariates mechanically.

In applications the paper therefore reports baseline HAC and regression-adjusted HAC side by side, describes the timing and content of the covariate set, and interprets adjustment through whether the covariates are credible proxies for predictable date-specific drift. A convenient moment-level scalar summary is $\mathrm{Gap}:=[\mathrm{tr}(\widehat\Omega_R^{+})-\mathrm{tr}(\widehat\Omega_R^{+}(\hat r))]/\mathrm{tr}(\widehat\Omega_R^{+})$. For a selected parameter block, the same trace ratio can be computed after applying the GMM sandwich and the corresponding selection matrix. Large values of $\mathrm{Gap}$ indicate substantial predictable movement in date-specific mean moments captured by the chosen covariate set; they do not, without the stronger conditions, estimate $\Omega_\mu$ itself.

Monte Carlo evidence

In each Monte Carlo setting below, the realized state and disturbance paths that define the conditioning environment are held fixed, and the assignment shocks are repeatedly re-randomized. This makes the sample-period estimand and its design variance directly observable by simulation. The data-generating process is: \[ y_t = 0.85 y_{t-1} + \tau_t W_t + u_t, \qquad u_t = 0.5 u_{t-1} + \sqrt{1-0.5^2}\,\eta_t, \qquad c_t = 0.9 c_{t-1} + \sqrt{1-0.9^2}\,\varepsilon_t, \] with i.i.d.\ standard-normal $(\varepsilon_t,\eta_t)$, i.i.d.\ Rademacher $W_t$, and a burn-in of $500$ observations. Here, $W_t$ is the re-randomized shock, and $\tau_t$ is the date-specific effect of that shock. Thus time variation in $\tau_t$ is a source of predictable drift in the local-projection moment: even though the shock is randomized, the effect of the shock can be larger or smaller at dates with different realized states.

The main text reports two designs chosen to isolate when regression adjustment should help. Design A sets $\tau_t=1+c_t$. Here the shock effect is larger when the persistent state $c_t$ is high, and the same state $c_t$ is used as the predetermined adjustment variable. This is the aligned case: the adjustment variable is deliberately chosen to track the predictable component of the mean path, so regression-adjusted HAC should move the feasible variance estimate toward the oracle benchmark. Design B sets $\tau_t=1+c_t+0.75(c_t^2-1)$, but still adjusts only with $c_t$. The term $c_t^2-1$ is a population-centered nonlinear state component, so Design B keeps the same basic state-dependence while adding mean-path variation that a linear adjustment in $c_t$ cannot span. This design asks whether regression adjustment can still provide benefits when the adjustment only imperfectly spans the predictable mean-path variation.

At each horizon $h=0,\ldots,12$, the simulation estimates the local projection \[ y_{t+h} = \alpha_h + \beta_h W_t + \gamma_h y_{t-1} + \delta_h c_t + \mathrm{error}_{t,h}, \] where $\beta_h$ is the LP coefficient whose sample-period target is $\beta_{T,h}^\star$, $\alpha_h$ absorbs the average level of $y_{t+h}$, $\gamma_h$ controls for lagged outcome persistence, and $\delta_h$ controls for the observed state $c_t$. The number of usable observations at horizon $h$ is $T_h:=T-h$, and all feasible procedures use Bartlett HAC with bandwidth $L_h=\lfloor T_h^{1/3}\rfloor$. The results below use $T=240$. For each environment and horizon, the reference simulation estimates the sample-period estimand $\beta_{T,h}^\star$ together with the design variance $V_{R,h}$, the raw HAC limit $V_{H,h}^{+}$, the regression-adjusted HAC limit $V_{RA,h}^{+}$, and the infeasible oracle variance $V_{O,h}$ obtained after removing the true mean path. The simulation then evaluates $90\%$ normal intervals for $\beta_{T,h}^\star$ under repeated shock re-randomizations. Reported entries are averages across environments, and parenthesized values are standard errors across environments. The baseline designs use $J=20$ fixed environments, $R=250$ shock re-randomizations per environment, and $M=5000$ reference draws per environment. The misspecified-design appendix keeps $\tau_t=1+c_t$ but replaces $c_t$ in the adjustment by $y_{t-1}$. The appendix also reports a no-drift diagnostic with $\tau_t\equiv 1$, a sample-size comparison, and a comparison of the same LP estimator with several conventional local-projection inference choices. The no-drift, misspecified, and local-projection comparison appendices use the same $J$, $R$, and $M$ counts; the sample-size appendix uses a lighter grid with $J=10$, $R=120$, and $M=2000$.

The variance ratios show that the mean-path term is quantitatively important mainly at short horizons in these designs. Table (ref) reports selected values at horizons $h=0,4,8,12$. In Design A, HAC is notably conservative at short horizons: at $h=0$, $(V_H^{+}/V_R,V_{RA}^{+}/V_R,V_O/V_R)=(6.026,1.633,0.999)$. In Design B the same short-horizon conservativeness is more pronounced, with $(9.939,4.821,0.998)$ at $h=0$. By $h=4$ the corresponding triples are $(1.136,1.028,1.011)$ in Design A and $(1.170,1.082,1.006)$ in Design B, and by $h=8$ the gap is considerably smaller. Figure (ref) shows the full horizon path. In these AR designs the quantitative relevance of the mean-path term is therefore concentrated at short horizons.

table[table omitted — 830 chars of source]

The interval diagnostics show that these variance-ratio differences translate into shorter intervals when adjustment is aligned with the drift. Table (ref) reports the corresponding coverage and length results. In Design A at $h=0$, HAC coverage is $0.999$ with mean length $0.498$, whereas RA delivers coverage $0.938$ and length $0.258$, close to the oracle values $0.893$ and $0.210$. In Design B at $h=0$, HAC gives $(0.998,0.621)$ for (coverage, length), RA gives $(0.994,0.439)$, and the oracle gives $(0.898,0.218)$. By $h=4$ in Design B the remaining gap is considerably smaller: coverage is $(0.870,0.858,0.866)$ and mean length is $(0.770,0.734,0.744)$ for HAC, RA, and the oracle. The residual deviation from nominal coverage at medium horizons is therefore shared by RA and the oracle and is consistent with finite-sample approximation error common to both procedures. Figure (ref) shows that the main differences operate through short-horizon interval length rather than through instability at longer horizons.

table[table omitted — 1,421 chars of source]
figure[figure omitted — 331 chars of source]
figure[figure omitted — 669 chars of source]

The simulations illustrate two conclusions under these designs. First, conventional HAC can be materially conservative for finite-history estimands, especially at short horizons. Second, predetermined projection can remove a substantial part of that conservativeness when the adjustment variable tracks the mean path, but the improvement is limited when it does not. These designs diagnose the variance decomposition and show how the variance limits move when the adjustment span is aligned or misaligned with the mean path. The appendix comparisons with misspecified adjustments and with common local-projection inference choices make the same distinction visible.

The deliberately adverse state-feedback design in Appendix (ref) reinforces the same qualification. There the adjustment variable is affected by the shocks that are re-randomized, and the adjusted variance ratio falls to $0.033$ with coverage of $0.238$, so the failure is not a small finite-sample deterioration. The example emphasizes that timing alone is not enough for the conservative adjusted interpretation: adjustment variables must either be fixed by the design, or satisfy the long-run orthogonality conditions needed for Proposition (ref).

Empirical application

The empirical application asks whether the same distinction that appears in the simulations matters in a familiar macroeconomic setting. It studies monthly U.S. monetary-policy shocks using the updated Jaroci\'nski--Karadi shock series and public monthly macro data, with a simple diagnostic logic: in local projections, the shock is the object whose assignment variation identifies the response, but the moment used for inference can still vary predictably over time because the macroeconomic environment changes. If information available before the shock captures this predictable variation, then residualizing the HAC moment on that information should matter. The application asks how much of the HAC moment variation in standard monetary-shock regressions can be explained by economically interpretable pre-shock macro variables, and it treats the resulting standard-error reductions primarily as a diagnostic for predictable mean-moment variation unless the stronger conservative-adjustment conditions are maintained.

The monthly shock is the updated MP_median monetary-policy shock from JarocinskiKaradiUpdate2024, and the macro block is built from public monthly U.S. series in the spirit of McCrackenNg2016. The working sample is the pre-COVID overlap 1990:02--2019:08. The headline figures focus on the Moody's BAA--AAA corporate bond spread and log CPI; the main summary table also reports the 1-year Treasury rate, and the appendix adds log industrial production. For each horizon $h=0,\ldots,18$, the baseline linear local projection is \[ y_{t+h}=\alpha_h+\beta_h x_t+\gamma_h^\top C_{t-1}+u_{t,h}. \] Here $y_{t+h}$ is the macro outcome at horizon $h$, $x_t$ is the monetary-policy shock, and $C_{t-1}$ is information dated before the shock. The coefficient $\beta_h$ is the horizon-$h$ response to the shock after controlling for lagged macro conditions. The application also estimates a state-dependent specification, \[ y_{t+h}=\alpha_h+\beta_{L,h}(1-S_{t-1})x_t+\beta_{H,h}S_{t-1}x_t+\gamma_h^\top C_{t-1}+u_{t,h}. \] The indicator $S_{t-1}$ equals one when the unemployment rate is at least 6.5 percent. Thus $\beta_{H,h}$ is the response in slack periods, and $\beta_{L,h}$ is the response outside slack periods. This specification asks whether the response to monetary-policy shocks differs across macroeconomic states.

The baseline control vector $C_{t-1}$ contains six monthly lags of the dependent variable and of each of log industrial production, log CPI, the 1-year Treasury rate, the BAA--AAA spread, and the monetary-policy shock. These controls define the familiar local-projection specification. They are distinct from the regression-adjustment covariates used to study the HAC moment. The shock is scaled as in the published series; all reported responses and standard errors therefore use the units of that series. A one-unit response is a response to a one-unit movement in the published MP_median shock series, and the replication files record the raw-data scaling used in the analysis.

The regression-adjustment covariates are organized in three nested predetermined sets. The first, parsimonious set contains low-order time trends and, in state-dependent regressions, the state indicator. This set asks whether simple deterministic features of the sample explain part of the HAC moment. The second, macro-lag set adds the same lag block that appears in the local projection. This set asks whether the standard pre-shock controls already contain information about predictable moment variation. The third, rich macro set adds additional lags of equity prices, labor-market slack, the short rate, and commodity prices, implemented with the S&P 500, the unemployment rate, the federal funds rate, and a broad commodity-price series. This set asks whether broader macro-financial information available before the shock further explains the moment variation. All regression-adjustment variables are dated before the monetary-policy shock whose moment is adjusted. That timing makes them predetermined for implementation, but because lagged macro variables may themselves reflect earlier shocks, the conservative variance interpretation rests on the additional orthogonality logic in Section (ref) rather than on timing alone. The covariate sets are fixed before inspecting horizon-by-horizon significance and are not selected separately by outcome or horizon. The point is therefore not to search for the most favorable adjustment variable, but to compare economically interpretable information sets of increasing richness.

The published high-frequency series is treated as the measured assignment shock used in the design, not as an estimated first-stage object. For the reported calculations, the conditioning environment fixes the sample window, the local-projection and residualization specifications, the observed lagged-information matrices used in $C_{t-1}$, the adjustment-covariate paths, and the measurement convention for $x_t$. The re-randomized object is the current measured monetary surprise entering the moment, with HAC providing a feasible short-memory long-run-variance approximation for the resulting moment sequence. Under this convention, the calculations do not add uncertainty from constructing the high-frequency shock series. A design that conditions on the realized shock path itself would leave no shock-assignment uncertainty to approximate, while a design that treats the high-frequency construction as random would require replacing the local-projection moment by the augmented moment vector in Appendix (ref). The empirical calculations below are therefore conditional on the measurement convention for the shock series and do not require monetary surprises to be literal randomized policy experiments.

The main pattern is that richer predetermined macro covariates produce larger standard-error reductions. Table (ref) reports the corresponding empirical results. In the linear projections, the average HAC standard-error reduction for log CPI rises from 3.71% under the parsimonious set to 16.12% under the rich macro covariate set; for the BAA--AAA spread it rises from 1.57% to 22.02%; and for the 1-year Treasury rate it rises from 8.95% to 32.27%. In the slack-state CPI specification, the high-slack coefficient has the largest reduction, with a 19.96% average decline in the standard error. These reductions are consistent with the Loewner-reduction result in Proposition (ref): larger sets of predetermined covariates reduce the reported HAC variance limit when they track predictable components of the mean path. Under the stronger conditions of Proposition (ref), the same reductions can be interpreted as tightening a conservative bound for the design variance; absent those conditions, the table is a diagnostic for movement in the date-specific mean path. It identifies neither $\Omega_\mu$ nor the exact design variance; it shows that the chosen covariate span explains part of the raw HAC moment variation.

Figure (ref) reports the linear projections using the rich macro covariate set for the two main outcomes. The point estimates are essentially unchanged, while the reported bands based on the adjusted HAC limit are narrower. For the spread, at $h=5$ the estimated response is 0.015 with a HAC standard error of 0.011 and an RA standard error of 0.008. For log CPI, the largest differences occur at longer horizons; at $h=15$ and $h=18$, the adjusted-HAC pointwise interval excludes zero while the baseline HAC interval does not. These pointwise differences are descriptive evidence about how residualization changes the reported variance limit, and their interpretation as tighter valid design intervals requires the stronger conditions discussed above. Fixed-horizon simultaneous statements can use either Bonferroni or \v{S}id{\'a}k adjustments under the fixed-$J$ joint Gaussian limit in Appendix (ref); the appendix reports Bonferroni as the conservative reference.

Figure (ref) shows the state-dependent CPI responses under the slack split. The rich macro covariate set has limited effect at short horizons and a larger effect at medium and longer horizons, especially in the high-slack state. At $h=13$ the high-slack coefficient is $-0.191$ with HAC and RA standard errors 0.099 and 0.085, and at $h=18$ it is $-0.206$ with corresponding standard errors 0.122 and 0.100. In this application, regression adjustment leaves the response profile largely unchanged and narrows the reported bands where the predictable moment component appears quantitatively important, subject to the same conservative-interpretation qualification.

The appendix reports the remaining monetary results: Table (ref) adds log industrial production to the linear results, Table (ref) reports the slack and zero-lower-bound (ZLB) specifications for both CPI and the spread, and Table (ref) reports a vector autoregression with exogenous controls (VARX) impulse-response analysis on the same data. In the VARX results, the rich macro covariate set reduces average standard errors by 25.71% for cumulative inflation responses, 19.73% for BAA--AAA spread responses, 15.77% for 1-year Treasury-rate responses, and 14.86% for cumulative industrial-production responses. The appendix also reports bandwidth and covariate-set sensitivity checks. The reductions remain positive across conventional HAC bandwidth choices and are driven primarily by larger sets of predetermined macro covariates rather than by demeaning or low-order trends alone. The appendix reports selected descriptive horizons at which pointwise significance changes when regression-adjusted HAC standard errors replace baseline HAC standard errors and shows that most do not survive fixed-horizon simultaneous adjustment. The empirical conclusion concerns predictable movement in date-specific mean moments and conservative variance limits rather than uniform rejection across response horizons.

table[table omitted — 2,100 chars of source]
figure[figure omitted — 476 chars of source]
figure[figure omitted — 482 chars of source]

\FloatBarrier

Conclusion

This paper develops a design-based interpretation of time-series GMM inference. The central idea is that macroeconomic time series are often used to study what happened in one realized economy, not to average over an imagined population of economies. Conditional on the realized economic environment, the finite-history estimand is fixed and the relevant uncertainty comes only from the assignment variation in the shocks. While the design long-run variance for that uncertainty is $\Omega_R$, conventional HAC procedures generally estimate a different object, $\Omega_R^+=\Omega_R+\Omega_\mu$, where $\Omega_\mu$ is the long-run variance of the centered date-specific mean path. The decomposition shows exactly what HAC adds: besides innovation uncertainty, it also counts predictable drift in the moments across dates. Under the maintained conditions, this makes HAC conservative for scalar functions of the finite-history estimand rather than invalid.

The empirical takeaway from this paper is that baseline HAC intervals remain a useful conservative benchmark, but they need not be the end of the analysis. When predetermined variables explain predictable movement in the moment, regression-adjusted HAC can lower the conservative HAC limit, and under the long-run orthogonality conditions it can be interpreted as moving the bound closer to the design variance. Without those conditions, the reduction remains a diagnostic for predictable moment variation rather than a standalone validity claim. The simulations show why alignment matters: adjustment works well when the covariates track the mean path, though they naturally work less well when they only imperfectly span it. The monetary-shock application shows the same force in data, with rich macro information delivering material standard-error reductions and little movement in point estimates. More broadly, the paper separates uncertainty about shock assignment within the realized economy from uncertainty over alternative possible economies, and uses that distinction to clarify when familiar HAC tools are conservative and when additional structure can make them sharper. Future work could extend the framework to high-dimensional adjustment, fixed-$b$ inference, generated shocks, and overidentified GMM under local misspecification.