EconBase
← Back to paper

Forecasting on the Accuracy-Timeliness Frontier: Two Novel `Look Ahead' Predictors

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

96,893 characters · 17 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Forecasting on the Accuracy–Timeliness Frontier: Two Novel `Look Ahead' Predictors

abstractWe re-examine the traditional Mean-Squared Error (MSE) forecasting paradigm by formally integrating an accuracy–timeliness trade-off: accuracy is defined by MSE (or target correlation) and timeliness by advancement (or phase excess). While MSE-optimized predictors are accurate in tracking levels, they sacrifice dynamic lead, causing them to lag behind changing targets. To address this, we introduce two `look-ahead' frameworks—Decoupling-from-Present (DFP) and Peak-Correlation-Shifting (PCS)—and provide closed-form solutions for their optimization. Notably, the classical MSE predictor is shown to be a special case within these frameworks. Dually, our methods achieve maximum advancement for any given accuracy level, so our approach reveals the complete efficient frontier of the accuracy–timeliness trade-off, whereas MSE represents only a single point. We also derive a universal upper bound on lead over MSE for any linear predictor under a consistency constraint and prove that our methods hit this ceiling. We validate this approach through applications in forecasting and real-time signal extraction, introducing a leading-indicator criterion and tailored linear benchmarks.

Introduction

Petropoulos et al. (2022), in their comprehensive treatment of forecasting theory and practice, contend that “the theory of forecasting appears mature today … the fact that forecasting is mature does not mean that all has been done.” Building on this perspective, we propose new directions by formalizing a novel forecast accuracy–timeliness dilemma. \\

Forecasting involves several partly competing objectives: accuracy (predicting future levels), timeliness (lead–lag behavior relative to a benchmark, i.e., retardation or advancement), and smoothness (suppressing spurious high-frequency noise). In this sense, forecasting and real-time signal extraction (nowcasting) rest on common methodological foundations, as our examples illustrate.\\

Ideally, one would optimize accuracy, smoothness, and timeliness (AST) in a single objective, producing predictors/nowcasts that track the latent level, detect turning points without systematic delay, and avoid spurious high-frequency noise. The AST setting, however, implies a formal trilemma: improving any one component necessarily worsens at least one of the others. Wildi (2005) and Wildi and McElroy (2019) propose a framework based on this trilemma, but their criterion lacks a closed-form solution. Wildi (2024), (2026a), and (2026b) further study a two-way accuracy–smoothness trade-off. Here, we instead focus on the explicit accuracy–timeliness trade-off in prediction.\\

We introduce two forecasting frameworks—Decoupling-From-Present (DFP) and Peak-Correlation-Shifting (PCS)—and derive exact closed-form solutions to their respective optimization problems. Both generalize the classical minimum-MSE predictor, recovered as a special case. We prove a duality result: for any target accuracy, our predictors achieve the maximum possible phase excess (lead) relative to the MSE benchmark, yielding an efficient frontier that quantifies the accuracy–timeliness trade-off. While MSE is a single point on this frontier, DFP and PCS trace the full curve. We derive an upper bound on the maximum `meaningful' lead that can be achieved by a linear predictor under a consistency constraint; our predictors attain it, restricting hyperparameters to a valid admissible region.\\

Our method deliberately trades accuracy at selected horizons for a systematic time lead over the MSE-optimal predictor. Unlike classic forecast combinations that pool common information across competing forecasts at a fixed horizon (e.g., Hsiao and Wan, 2014; Makridakis et al., 2018), we exploit dissimilarities in predictors across different horizons. Counterintuitively, decoupling the predictor from the latest observation—usually considered the most influential—can increase the lead, and overweighting older data can extend it further. In some scenarios where timeliness is paramount, the conventional goal of minimizing forecast error can be inverted to maximize it within a controlled bound. This demonstrates how prioritizing speed can challenge conventional forecasting wisdom. We focus on stationary univariate results for brevity, but the approach extends to nonstationary and multivariate cases.\\

We begin in Section (ref) by documenting the limitations of the traditional MSE forecasting model. Section (ref) then introduces the DFP and PCS predictors, deriving closed-form solutions and a dual formulation that yields a new accuracy–timeliness efficient frontier. Lead time is defined in Section (ref), where we relate it to the key hyperparameters and derive, under a consistency constraint, an upper bound on admissible lead. Applications to time-series forecasting and signal extraction—including the construction of leading indicators and the customization of linear predictors—are presented in Section (ref). Section (ref) then summarizes the main conclusions.

Limitations of Classic Forecast Approach: a Case Study

For illustration we consider forecasting an MA($q$) process \[x_t=\mu+\gamma_0\epsilon_t+\gamma_1\epsilon_{t-1}+...+\gamma_q\epsilon_{t-q},\] where $\gamma_0=1$. To simplify exposition, we assume that $\mu=0$, that $\gamma_1,...,\gamma_q$ are known and that the white noise innovations $\epsilon_t$ are observed.\\

The MSE $h$-step ahead forecast $\hat{x}_{hT}^{MSE}$ of $x_{T+h}$ at the sample end $t=T$ is

eqnarray*[eqnarray* omitted — 159 chars of source]

For $h\leq q$ and $t\in\{q-h+1, \ldots ,T\}$, consider the forecast filter $\hat{x}_{ht}^{MSE}=\sum_{k=0}^{q-h} \gamma_{k+h}\epsilon_{t-k}$ with weights $\gamma_{h}, \ldots ,\gamma_{q}$. The cross correlation function (CCF) between $x_{t+\delta}$ and $\hat{x}_{ht}^{MSE}$ at lead $\delta$ ($q\geq \delta>0$) or at lag $\delta$ ($h-q\leq \delta\leq 0$) is given by

eqnarray*[eqnarray* omitted — 202 chars of source]

For illustration, Fig. (ref) presents the CCF for an exponentially weighted MA(9) process, $x_t=\sum_{k=0}^{9}0.9^k\epsilon_{t-k}$, together with its optimal $h=5$ step-ahead forecast, $\hat{x}_{ht}^{MSE}=\hat{x}_{5t}^{MSE}=\sum_{k=0}^{4}0.9^{k+5}\epsilon_{t-k}$, evaluated at leads and lags $\delta\in\{-4,-3,...,9\}$ (top panel).\\

By design, $\hat{x}_{5t}^{MSE}$ maximizes the CCF at the designated forecast horizon $h = 5$, as indicated by the kink at the green vertical line. The correlation also increases as the lead shrinks ($\delta \to 0$, with $0\le \delta \le 5$), peaking at 0.86 when $\delta=0$ (black vertical line). The lower panel plots $x_{t}$ (black line) and its $h=5$-step-ahead predictor, $\hat{x}_{5t}^{MSE}$ (green), for $t = 1, …, 100$. Instead of clearly leading the series, the predictor is largely time-aligned with the process it is meant to forecast. This is at odds with the intent of a truly forward-looking predictor and stems from the tight coupling between $\hat{x}_{ht}^{MSE}$ and $x_t$ implied by MSE optimality, as the CCF makes clear. Increasing $h$ does not weaken this dependence, so changing the forecast horizon does not fix the problem. Although deliberately simple, the example reflects common challenges in economic forecasting, including real-time signal extraction and trend nowcasting (see Section (ref)). Our new methods extend this framework to general forecasting problems and address the issue by explicitly incorporating a left shift (an effective lead) in forecasts for general stationary processes (extensions to nonstationary integrated processes are omitted for brevity).\\

figure[figure omitted — 369 chars of source]

We now move beyond the MA(9) example and consider a stationary linear process

eqnarray[eqnarray omitted — 80 chars of source]

where the impulse response sequence $\gamma_k$ is square-summable, $\sum_{k=-\infty}^{\infty}\gamma_k^2<\infty$. We assume iid innovations $\epsilon_t$, though the results extend to uncorrelated white noise when focusing on best linear prediction. For simplicity, we take $\epsilon_t$ to be observed\footnote{We also analyze AR inversions, but the MA impulse-response representation is generally more informative about predictor dynamics.}; otherwise, $\epsilon_t$ can be recovered using standard inversion methods (see Section (ref) for an example). Allowing $k < 0$ in (ref) (i.e., inclusion of future innovations $\epsilon_{t-k}$) accommodates signal-extraction settings with two-sided bi-infinite filters (cf. Section (ref)). Consider a generic (not necessarily MSE-optimal) predictor for $x_{t+h}$ generated by a causal finite-length filter of order $L$ with coefficient vector $\mathbf{b}=(b_0, …, b_{L-1})'$: \[ \hat{x}_{ht} = \mathbf{b}'\boldsymbol{\epsilon}_t=\sum_{k=0}^{L-1} b_k\epsilon_{t-k}, \] where $\mathbf{b}=\mathbf{b}(h)$ generally depends on the forecast horizon $h$. Define the length-$L$ coefficient vector for horizon $h$ as $\boldsymbol{\gamma}_h:=(\gamma_h,\gamma_{h+1},...,\gamma_{h+L-1})'$, which yields the (length-$L$) MSE-optimal linear predictor \[ \hat{x}_{ht}^{MSE}=\sum_{k=0}^{L-1}\gamma_{h+k}\epsilon_{t-k}. \] For brevity, we refer to the`MSE predictor at horizon $h$' interchangeably as the weight vector $\boldsymbol{\gamma}_h$ or its implied predictor $\hat{x}_{ht}^{MSE}$. When $h=0$, this reduces to the nowcast, and $\boldsymbol{\gamma}_0$ denotes the contemporaneous coefficient vector. In a stationary ARMA($p,q$) setting, the nowcast coincides with a truncated MA inversion (Wold decomposition, assuming no deterministic component). As mentioned, we restrict attention to univariate stationary processes, noting that the approaches extend to integrated and multivariate settings (not shown).

Look-Ahead Forecast Optimization Criteria

We develop two variants—Decoupling from Present (DFP) and Peak Correlation Shifting (PCS)—designed to advance the predictor relative to the benchmark MSE design (i.e., induce a controllable lead via a leftward shift).

Decoupling from Present

As established in Section (ref), the classical MSE predictor in this setting achieves its CCF maximum at the contemporaneous lag $\delta=0$, yielding a coincident rather than a leading forecast. This property is essentially invariant to increases in the forecast horizon $h\leq q$. To induce a systematic lead, we therefore seek to decouple the predictor from the present and formulate the following decoupling-from-present (DFP) problem:

eqnarray[eqnarray omitted — 212 chars of source]

where $\|\boldsymbol{\gamma}_0\|=\sqrt{\boldsymbol{\gamma}_{0}'\boldsymbol{\gamma}_{0}}$. The unit-norm constraint, $\mathbf{b}'\mathbf{b}=1$, implies that the linear objective $\boldsymbol{\gamma}_h'\mathbf{b}$ is proportional to the correlation $\rho(x_{t+h},\hat{x}_{ht})$ at $\delta=h$. This is computationally advantageous: $\rho(x_{t+h}, \hat{x}_{ht})$ is a nonlinear function of $\mathbf{b}$, whereas the objective in (ref) is linear, yielding a tractable optimization while preserving the correlation-based interpretation at the target horizon. Likewise, under the unit-length constraint $\alpha_0$ specifies the contemporaneous correlation $\rho(x_{t},\hat{x}_{ht})$ at $\delta=0$. Setting $\alpha_0=0$ achieves complete decoupling of the predictor from the present; see Section (ref) for an example. Observe that $\alpha_0=\cos(\theta_{0b})$, where $\theta_{0b}$ denotes the angle between $\boldsymbol{\gamma}_{0}$ and $\mathbf{b}$. Hence, $\alpha_0$ parametrizes the phase $\theta_{0b}$ between the predictor $\mathbf{b}$ and the nowcast $\boldsymbol{\gamma}_{0}$ at $\delta=0$, see Fig.(ref). Conversely, $\theta_{0b}$ is determined by $\alpha_0$ up to its sign. Typically, only one of the two solutions ($\pm \theta_{0b}$) is genuinely forward-looking; the alternative is backward-looking and entails a corresponding lag. Conceptually, this connects to Wildi and McElroy (2019), who frame a trilemma in terms of time-shift and amplitude functions. In what follows, we assume $|\alpha_0|\le 1$ (feasibility) and that the implied solution delivers a strictly positive objective value (strict positivity), as is typical in applications (the criterion can vanish, for example, for an MA($q$) process when $h>q$). \\

Remark: Decoupling the predictor from the present may seem counterintuitive, as the most recent observation $x_T$ at the sample end $t = T$ is typically regarded as highly informative for forecasting. However, decoupling does not imply that $x_T$ is irrelevant. Rather, it acknowledges that achieving an effective lead of the predictor requires a dynamic pattern that differs from the contemporaneous behavior of $x_T$. In this sense, a degree of decoupling—i.e., loosening the tie to the present—is intrinsic to constructing a genuinely forward-looking predictor.\\

The following Theorem provides the closed-form solution to the DFP criterion.

TheoremAssume $\boldsymbol{\gamma}_0$ and $\boldsymbol{\gamma}_h$ ($h>0$) are linearly independent and $|\alpha_0|< 1$. Then the solution to criterion (ref) lies in $\textrm{span}\{\boldsymbol{\gamma}_0,\boldsymbol{\gamma}_h\}$: \begin{eqnarray} \mathbf{{b}}=\lambda_1\boldsymbol{\gamma}_h+\lambda_2\boldsymbol{\gamma}_0 \end{eqnarray} for some scalars $\lambda_1,\lambda_2$. If $\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_h\neq 0$, then \begin{eqnarray*} \lambda_1&=&\frac{\alpha_0\|\boldsymbol{\gamma}_0\|-\lambda_2\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_0}{\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_h} \end{eqnarray*} and $\lambda_2$ is determined by the quadratic \begin{eqnarray} \lambda_2&=&\frac{-b\pm\sqrt{b^2-4ac}}{2a}, \end{eqnarray} with coefficients \begin{eqnarray*} a&=&\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_0\left(\frac{\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_0\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_h}{\boldsymbol{(\gamma}_0'\boldsymbol{\gamma}_h)^2}-1\right)\\ b&=&2\alpha_0\|\boldsymbol{\gamma}_0\|\left(1-\frac{\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_0\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_h}{(\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_h)^2}\right)\\ c&=&\frac{\alpha_0^2\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_0\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_h}{(\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_h)^2}-1. \end{eqnarray*} Among the two candidate solutions for $\lambda_2$, choose the root that maximizes the objective $\boldsymbol{\gamma}_h'\mathbf{b}$ in (ref). On the other hand, if $\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_h=0$, then \begin{eqnarray*} \lambda_2&=&\frac{\alpha_0}{\|\boldsymbol{\gamma}_0\|}\\ \lambda_1&=&\pm\sqrt{\frac{1-\alpha_0^2}{\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_h}}. \end{eqnarray*} where the appropriate sign maximizes the objective $\boldsymbol{\gamma}_h'\mathbf{b}$. Finally, if $|\alpha_0|= 1$, then the solution is $\mathbf{{b}}=\operatorname{sign}(\alpha_0)\boldsymbol{\gamma}_0/\|\boldsymbol{\gamma}_0\|$.

Proof: We first assume $|\alpha_0|< 1$. Then, under criterion (ref), the feasible set is the intersection of the unit sphere ($\mathbf{b}'\mathbf{b}=1$) and the circular cone defined by the decoupling constraint $\boldsymbol{\gamma}_{0}'\mathbf{b}=\alpha_0\|\boldsymbol{\gamma}_0\|\|\mathbf{b}\|$ (where $\|\mathbf{b}\|=1$ is relaxed), i.e., the cone with axis $\boldsymbol{\gamma}_{0}$ and semi-angle $\theta_{0b}=\arccos(\alpha_0)$. Moreover, maximizing the projection $\boldsymbol{\gamma}_h'\mathbf{b}$ implies that the optimizer—i.e., the admissible unit vectors on the cone—lies in the two-dimensional subspace spanned by the axis $\boldsymbol{\gamma}_0$ of the cone and the objective direction $\boldsymbol{\gamma}_h$ (see Fig. (ref)). Consequently, \[ \mathbf{{b}}=\lambda_1\boldsymbol{\gamma}_h+\lambda_2\boldsymbol{\gamma}_0, \] where $\lambda_1$ and $\lambda_2$ are determined by the decoupling and unit-norm constraints. Assuming $\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_h\neq 0$, the decoupling constraint implies

eqnarray[eqnarray omitted — 180 chars of source]

From the unit-length constraint we deduce

eqnarray*[eqnarray* omitted — 169 chars of source]

Substituting (ref) into the unit‑length condition and simplifying yields the quadratic (ref) in $\lambda_2$; the appropriate branch (sign) is then selected by maximizing the objective function. \\ On the other hand, if $\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_h= 0$, then the decoupling constraint implies

eqnarray[eqnarray omitted — 171 chars of source]

Inserted into the length constraint, we obtain \[ 1=\mathbf{b}'\mathbf{b}=(\lambda_1\boldsymbol{\gamma}_h+\lambda_2\boldsymbol{\gamma}_0)'(\lambda_1\boldsymbol{\gamma}_h+\lambda_2\boldsymbol{\gamma}_0)=\lambda_1^2\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_h+\lambda_2^2\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_0. \] Inserting (ref) then implies \[ \lambda_1=\pm\sqrt{\frac{1-\alpha_0^2}{\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_h}}, \] where the appropriate sign is selected to satisfy the objective (e.g., maximize alignment with $\boldsymbol{\gamma}_h $). Finally, if $|\alpha_0|= 1$, then the cone degenerates to a ray along $\boldsymbol{\gamma}_0$ and the decoupling constraint together with the unit-length uniquely determine \[ \mathbf{{b}}=\operatorname{sign}(\alpha_0)\boldsymbol{\gamma}_0/\|\boldsymbol{\gamma}_0\|, \] as claimed.\qed\\

figure[figure omitted — 1,054 chars of source]

Remark: Figure (ref) shows that, for the solution based on the root $\lambda_2<0$, the angle $\theta_{0b}$ between the DFP predictor $\mathbf{{b}}$ and the nowcast $\boldsymbol{\gamma}_0$ exceeds the angle $\theta_{0h}$ between $\boldsymbol{\gamma}_h$ and $\boldsymbol{\gamma}_0$. The difference $\theta_{0b}-\theta_{0h}>0$ represents the excess phase induced by decoupling and is directly related to the predictor’s lead under a `standard' configuration (see Section (ref)). The alternative root $\lambda_2>0$, corresponding to the branch at $-\theta_{0b}$, produces a phase-reversed solution and thus corresponds to a lag (i.e., a negative lead). Complete decoupling occurs when $\theta_{0b}=\pi/2$, in which case $\mathbf{{b}}$ is orthogonal to the nowcast $\boldsymbol{\gamma}_0$. \\

We can derive the distribution of the DFP predictor conditional on the distribution of an estimate $\boldsymbol{\hat{\gamma}}_0$ of the nowcast $\boldsymbol{\gamma}_0$. To this end, we approximate Equation (ref) with

eqnarray[eqnarray omitted — 117 chars of source]

where $\mathbf{F}$ is the $L\times L$ forward operator and $\mathbf{I}$ is the identity matrix. The matrix $\mathbf{F}$ has ones on its first superdiagonal and zeroes elsewhere. Note that the last $h$ entries of $\mathbf{F}^h\boldsymbol{\gamma}_0$ are zero, so the approximation in (ref) is valid provided $\gamma_{k}\approx 0$ for $L-1\geq k\geq L-1-h$, which holds for sufficiently large $L$ under stationarity.

CorollaryLet $\hat{\boldsymbol{\gamma}}_0$ be an estimator of $\boldsymbol{\gamma}_0$ with mean $\boldsymbol{\mu}_{\gamma_0}$ and covariance matrix $\boldsymbol{\Sigma}_{\gamma_0}$ and assume $\lambda_1,\lambda_2$ are fixed. Then the DFP predictor in Equation (ref) has mean and covariance \begin{eqnarray} \boldsymbol{\mu}_b\approx(\lambda_1 \mathbf{F}^h+\lambda_2\mathbf{I})\boldsymbol{\mu}_{\gamma_0} and \boldsymbol{\Sigma}_b\approx(\lambda_1 \mathbf{F}^h+\lambda_2\mathbf{I})\boldsymbol{\Sigma}_{\gamma_0}(\lambda_1 \mathbf{F}^h+\lambda_2\mathbf{I})', \end{eqnarray} where the approximations are valid for $L$ sufficiently large. Moreover, if $\hat{\boldsymbol{\gamma}}_0$ is (asymptotically) Gaussian, then $\mathbf{\hat{b}}$ is (asymptotically) Gaussian as well.

The result follows immediately from (ref) and standard properties of linear transformations. For the (asymptotic) distribution of the classic MSE estimate $\hat{\boldsymbol{\gamma}}_0$ of $\boldsymbol{\gamma}_0$, see, for instance, Brockwell and Davis (1993) and Cox (1990)\footnote{For an ARMA-process, the distribution of the MA-inverted weights $\gamma_k$, $k=1,...,L-1$ (noting that $\gamma_0=1$) can be derived from the distribution of the AR- and MA-estimates via the delta method as elaborated in Cox (1990), cf. Appendix (ref).}. \\

A limitation of the corollary is that it treats $\lambda_1$ and $\lambda_2$ as fixed, even though both depend on $\boldsymbol{\gamma}_0$ through Theorem (ref). We address this issue below. \\

For further consideration, we propose an alternative DFP-MSE decoupling criterion given by

eqnarray[eqnarray omitted — 187 chars of source]

This objective penalizes the squared deviation of $\mathbf{b}$ from the MSE-predictor $\boldsymbol{\gamma}_h$ and thus minimizes the mean-squared prediction error for $x_{t+h}$ under the DFP predictor. In particular, if $\alpha_0=\boldsymbol{\gamma}_{0}'\boldsymbol{\gamma}_{h}$, the solution recovers $\mathbf{{b}}=\boldsymbol{\gamma}_h$. In this formulation, the unit-norm constraint from the earlier DFP criterion can be omitted, simplifying both the geometry and the optimization, but the hyperparameter $\alpha_0$ loses its intuitive interpretation as a correlation and becomes harder to interpret. The objective captures accuracy, while lead is enforced by the decoupling constraint, thereby formalizing the accuracy–timeliness trade-off in Criterion (ref) (see Section (ref)). Although minimizing (ref) is natural, some non-standard configurations instead require maximizing the MSE (under boundedness) to obtain a sizeable lead from $\mathbf{b}$, see Appendix (ref).

PropositionAssume that $\boldsymbol{\gamma}_0$ and $\boldsymbol{\gamma}_h$ ($h>0$) are linearly independent. Then the solution to criterion (ref) is \begin{eqnarray} \mathbf{{b}}=\boldsymbol{\gamma}_h+{\lambda}\boldsymbol{\gamma}_0, \end{eqnarray} where the scalar ${\lambda}$ is determined by \begin{eqnarray} {\lambda}=\frac{\alpha_0-\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_0}{\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_0}. \end{eqnarray}

Proof: The solution to (ref) lies on the affine hyperplane defined by the decoupling constraint $\boldsymbol{\gamma}_{0}'\mathbf{b}=\alpha_0$. By the least-squares optimality principle, it is the orthogonal projection of $\boldsymbol{\gamma}_{h}$ onto this hyperplane. Because $\boldsymbol{\gamma}_{0}$ is normal to the hyperplane, the optimizer must have the form \[ \mathbf{{b}}=\boldsymbol{\gamma}_h+{\lambda}\boldsymbol{\gamma}_0 \] for some scalar ${\lambda}$, which is determined by enforcing the decoupling constraint: \[ \boldsymbol{\gamma}_{0}'(\boldsymbol{\gamma}_h+{\lambda}\boldsymbol{\gamma}_0)=\alpha_0, \] as claimed. \qed\\

When $\alpha_0=\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_0$ in (ref), we have $\lambda=0$ and $\mathbf{{b}}=\boldsymbol{\gamma}_h$ reproduces the MSE predictor, as noted earlier. The distribution of the MSE-DFP predictor follows from Corollary (ref) by setting $\lambda_1=1$ and $\lambda_2=\lambda$, treating $\lambda$ as fixed. We now extend this result to incorporate the randomness of $\lambda$, which arises when $\lambda$ is defined as a function of an estimator $\boldsymbol{\hat{\gamma}}_0$ of $\boldsymbol{\gamma}_0$, which is substituted into (ref).

CorollarySuppose $\boldsymbol{\gamma}_0$ and $\boldsymbol{\gamma}_h$ ($h>0$) are linearly independent. Let $\hat{\boldsymbol{\gamma}}_0$ be an estimator of $\boldsymbol{\gamma}_0$ with mean $\boldsymbol{\mu}_{\gamma_0}\to\boldsymbol{\gamma}_0\neq \mathbf{0}$ and covariance matrix $\boldsymbol{\Sigma}_{\gamma_0}$ with $\boldsymbol{\Sigma}_{\gamma_0}\to\mathbf{0}$ for $L\to\infty$. Then, for filter length $L$ and sample size $T$ sufficiently large, the mean $\boldsymbol{\mu}_b$ and covariance matrix $\boldsymbol{\Sigma}_b$ of the DFP-MSE (ref), based on stochastic $\lambda=\lambda(\hat{\boldsymbol{\gamma}}_0)$ in (ref), can be approximated by \begin{eqnarray} \boldsymbol{\mu}_b&\approx&(\mathbf{F}^h+{\tilde{\lambda}}\mathbf{I})\boldsymbol{\gamma}_0\\ \boldsymbol{\Sigma}_{b}&\approx&\mathbf{J}\boldsymbol{\Sigma}_{\gamma_0}\mathbf{J}' \end{eqnarray} where \begin{eqnarray} \tilde{\lambda}&=&\frac{\alpha_0-(\mathbf{F}^h\boldsymbol{\gamma}_0)'\boldsymbol{\gamma}_0}{\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_0}\\ \mathbf{J}&=&\mathbf{F}^h+\boldsymbol{\gamma}_0\partial\boldsymbol{\tilde{\lambda}}'+\tilde{\lambda}\mathbf{I}\\ {d}{\tilde{\lambda}}_k&=&-\frac{I_{\{k\geq h\}}\gamma_{k-h}+I_{\{k\leq L-1-h\}}\gamma_{k+h}}{\|\boldsymbol{\gamma}_0\|}-\frac{\alpha_0-(\mathbf{F}^h\boldsymbol{\gamma}_0)'\boldsymbol{\gamma}_0}{\|\boldsymbol{\gamma}_0\|^2}2\gamma_k \end{eqnarray} with ${d}{\tilde{\lambda}}_k$ denoting the $k$-th element of the row vector $\partial\boldsymbol{\tilde{\lambda}}'$ and $I_{\{\}}$ the indicator function. If the first weight of $\boldsymbol{\gamma}_0$ is held fixed ($\gamma_0=1$), then the first row and the first column of $\boldsymbol{\Sigma}_{\gamma_0}$ are zero; additionally, in (ref) the first column of $\boldsymbol{\gamma}_0\partial\boldsymbol{\tilde{\lambda}}'+\tilde{\lambda}\mathbf{I}$ is set to zero.

Proof: When $\lambda=\lambda(\hat{\boldsymbol{\gamma}}_0)$ in the DFP predictor is not treated as fixed, Equation (ref) follows by continuity and by using the approximation $\boldsymbol{\gamma}_h\approx\mathbf{F}^h\boldsymbol{\gamma}_0$ when $L$ is sufficiently large. To obtain the covariance matrix in (ref), we apply the delta method (see Cox, 1990) and use the first-order Taylor expansion \[ \mathbf{b}(\hat{\boldsymbol{\gamma}}_0)\approx \mathbf{b}({\boldsymbol{\gamma}_0})+\mathbf{J}( \hat{\boldsymbol{\gamma}}_0-{\boldsymbol{\gamma}_0}), \] assuming $T$ is large. The Jacobian $\mathbf{J}$ has entries

eqnarray[eqnarray omitted — 213 chars of source]

with $\tilde{\lambda}=\tilde{\lambda}(\boldsymbol{\gamma}_0)$ as specified in (ref) and $\mathbf{I}_{ij}=\delta_{ij}$ the Kronecker symbol. Hence, $\partial\tilde{\lambda}/\partial\gamma_{j-1}=d\tilde{\lambda}_{j}$ as in (ref). It follows that \[ \boldsymbol{\Sigma}_{b}\approx\mathbf{J}\boldsymbol{\Sigma}_{\gamma_0}\mathbf{J}', \] which establishes the claim. If the first component $\gamma_0$ of $\boldsymbol{\gamma}_0$ is held fixed, then the first row and the first column of $\boldsymbol{\Sigma}_{\gamma_0}$ vanish. In addition, the partial derivatives with respect to $\gamma_0$ in (ref) disappear; equivalently, the first column of $\boldsymbol{\gamma}_0\partial\boldsymbol{\tilde{\lambda}}'+\tilde{\lambda}\mathbf{I}$ is set to zero (the first row does generally not vanish because $\tilde{\lambda}$ depends on $\gamma_k$, $L>k>0$). \qed\\

Remark: an extension of Corollary (ref) to the case of random $\lambda_1,\lambda_2$ can be obtained in the same way. In that setting, however, $\lambda_i(\boldsymbol{\gamma}_0)$, for $i=1,2$—and therefore $\mathbf{{b}}=\lambda_1\boldsymbol{\gamma}_h+\lambda_2\boldsymbol{\gamma}_0$—become more involved functions of $\boldsymbol{\gamma}_0$; see Theorem (ref).\\

Both DFP predictors—those defined by criteria (ref) and (ref)—are linear combinations of $\boldsymbol{\gamma}_h$ and $\boldsymbol{\gamma}_0$, but entail distinct trade-offs. The mean-squared formulation (ref) yields a geometrically simple construction with a unique solution and is especially convenient for extensions to integrated processes (not pursued here). By contrast, (ref) induces a richer geometry and reduces to quadratics with two candidate solutions. Its chief advantage is interpretability: under the unit-norm constraint $\mathbf{b}'\mathbf{b}=1$, the decoupling parameter has a direct correlation meaning at $\delta=0$, namely $\alpha_0=\rho(x_t,\hat{x}_{ht})$, which facilitates hyperparameter selection. Criterion (ref) can be adapted to endow its hyperparameter with a correlation interpretation as well, albeit at the cost of a more complex solution structure.\\

Let $\theta_{0b}$ and $\theta_{0h}$ represent the angles formed between $\boldsymbol{\gamma}_{0}$ and $\mathbf{{b}}$, and between $\boldsymbol{\gamma}_{0}$ and $\boldsymbol{\gamma}_{h}$, respectively, see Fig.(ref). For illustration we assume the nowcast and the MSE predictor to lie in the first quadrant, so that $\theta_{0h}<\pi/2$ (the general case is addressed in Section (ref)). When $\lambda < 0$, a phase excess occurs where $\theta_{0b} > \theta_{0h}$. Intuitively, the MSE predictor’s lead over the nowcast is tied to $\theta_{0h}$ , so doing more of the same—i.e., taking $\theta_{0b}>\theta_{0h}$—further increases $\mathbf{b}$'s lead; we formalize this later, noting that this intuition does not always hold. Strict positivity $\boldsymbol{\gamma}_h'\mathbf{b}>0$ implies $\beta=\theta_{hb}<\pi/2$. In some circumstances, when $\theta_{0h}<\pi/2$ (as assumed), it is meaningful to formulate a more stringent positivity rule

eqnarray[eqnarray omitted — 60 chars of source]

for the phase excess. This guarantees that the DFP predictor correlates positively with both the $h$-step ahead MSE predictor and the nowcast (or the original process $x_t$).\\

The Appendix states formal results that relate these angles to $\lambda$ and to the hyperparameter $\alpha_0$ (Proposition (ref) and Corollary (ref)). Unlike Fig.(ref), Fig.(ref) emphasizes the so-called DFP-triangle with vertices $\mathbf{0}$, $\boldsymbol{\gamma}_h$ and $\mathbf{b}$; this configuration will be useful when connecting $\lambda$ to an effective measure of the DFP predictor’s lead (or advancement); cf. Section (ref). \\

figure[figure omitted — 380 chars of source]

Complementing the MA form (ref), an AR inversion further clarifies the structure of the DFP principle. For a stationary, invertible zero-mean ARMA process $x_t$, let $\boldsymbol{\phi}$ ($\phi_0=1$) be a finite (length-$L$) AR inversion. For sufficiently large $L$, $\boldsymbol{\phi}'\mathbf{x}_t\approx\epsilon_t$, with $\mathbf{x}_t:=(x_t,...,x_{t-(L-1)})'$. The AR-form DFP weights $\mathbf{a}=\mathbf{a}(\lambda)$ follow from convolving $\boldsymbol{\phi}$ with $\mathbf{b}(\lambda)$:

eqnarray[eqnarray omitted — 273 chars of source]

where $\boldsymbol{\phi} \cdot \boldsymbol{\gamma}_h$ is the AR-inverted MSE predictor and $\mathbf{e}_1=(1, 0, \dots, 0)$. Since for large $L$ the AR inversion $\boldsymbol{\phi}$ approximately cancels the MA inversion $\boldsymbol{\gamma}_0$, $\lambda$ affects only the lag-zero weight of the MSE predictor in the AR representation.\\

Equation (ref) neatly captures the DFP criterion's dual purpose: it matches the standard MSE predictor by replicating lags $k>0$ in $\boldsymbol{\phi} \cdot \boldsymbol{\gamma}_h$, and enforces decoupling by adjusting only lag zero in $\lambda \mathbf{e_1}$. The magnitude $|a_0|$ of the weight on $x_t$ can decrease or increase, as $\lambda<0$ (phase excess) becomes more negative, depending on the sign of the lag-zero weight, i.e., $\operatorname{sign}((\boldsymbol{\phi}\cdot\boldsymbol{\gamma}_h)_0)$, of the MSE predictor (assuming $|\lambda|<|(\boldsymbol{\phi}\cdot\boldsymbol{\gamma}_h)_0|$). Although the AR form (ref) is simple, the MA impulse response (ref) is more informative about the predictor’s dynamics, and the look-ahead predictor in the next section offers no comparably simple AR form; hence the impulse response (MA-) representation $\mathbf{b}(\lambda)$ of the prediction problem remains our focus.\\

To conclude, we briefly comment on the linear-independence assumption for the finite-length predictors $\boldsymbol{\gamma}_0$ and $\boldsymbol{\gamma}_h$ (length $L$). If it fails (i.e., $\boldsymbol{\gamma}_0\propto\boldsymbol{\gamma}_h$), the DFP problem becomes degenerate since the objective is fixed by the constraint. A practical remedy is to slightly perturb one or both of $\boldsymbol{\gamma}_0,\boldsymbol{\gamma}_h$ to restore linear independence.\footnote{If at least one entry $\gamma_k\neq 0$ for $k=L-h,\ldots,L-1$, then $\mathbf{F}^h\boldsymbol{\gamma}_0$ (approximately $\boldsymbol{\gamma}_h$) and $\boldsymbol{\gamma}_0$ are linearly independent.} These perturbations can be made arbitrarily small, leaving the performance of the underlying MSE predictor (or nowcast) essentially unchanged; we do not elaborate further for brevity. Cases where $\boldsymbol{\gamma}_0\propto\boldsymbol{\gamma}_h$ can arise naturally: when the DGP is an AR(1) process, $\boldsymbol{\gamma}_0$ is proportional to $\boldsymbol{\gamma}_h$ for all $h$; however, proportionality might also appear for periodic or seasonal structures at specific $h$; finally, collinearity could be accidental, without deeper implications for the DGP. In all such cases, the perturbation scheme can be used to induce effective leads with the modified DFP predictors.

Peak Correlation Shifting

We propose an alternative mechanism to attenuate indirectly the MSE predictor’s contemporaneous coupling with $\boldsymbol{\gamma}_0$ by relocating the peak of its CCF. As shown in Fig. (ref), the MSE predictor peaks at $\delta = 0$. By shifting this peak to the target forecast horizon $\delta=h$, the predictor becomes genuinely forward-looking. Accordingly, we formulate the Peak Correlation Shifting (PCS) problem:

eqnarray[eqnarray omitted — 218 chars of source]

where $\beta_h$ is a hyperparameter that controls the shape of the CCF. As with the DFP criterion, we assume feasibility—i.e., $|\beta_h|/\|\boldsymbol{\gamma}_{h-1}-\boldsymbol{\gamma}_{h}\|\leq 1$—and positivity.\\

For interpretation, note that in many applications the CCF attains its maximum at a lead smaller than $h$ (i.e., the peak lies to the left of $\delta=h$). As a result—often, though not always—$(\boldsymbol{\gamma}_{h-1}-\boldsymbol{\gamma}_{h})'\mathbf{b}>0$ at $\delta=h$ (cf. Fig.(ref)). The PCS criterion (ref) directly controls this term, which measures the change in the forecast CCF from lead $h-1$ to $h$. A necessary (though not sufficient) condition to move the peak to $\delta=h$ is $(\boldsymbol{\gamma}_{h-1}-\boldsymbol{\gamma}_{h})'\mathbf{b}<0$, opposite to Fig.(ref); this can be enforced by choosing $\beta_h$ in (ref). \\ In some settings, however, constraining only the single lead $\delta=h$ may be too weak to generate a meaningful peak shift. A stronger alternative is to impose the constraint over a neighborhood of leads (Appendix (ref)), but for clarity we proceed with the simple PCS specification above. \\

The proposed PCS criterion (ref) is structurally similar to the DFP formulation (ref). In particular, the solution to (ref) follows directly from Theorem (ref) upon the substitutions \[ \boldsymbol{\gamma}_0\to\boldsymbol{\gamma}_{h-1} - \boldsymbol{\gamma}_h, \textrm{~~} \alpha_0\to \beta_h/\|\boldsymbol{\gamma}_{h-1}-\boldsymbol{\gamma}_h\|, \] assuming $\boldsymbol{\gamma}_{h-1}$ and $\boldsymbol{\gamma}_h$ are linearly independent. Consequently,

eqnarray[eqnarray omitted — 233 chars of source]

where the approximation holds for sufficiently large $L$ under stationarity. The distribution of the PCS predictor can be obtained from Corollary (ref) (assuming fixed $\lambda_1,\lambda_2$) with the corresponding adjustments. \\

By analogy with (ref), we can define an MSE-PCS criterion as

eqnarray[eqnarray omitted — 214 chars of source]

where we omit the unit-norm constraint $\mathbf{b}'\mathbf{b}=1$. The solution is

eqnarray[eqnarray omitted — 212 chars of source]

and it follows from Proposition (ref) with routine adjustments. The distribution of $\mathbf{b}$ is obtained as in Corollary (ref), allowing for stochastic $\lambda=\lambda(\hat{\boldsymbol{\gamma}}_0)$.\\

Because the DFP and PCS predictors lie in $\textrm{span}\{\boldsymbol{\gamma}_0,\boldsymbol{\gamma}_h\}$ and $\textrm{span}\{\boldsymbol{\gamma}_{h-1},\boldsymbol{\gamma}_h\}$, respectively, they generally differ unless these subspaces coincide—i.e., unless $\boldsymbol{\gamma}_{h-1}$ is a linear combination of $\boldsymbol{\gamma}_{h}$ and $\boldsymbol{\gamma}_0$.\footnote{For an AR(2) process, the Yule–Walker equations imply this collinearity, so DFP and PCS coincide when $\alpha_0$ and $\beta_h$ match (not shown).} Figure (ref) illustrates the geometry of the PCS solution $\mathbf{b}$ in two cases, labeled $\mathbf{b}_1$ and $\mathbf{b}_2$, corresponding to $\beta_h=0$ (red) and $\beta_h<0$ (blue), respectively. \\ When $\beta_h=0$ , the PCS constraint implies

equation[equation omitted — 213 chars of source]

where we use $\|\mathbf{b}_1\|=1$, $\theta_{hh-1}$ is the angle between $\boldsymbol{\gamma}_{h-1}$ and $\boldsymbol{\gamma}_{h}$, and $\theta_{hb1}$ is the angle between $\mathbf{b}_1$ and $\boldsymbol{\gamma}_{h}$. \\ If $L$ is large enough and $\gamma_{h-1}\neq 0$, then

equation[equation omitted — 186 chars of source]

since for stationary processes $\gamma_{L+h-1}\to 0$ as $L\to\infty$. Substituting this inequality into (ref) gives \[ \cos(\theta_{hb1}+\theta_{hh-1})< \cos(\theta_{hb1}), \] which implies that, within the plane spanned by the two MSE predictors, $\mathbf{b}_1$ lies on the side of $\boldsymbol{\gamma}_h$ opposite to $\boldsymbol{\gamma}_{h-1}$. Equivalently, $\theta_{hb1}>0$ represents a positive phase excess of $\mathbf{b}_1$ relative to $\boldsymbol{\gamma}_h$. \\ The same qualitative conclusion applies when $\beta_h<0$ (the blue case) or when $\gamma_{h-1}=0$ in (ref).\footnote{In this case (ref) becomes (asymptotically) an equality as $L\to\infty$, implying $\theta_{hb1}>0$ still holds; strictness follows from (ref) and $\theta_{hh-1}>0$, provided the two MSE predictors are not collinear.} The appendix provides a formal link between the hyperparameter $\beta_h$ and the phase excess $\theta_{hb}$ (Proposition (ref)). \\

Finally, it is important to note that the AR-representation of the PCS predictor does not share the simple structure of the DFP predictor shown in Eq. (ref), where the influence of $\lambda$ is strictly limited to the lag-zero weight because the convolution term $\boldsymbol{\phi}\cdot\boldsymbol{\gamma}_0\approx\mathbf{e}_1$ is an identity. In the PCS predictor (ref), however, the convolution term $\boldsymbol{\phi}\cdot(\boldsymbol{\gamma}_{h-1}-\boldsymbol{\gamma}_h)$ does not resolve to the identity. Consequently, the impact of $\lambda$ is not localized at lag zero; rather, it typically propagates across all lags, both within the MA as well as within the AR representations.

figure[figure omitted — 662 chars of source]

The Construction of Leading Indicators and Benchmark Customization

The (MSE-)PCS objective in (ref) can be adapted to tasks such as constructing macroeconomic leading indicators or customizing benchmarks. Since these settings share the same structure, we focus on the former.\\

Let $\boldsymbol{\gamma}$ be the causal, length-$L$ MA inversion of the target (coincident) indicator, e.g., the Wold decomposition of a stationary macro series (such as differenced GDP), or the causal right-half of a (possibly acausal) trend/cycle filter applied to that Wold decomposition (for brevity we do not discuss multivariate designs). We then solve

eqnarray[eqnarray omitted — 217 chars of source]

Unlike the original formulation (ref), criterion (ref) does not aim to approximate the $h$-step MSE predictor. Instead, it targets the impulse response $\boldsymbol{\gamma}$ of the indicator while enforcing a constraint that moves the CCF peak $h$ periods to the right, inducing leading behavior. The solution lies in the span of $\boldsymbol{\gamma}$ and $(\mathbf{F}^{h-1}-\mathbf{F}^h)\boldsymbol{\gamma}$, and therefore generally differs from DFP and the original PCS designs.\\

Benchmark customization means tracking a chosen target $\boldsymbol{\gamma}$ as closely as possible while adjusting its timeliness to satisfy the PCS constraint. This preserves a preferred design while allowing favorable real-time fine-tuning; see Wildi (2024–2026) for smoothness customizations.\\ When $\boldsymbol{\gamma}:=\boldsymbol{\gamma}_h$, PCS customizes the MSE predictor, but $\boldsymbol{\gamma}$ can be any target filter, including standard business-cycle tools such as Hodrick–Prescott (see Section (ref)), Christiano–Fitzgerald, Beveridge–Nelson, or Hamilton filters.

Accuracy-Timeliness Dilemma

For the unit-length DFP predictor in (ref), lead (advancement) appears as a positive phase excess, $\theta_{0b}>\theta_{0h}$ in Fig. (ref) (a formal connection between lead and phase excess is established in the next Section (ref)). In this setting, the decoupling parameter and phase are linked by $\alpha_0=\cos(\theta_{0b})$. Expressing the decoupling constraint in terms of the phase renders the accuracy–timeliness trade-off explicit: accuracy is encoded in the objective, while timeliness is governed by the phase-based constraint.\\ Geometrically, as the phase excess $\theta_{0b}-\theta_{0h}$ increases (with $\theta_{0b}>\theta_{0h}$ throughout)) in Fig.(ref), the projection of $\mathbf{b}$ onto $\boldsymbol{\gamma}_h$ declines. Under the unit-norm constraint $\|\mathbf{b}\|=1$, \[ \mathbf{b}'\boldsymbol{\gamma}_h = \cos(\theta_{0b}-\theta_{0h})\|\boldsymbol{\gamma}_h\|, \] so a larger phase excess reduces the objective value $\mathbf{b}'\boldsymbol{\gamma}_h$ (and vice versa). An analogous argument applies to the PCS predictor, where the objective $\mathbf{b}'\boldsymbol{\gamma}_h=\cos(\theta_{hb})\|\boldsymbol{\gamma}_h\|$ decreases as the phase excess $\theta_{hb}$ increases (a formal link between $\theta_{hb}$ and the hyperparameter $\beta_h$ is established in Appendix (ref), Proposition (ref)). This captures the inherent dilemma between timeliness (phase excess or lead) and accuracy (target correlation or MSE).

Dual Interpretation and Efficient (Accuracy-Timeliness) Frontier

Consider the dual DFP optimization criterion

eqnarray[eqnarray omitted — 217 chars of source]

It is obtained from the primal DFP problem by interchanging the objective and the correlation constraint and switching to minimization. While the algebra is analogous, the geometry changes: the feasible cone determined by the constraints is centered on $\boldsymbol{\gamma}_h$ (not $\boldsymbol{\gamma}_0$), i.e., the cone’s boundary rays are symmetric about $\boldsymbol{\gamma}_h$.\\

Let $\theta_{primal}=\theta_{0b}$ be the angle between $\mathbf{b}$ and $\boldsymbol{\gamma}_0$ in the primal geometry (cf. Fig.(ref)). In the dual problem, the constraint fixes $\theta_{dual}=\theta_{hb}$, the angle between $\mathbf{b}$ and $\boldsymbol{\gamma}_h$. Since $\|\mathbf{b}\|=1$, the correlation constraint gives $\alpha_h=\cos(\theta_{dual})$. Geometrically, $\theta_{0b}=\theta_{0h}+\theta_{hb}$, where $\theta_{0h}$ is the angle between $\boldsymbol{\gamma}_0$ and $\boldsymbol{\gamma}_h$. Therefore, \[ \theta_{primal}=\theta_{dual}+\theta_{0h}=\arccos(\alpha_h)+\theta_{0h}~,~ \alpha_0=\cos(\theta_{primal})=\cos(\arccos(\alpha_h)+\theta_{0h}). \] Thus $\alpha_0$ (primal) and $\alpha_h$ (dual) are linked via this identity. The dual solutions corresponding to $\pm\theta_{dual}$ yield objective values \[ \boldsymbol{\gamma}_0'\mathbf{b} = \cos(\pm\theta_{dual}+\theta_{0h})\|\boldsymbol{\gamma}_0\| \] in (ref). Selecting the minimizing branch in (ref) yields $+\theta_{dual}$, matching the original DFP solution, with $\theta_{primal}=\arccos(\alpha_h)+\theta_{0h}$. Hence, when $\alpha_0:=\cos(\arccos(\alpha_h)+\theta_{0h})$ (equivalently $\alpha_h:=\cos(\arccos(\alpha_0)-\theta_{0h})$), the primal and dual formulations coincide, giving a dual interpretation of DFP: it is the `fastest' predictor—i.e., the one with maximal positive phase excess $\theta_{dual}$—subject to a prescribed tracking accuracy $\alpha_h$, measured by the correlation $\boldsymbol{\gamma}_h'\mathbf{b}/\|\boldsymbol{\gamma}_h\|$ with the MSE predictor $\boldsymbol{\gamma}_h$.\\

The equivalence of primal and dual formulations demonstrates that the DFP criterion creates an efficient frontier between accuracy (target correlation/MSE) and timeliness (phase excess/lead). Unlike the MSE predictor, which occupies only one spot on this curve, the DFP criterion spans the full frontier. The PCS approach follows a similar rationale, though it defines timeliness via a different metric (CCF shift) instead of phase excess.

Connecting Hyperparemeters and Phase-Excess with Lead

For the PCS predictor, the concept of a lead is directly associated with a peak shift in its cross-correlation function (CCF). Specifically, a rightward displacement of the CCF peak constitutes a quantifiable lead over the MSE predictor (De Jong and Nijman, 1997). Conversely, the phase excess $\theta_{hb}>0$ exhibited by the DFP predictor does not lend itself to an immediate interpretation as a temporal lead. To establish a functional relationship between $\theta_{hb}$ (or the hyperparameter $\alpha_0$) and the relative lead of the DFP predictor over the MSE benchmark, we must first define the concept of a lead. Notably, this definition is generalizable and can be applied equally to the PCS predictor.

Time-Shift at Frequency Zero

Define the (complex) transfer function of the length-$L$ filter $\boldsymbol{\gamma}=(\gamma_{0},...,\gamma_{L-1})'$ by \[ \Gamma(\omega)=\sum_{k=0}^{L-1}\gamma_{k}\exp(-ik\omega),~ \omega\in [0,\pi]\footnote{Negative frequencies $\omega\in [-\pi,0[$ can be ignored due to symmetry.}. \] Write its polar form as \[ \Gamma(\omega)=A(\omega)\exp(i\Phi(\omega)), \] where $A(\omega)=|\Gamma(\omega)|\geq 0$ is the amplitude response and $\Phi(\omega)=\textrm{arg}(\Gamma(\omega))$ is the phase response (mod$(2\pi)$). For $\omega\neq 0$, define the associated time-shift

eqnarray[eqnarray omitted — 64 chars of source]

Let $x_t=\exp(it\omega)$. The filtered output is \[ y_t=\sum_{k=0}^{L-1}\gamma_kx_{t-k}=\sum_{k=0}^{L-1}\gamma_k\exp(i(t-k)\omega)=\Gamma(\omega)x_t=A(\omega)\exp(i(t+\tau(\omega))\omega)=A(\omega)x_{t+\tau(\omega)}. \] Hence, when a sinusoid $x_t=\sin(t\omega)$ passes through the filter, the output is scaled by $A(\omega)$ and time-shifted by $\tau(\omega)$: $\tau(\omega)>0$ indicates an advance (lead,left shift) and $\tau(\omega)<0$ indicates a retardation (lag,right shift). If $A(\omega)=0$, the sinusoid at frequency $\omega$ is completely suppressed, so the notion of a time-shift at that $\omega$ is not meaningful. \\ To extend the definition to $\omega=0$, differentiate $\Gamma(\omega)$ with respect to $\omega$: \[ \dot{\Gamma}(\omega)=\dot{A}(\omega)\exp(i\Phi (\omega))+iA (\omega)\exp(i\Phi (\omega))\dot{\Phi} (\omega), \] where the dot denotes the first derivative with respect to $\omega$. Enforcing $\Gamma (0)>0$ rules out signal suppression or phase inversion, which yields $A(0)=\Gamma(0)>0$ and $\Phi(0)=0$. Since $A$ is even, $\dot{ A}(0)=0$, hence \[ \dot{\Gamma}(0)=i \Gamma(0)\dot{\Phi}(0), \textrm{ ~so~} \dot{\Phi}(0)=-i\frac{\dot{\Gamma}(0)}{\Gamma(0)}. \] Consequently, the time-shift at $\omega=0$ is obtained from

eqnarray[eqnarray omitted — 278 chars of source]

If the filter is applied to a zero-frequency linear trend $x_t=t$, then \[ y_t=\sum_{k=0}^{L-1}\gamma_{k}x_{t-k}=\sum_{k=0}^{L-1}\gamma_{k}(t-k)=t\sum_{k=0}^{L-1}\gamma_{k}-\sum_{k=0}^{L-1}k\gamma_{k}=(t+\tau(0))\sum_{k=0}^{L-1}\gamma_{k}=A(0)x_{t+\tau(0)}, \] which confirms Equation (ref). Since $\tau(\omega)$ varies with $\omega$, we fix a reference frequency $\omega_0$, taking $\omega_0=0$ to emphasize low-frequency (trend) components relevant for forecasting/nowcasting (any $\omega_0$ is possible via (ref)). We assume a `standard' scenario in which the MSE predictor leads the nowcast, $\tau_h(0) > \tau_0(0)$ (the non-standard case $\tau_h(0) \leq \tau_0(0)$ is addressed in Appendix (ref)). If, at this zero frequency, the DFP predictor also leads the MSE predictor, then $\tau_b(0) > \tau_h(0) > \tau_0(0)$; we later show how to tune $\lambda$ within $\mathbf{b}(\lambda)$ to hit a desired $\tau_b(0)$. This ordering aligns the frequency-domain triangle in the complex plane (with vertices $\mathbf{0}$, $\Gamma_h(\omega)$, $\Gamma_b(\omega)$) with the time-domain DFP triangle in Fig. (ref), simplifying the exposition; see details below. \\

Let $\Gamma_0(\omega), \Gamma_{h}(\omega),\Gamma_{b}(\omega)$ denote the transfer functions of the nowcast and the MSE and DFP predictors, with corresponding amplitude and phase functions $A_0,A_h,A_b$ and $\Phi_0,\Phi_h,\Phi_b$. Throughout this section, we assume that $\Gamma_{0}(0), \Gamma_{h}(0)$ and $\Gamma_{b}(0)$ are strictly positive at $\omega_0=0$, thereby ruling out signal suppression or `trend reversal'. In cases where this does not hold, an alternative reference frequency $\omega_0 > 0$ may be chosen, though we do not discuss that case here.

Linking Hyperparameters to the Time-Shift

We analyze the DFP-MSE predictor in (ref), given by $\mathbf{b}=\boldsymbol{\gamma}_h+\lambda\boldsymbol{\gamma}_0$, and begin by linking $\lambda$ to the phase gap $\Phi_b(\omega)-\Phi_h(\omega)$ between the DFP and MSE predictors. The same reasoning applies to the PCS predictor and is therefore omitted. Unless stated otherwise, we assume the following baseline assumptions: $\boldsymbol{\gamma}_0\not\propto\boldsymbol{\gamma}_h$ (rank two), $\theta_{0b}>\theta_{0h}>0$ (phase excess), $\Gamma_0(0)>0$, $\Gamma_h(0)>0$, $\Gamma_b(0)>0$ (no trend reversal) and $\tau_h(0)>\tau_0(0)$ (standard scenario).

PropositionLet the baseline assumptions hold and define $\beta(\omega):=\Phi_b(\omega)-\Phi_h(\omega)$. Then, for $\omega$ in a neighborhood of zero, \begin{eqnarray} \beta(\omega)=\arctan\left(\frac{\sin\left(\gamma(\omega)\right)}{a(\omega)/b(\omega)-\cos\left(\gamma(\omega)\right)}\right)\approx\frac{\gamma(\omega)}{a(0)/b(0)-1}, \end{eqnarray} where $\gamma(\omega):=\Phi_h(\omega)-\Phi_0(\omega)$, $b=|\lambda|A_0(\omega)$, and $a=A_h(\omega)$.

Proof: From $\mathbf{b}=\boldsymbol{\gamma}_h+\lambda\boldsymbol{\gamma}_0$ it follows that \[ \Gamma_b(\omega)=\Gamma_h(\omega)+\lambda\Gamma_0(\omega). \] Moreover, $\theta_{0b}>\theta_{0h}>0$ implies $\lambda< 0$ (cf. Fig.(ref)). Hence, the three transfer functions form a frequency-dependent DFP-triangle in the complex plane with side lengths $a(\omega)=A_h(\omega)$, $b(\omega)=|\lambda|A_0(\omega)$, and $c(\omega)=A_b(\omega)$\footnote{The fixed time-domain vectors $\mathbf{b}$,$\boldsymbol{\gamma}_0$ and $\boldsymbol{\gamma}_h$ in the figure are replaced by their transfer-function counterparts which now depend on $\omega$. Moreover, the abscissa and ordinate correspond to the real and imaginary axis, respectively.}. For $|\omega|$ sufficiently small (and $\lambda<0$), $\Gamma_h(\omega)$ lies between $\Gamma_0(\omega)$ and $\Gamma_b(\omega)$, i.e.,

eqnarray[eqnarray omitted — 84 chars of source]

The second inequality follows from \[ \Phi_h(\omega)\approx\dot{\Phi}_h(0)\omega=\tau_h(0)\omega>\tau_0(0)\omega=\dot{\Phi}_0(h)\omega\approx\Phi_0(\omega) \] where inequality and Taylor expansions are justified by the baseline assumptions $\Gamma_k(0)>0$ (so $\Phi_k(0)=0$) and $\tau_h(0)>\tau_0(0)$. The first inequality in (ref), $\Phi_b(\omega)\geq\Phi_h(\omega)$, is implied by $\lambda<0$ (reflecting phase excess $\theta_{0b}>\theta_{0h}$). When $\lambda<0$, the ordering in (ref) continues to hold as long as $\Phi_h(\omega)\leq\Phi_0(\omega)+\pi$.\footnote{Geometrically, for $\lambda\in]-\infty,0]$, the argument $\Phi_b$ of $\Gamma_b=\Gamma_h+\lambda\Gamma_0$ lies in $[\Phi_h,\Phi_0+\pi]$, with the upper bound attained as $\lambda\to-\infty$. If $\Phi_h>\Phi_0+\pi$, then $\Phi_b<\Phi_h$ which affects the ordering in (ref).} This requirement holds for $\omega$ sufficiently close to zero: the phases are well-defined at $\omega=0$ (all transfer functions are non-vanishing), continuous in $\omega$, and obey $\Phi(0)=0$ because the transfer functions are strictly positive at zero frequency (no trend reversal). \\ Assume first the strict ordering $\Phi_b(\omega)>\Phi_h(\omega)>\Phi_0(\omega)$ in (ref) which yields a well-defined, nondegenerate triangle. The angles opposite the sides $a,b,c$ are $\alpha(\omega)$, $\beta(\omega)=\Phi_b(\omega)-\Phi_h(\omega)$ (i.e., $\theta_{0b}-\theta_{0h}$ in Figure (ref)), and $\gamma(\omega)=\Phi_h(\omega)-\Phi_0(\omega)$ (i.e., $\theta_{0h}$). Given the side lengths $a(\omega),b(\omega)$ and their included angle $\gamma(\omega)$, we have

equation[equation omitted — 274 chars of source]

where the last step holds for sufficiently small $|\omega|$, using first-order Taylor expansions of the trigonometric terms. This approximation is well-defined because

equation[equation omitted — 87 chars of source]

otherwise $|\lambda|=\Gamma_h(0)/\Gamma_0(0)$ would (with $\lambda<0$) force $\Gamma_b(0)=0$. \\ Next, fix $\omega$ sufficiently close to zero and assume $\gamma(\omega)=\Phi_h(\omega)-\Phi_0(\omega)=0$. In this case the triangle degenerates into a line segment, so $\beta(\omega)\in\{0,\pi\}$ depending on $\lambda$. In particular, if $\lambda<0$ (phase excess) and $|\lambda|<\|\boldsymbol{\gamma}_h\|/\|\boldsymbol{\gamma}_0\|$, then the side lengths satisfy $b(\omega)<a(\omega)$, which implies $\beta_h=0$. If instead $|\lambda|\geq\|\boldsymbol{\gamma}_h\|/\|\boldsymbol{\gamma}_0\|$, then $\beta(\omega)$ is either undefined or equal to $\pi$.\\ Since phase reversal by the DFP predictor is excluded at $\omega=0$ (all transfer functions are strictly positive), the same `no phase reversal' condition, $\arg(\Gamma_b(\omega))-\arg(\Gamma_h(\omega))\neq \pm\pi$, must hold at the fixed $\omega$, due to continuity (provided $\omega$ is sufficiently close to zero). Therefore, the second possibility $\beta(\omega)=\pi$ is ruled-out and $\beta(\omega)=0$ must hold, which verifies Equation (ref) in the first degenerate case. A similar argument applies when $\beta(\omega)=\Phi_b(\omega)-\Phi_h(\omega)=0$ (second degenerate case: the triangle collapses again to a line segment): for $\lambda<0$ this implies that the side-lengths satisfy $c(\omega)<a(\omega)$, implying $\gamma(\omega)=\Phi_h(\omega)-\Phi_0(\omega)=0$, so the arctan term in (ref) is zero, completing the proof. \qed\\

If $L$ is sufficiently large, then $\boldsymbol{\gamma}_h\approx\mathbf{F}^h\boldsymbol{\gamma}_0$, so $\Phi_h(\omega),A_h(\omega)$ are fully determined by $\boldsymbol{\gamma}_0$ (up to negligible deviations). Consequently, $\beta(\omega)$ in (ref) depends only on $\lambda$ and $\boldsymbol{\gamma}_0$.\\

Remark: A key technical difficulty is that the frequency-domain triangle with vertices $\mathbf{0}$, $\mathbf{\Gamma}_h$ and $\mathbf{\Gamma}_b$ (in the complex plane) depends on $\omega$ and therefore deforms with $\omega$, unlike the fixed time-domain triangle in Fig. (ref). We thus restrict $\omega>0$ in a sufficiently small neighborhood of zero so the deformation is limited and the ordering in (ref) remains unchanged. This ordering is anchored by the baseline assumptions, which fix the limiting configuration as $\omega\to 0$ approaches the reference frequency; at $\omega=0$ the frequency-domain triangle degenerates to a line segment.\\

The proposition's (final) baseline assumption $\tau_h(0)>\tau_0(0)$ is not essential for deriving $\Phi_b(\omega)-\Phi_h(\omega)$, but it streamlines the analysis and the derivation. Appendix (ref) treats the nonstandard case $\tau_h(0)<\tau_0(0)$. Obtaining larger leads would then typically require either flipping signs (turning a maximization into a minimization) or swapping the nowcast and MSE predictor in the objective and constraint, which runs counter to common sense.\\

Next, we connect $\lambda$ to the relative shift $\tau:=\tau_b(0)-\tau_h(0)$ of the DFP-MSE predictor with respect to the MSE benchmark at frequency zero (the same approach extends to the unit-length DFP, the PCS predictors, and nonzero frequencies). A relative lead at frequency zero is produced by enforcing $\tau>0$; we now show how to choose $\lambda=\lambda(\tau)$ as a function of $\tau$.

TheoremLet the baseline assumptions hold and let $\tau_b(0)-\tau_h(0)=\tau>0$ denote a target lead of the DFP relative to the MSE predictor at frequency zero, with $\tau\neq \tau_0(0)-\tau_h(0)$. Then, $\lambda$ is determined by $\tau$ as \begin{eqnarray} \lambda=-\frac{\tau \Gamma_h(0)}{(\tau+\tau_h(0)-\tau_0(0))\Gamma_0(0)}. \end{eqnarray}

Proof: Observe that

eqnarray*[eqnarray* omitted — 384 chars of source]

where $\beta(\omega),\gamma(\omega)$ are as in Proposition (ref). In the present setting, with $\omega\to 0$, the Taylor approximation on the right hand side of (ref) is exact. The quotient is well-defined when $\Gamma_b(0)>0$; see (ref). Solving for $\lambda$ gives

eqnarray*[eqnarray* omitted — 89 chars of source]

This is well defined because, by assumption, $\tau\neq \tau_0(0)-\tau_h(0)$. Moreover, under the baseline (standard-case) assumptions, $\lambda$ must be negative to induce a lead.\\ An interesting situation occurs when $\tau_0(0)-\tau_h(0)= 0$ (which would contradict the final baseline assumption). Then the formula for $\lambda$ simplifies to \[ \lambda= -\Gamma_h(0)/\Gamma_0(0), \] so it no longer depends on $\tau$. However,

equation[equation omitted — 76 chars of source]

which contradicts the assumption $\Gamma_b(0)>0$. We may therefore assume $\tau_0(0)-\tau_h(0)\neq 0$, completing the proof. \\ It is nevertheless instructive to interpret this degenerate case geometrically. Since the derivative of phase differences \[ \dot{\Phi}_h(0)-\dot{\Phi}_0(0)=\tau_h(0)-\tau_0(0)= 0 \] vanishes at $\omega=0$, the constant and linear terms in the Taylor expansion of $\Phi_h(\omega)-\Phi_0(\omega)$ cancel (the constant term is zero because the transfer functions are strictly positive), and we obtain \[ \Phi_h(\omega)-\Phi_0(\omega)=\textrm{O}(\omega^2). \] In the DFP-triangle with vertices $\mathbf{0},\Gamma_b(\omega),\Gamma_h(\omega)$ (cf. Fig.(ref) and footnote (ref)), the angle at $\Gamma_h(\omega)$ is \[ \gamma(\omega)=\Phi_h(\omega)-\Phi_0(\omega)=O(\omega^2), \] i.e., the triangle becomes `super flat' as $\omega\to 0$ (in Fig. (ref), $\gamma(\omega)=\Phi_h(\omega)-\Phi_0(\omega)$ corresponds to $\gamma=\theta_{0h}$). If $\Gamma_b(0)> 0$ (as assumed), then $\beta(\omega)=\Phi_b(\omega)-\Phi_h(\omega)$ must be of the same order, i.e., $\beta(\omega)=O(\omega^2)$\footnote{Note that $\beta(\omega)=\pi+O(\omega^2)$ is excluded because $\Gamma_b(0)>0$ (no phase reversal towards $\omega=0$).}. This would imply \[ \tau=\tau_b(0)-\tau_h(0)=\lim_{\omega\to 0}(\dot{\Phi}_b(\omega)-\dot{\Phi}_h(\omega))=0, \] contradicting $\tau>0$. Therefore, $\Gamma_b(0)=0$, as predicted by (ref). This analysis shows that one cannot have simultaneously $\Gamma_b(0)>0$ and $\tau>0$ when $\tau_h(0)-\tau_0(0)= 0$ (which is ruled-out by our assumptions). \qed\\

Remark: The restriction $\tau\neq \tau_0(0)-\tau_h(0)$ in the theorem guarantees that the DFP predictor does not offset or eliminate the lead $\tau_h(0)-\tau_0(0)$ that the MSE benchmark has over the nowcast, which is a natural and intuitively desirable requirement.\\

For $\tau=0$, Eq.(ref) gives $\lambda=0$ since $\tau_h(0)>\tau_0(0)$ (baseline assumption), and thus $\mathbf{b}=\boldsymbol{\gamma}_h$, as expected. As $\tau\to\infty$, $\lambda\to-\Gamma_h(0)/\Gamma_0(0)$, so \[ \Gamma_b(0)=\Gamma_h(0)+\lambda\Gamma_0(0)\to 0, \] i.e., the DFP predictor asymptotically removes linear trends. Although this determines the limit of $\lambda=\lambda(\tau)$, as a function of $\tau$, one may choose $\lambda<-\Gamma_h(0)/\Gamma_0(0)$ if trend reversal is permitted; then the left-shift typically keeps increasing (see the next section), but $\tau_b(0)$ is no longer well-defined at zero frequency (Equation (ref) instead gives the time-shift of the sign-reversed filter).\\

We can also express $\alpha_0$ in the DFP constraint, as well as the phase excess $\theta_{hb}$ between the DFP predictor and the MSE predictor, in terms of $\tau=\tau_b(0)-\tau_h(0)$.

CorollaryUnder the assumptions of the theorem \begin{eqnarray} \alpha_0&=&\boldsymbol{\gamma}_0'\left(\boldsymbol{\gamma}_h+\lambda(\tau)\boldsymbol{\gamma}_0\right)\\ \theta_{hb}&=&\arccos\left(\frac{\boldsymbol{\gamma}_h'(\boldsymbol{\gamma}_h+\lambda(\tau)\boldsymbol{\gamma}_0}{\|\boldsymbol{\gamma}_h\|\|\boldsymbol{\gamma}_h+\lambda(\tau)\boldsymbol{\gamma}_0\|}\right)\approx \arccos\left(\frac{(\mathbf{F}^h\boldsymbol{\gamma}_0)'\left(\mathbf{F}^h+\lambda(\tau)\mathbf{I}\right)\boldsymbol{\gamma}_0}{\|\mathbf{F}^h\boldsymbol{\gamma}_0\|\|(\mathbf{F}^h+\lambda(\tau)\mathbf{I})\boldsymbol{\gamma}_0\|}\right) , \end{eqnarray} where $\lambda(\tau)$ is defined in (ref).

The result follows from (ref) and the identity \[ \cos(\theta_{hb})=\frac{\boldsymbol{\gamma}_h'\mathbf{b}}{\|\boldsymbol{\gamma}_h\|\|\mathbf{b}\|}=\frac{\boldsymbol{\gamma}_h'(\boldsymbol{\gamma}_h+\lambda(\tau)\boldsymbol{\gamma}_0)}{\|\boldsymbol{\gamma}_h\|\|\boldsymbol{\gamma}_h+\lambda(\tau)\boldsymbol{\gamma}_0\|}, \] using $\boldsymbol{\gamma}_h\approx \mathbf{F}^h\boldsymbol{\gamma}_0$ for $L$ sufficiently large. \\

At $\tau=0$, $\lambda(\tau)=0$ and $\cos(\theta_{0b})=\frac{\boldsymbol{\gamma}_0'\boldsymbol{\gamma}_h}{\|\boldsymbol{\gamma}_0\|\|\boldsymbol{\gamma}_h\|}$, so $\theta_{0b}=\theta_{0h}$ and $\theta_{hb}=0$ (no phase excess because $\mathbf{b}=\boldsymbol{\gamma}_h$). Since $\lambda(\tau)$ is strictly decreasing for $\tau\geq 0$, $\theta_{hb}$ in (ref) is strictly increasing. Thus the dual interpretation in Section (ref) carries over from phase excess $\theta_{hb}$ to lead $\tau$: DFP maximizes $\tau$ subject to given tracking-accuracy. By primal–dual equivalence, DFP traces the efficient frontier between timeliness $\tau$ and accuracy (max-target correlation or min-MSE). \\

The relationship between $\alpha_0$ and the lead time $\tau$ makes it possible to interpret the hyperparameter and to choose appropriate values using a formal decision rule, as illustrated in the applications below.

Limits on the Attainable Phase Excess and Lead under Strict Positivity

To illustrate, assume $\boldsymbol{\gamma}_0$ and $\boldsymbol{\gamma}_h$ are nearly (but not exactly) aligned. As $\lambda<0$ becomes more negative, $\mathbf{b}=\boldsymbol{\gamma}_h+\lambda\boldsymbol{\gamma}_0$ rotates to point almost opposite $\boldsymbol{\gamma}_h$, so pushing $\lambda$ too far to increase lead can (almost) flip the predictor’s sign. Componentwise, such sign flips can sometimes increase lead (see Section (ref)), but the limit is extreme, so a criterion is needed to rule it out.\\

A basic requirement for a meaningful DFP predictor is the strict positivity condition (Section (ref)): the target correlation must be strictly positive, $\boldsymbol{\gamma}_h'\mathbf{b}(\lambda)>0$, equivalently $\theta_{hb}<\pi/2$. Via $\lambda(\tau)$ in Equation (ref) and Equation (ref), this imposes bounds on admissible $\lambda$ (and hence $\tau$). We require \[ 0< \cos(\theta_{hb})=\frac{\boldsymbol{\gamma}_h'(\boldsymbol{\gamma}_h+\lambda\boldsymbol{\gamma}_0)}{\|\boldsymbol{\gamma}_h\|\|\boldsymbol{\gamma}_h+\lambda\boldsymbol{\gamma}_0\|}, \] which implies $\boldsymbol{\gamma}_h'(\boldsymbol{\gamma}_h+\lambda\boldsymbol{\gamma}_0)>0$. If $\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_0>0$ then \[ \lambda>-\frac{\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_h}{\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_0}=:\lambda_{\lim}, \] (with a similar bound $\lambda<\left|\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_h/\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_0\right|$ when $\boldsymbol{\gamma}_h'\boldsymbol{\gamma}_0<0$). If $\lambda_{\lim}\Gamma_0(0)+\Gamma_h(0)>0$, substituting into (ref) yields the upper bound: \[ \tau\leq-\frac{(\tau_h(0)-\tau0(0))\lambda_{\lim }\Gamma_0(0)}{\lambda_{\lim}\Gamma_0(0)+\Gamma_h(0)}=\tau_{\lim}>0. \] If instead $\lambda_{\lim}\Gamma_0(0)+\Gamma_h(0)\leq 0$, the resulting $\mathbf{b}(\lambda_{\lim})$ eliminates or reverses the trend, making $\tau_{\lim}$ ill-defined; in that case, strict positivity imposes no restriction on $\tau$ and any finite $\tau$ is admissible.\\

We contend that any predictive lead (`left-shift') achieved by violating positivity is ill-conditioned. While dropping positivity can seem to extend the forecast horizon (e.g., in periodic settings), positivity ensures the predictor’s dynamics stay aligned with those of the process at horizon $h$. Thus $\theta_{hb}<\pi/2$—equivalently the bounds $\lambda_{\lim}$ or $\tau_{\lim}$— places a hard cap on how far the DFP predictor $\mathbf{b}(\lambda(\tau))$ can genuinely outperform the MSE benchmark $\boldsymbol{\gamma}_h$ while remaining positively correlated with the target $x_{t+h}$. This coupling to $x_{t+h}$ prevents pushing the lead towards future points whose dynamics diverge from the target horizon, avoiding, among others, a sign-flip.\\

By primal–dual equivalence, for any fixed target correlation, no linear predictor can attain a larger lead than DFP. Therefore, any lead beyond $\tau_{\lim}$ must violate strict positivity, making $\tau_{\lim}$ a universal upper bound on lead among linear predictors—and DFP achieves it. The same applies to PCS, where lead manifests as a rightward shift of the CCF peak.\\

Even stronger bounds on $\lambda$ and $\tau$ can be established by enforcing the stricter positivity rule $\mathbf{b}(\lambda(\tau))'\boldsymbol{\gamma}_0 > 0$, i.e., $\theta_{0b} < \pi/2$, see Equation (ref) (details omitted).

Empirical Examples

We illustrate several aspects of the proposed designs. First, we apply the unit-length DFP criterion (ref) to the MA process introduced in Section (ref), treating $\epsilon_t$ as observed, and compare partial and complete decoupling. Second, we implement the MSE-DFP criterion (ref) for AR and ARMA processes, where the innovations $\epsilon_t$ are unobserved and are recovered by inversion. We also construct a predictor targeting a prespecified lead over the MSE benchmark (Theorem (ref)) and evaluate finite-sample performance via simulation. Finally, we apply the leading indicator PCS (ref) to business-cycle analysis using quarterly GDP and the Hodrick–Prescott filter (Hodrick and Prescott, 1997).

Illustration of the DFP-Criterion in the MA(9)-Case

We apply the unit-length DFP criterion (ref) to the MA(9) example with forecast horizon $h = 5$, treating $\epsilon_t$ as observed (see Section (ref)). Figure (ref) shows that the CCF $\rho(\hat{x}_{5t}^{MSE},x_{t+\delta})$ of the MSE predictor attains a peak of 0.86 at $\delta=0$, indicating a strong tie with the present $x_t$—i.e., a predominantly coincident predictor $\hat{x}_{5t}^{MSE}$. To mitigate this, we consider two DFP designs that attenuate the CCF at $\delta=0$: (i) partial decoupling, implemented by setting the CCF to $0.43$ (i.e., $\alpha_0 = 0.43$), and (ii) complete decoupling, imposing a vanishing CCF via $\alpha_0=0$. Solving the quadratic in (ref) and selecting the root that maximizes the objective yields the predictors

eqnarray*[eqnarray* omitted — 191 chars of source]

Figure (ref) compares predictors (top-left), CCF (top-right), and forecasts (bottom). Unlike the MSE design, the DFP predictor can place nonzero weights up to lag $k=L$ because $\mathbf{b}$ depends on $\boldsymbol{\gamma}_0$ (a similar result is reported in Wildi (2024) and (2026a), when prioritizing smoothness rather than lead). The DFP criterion enforces decoupling at $\delta=0$ (top-right) while maximizing forecast performance at the target horizon $h = 5$. Ideally, any degradation at $\delta=0$ does not propagate to $\delta=h$; the criterion is meant to minimize such spillover at the target horizon. The apparent CCF peak shift to $\delta=h$ under full decoupling (top right, red) is incidental in this example. By contrast, PCS (which also shifts the peak to $\delta=h$) differs because its CCF does not vanish at $\delta=0$ (not shown). As decoupling strengthens, the DFP forecasts (bottom) shift progressively left, indicating relative anticipation. Notably, this lead is achieved by placing more weight on higher-lag observations, which is counterintuitive. \\

This example highlights the core trade-off underlying the DFP criterion: timeliness (advancement via decoupling) versus accuracy (MSE performance at $\delta=h$). Table (ref) quantifies this trade-off through the CCF and the relative lead of the DFP design compared with the benchmark. Leads are measured as the shift at which the sample correlation between the MSE predictor and the (shifted) DFP predictor attains its maximum (cf. De Jong and Nijman, 1997).

table[table omitted — 917 chars of source]
figure[figure omitted — 484 chars of source]

Application of the Mean-Squared DFP to AR- and ARMA-Processes

We apply the mean-squared-error variant (ref) of the DFP predictor to the stationary AR(3) and ARMA(3,2) processes

eqnarray*[eqnarray* omitted — 186 chars of source]

For the first process, the characteristic polynomial admits three positive real roots. In the ARMA(3,2) case, it admits one positive real root and a complex-conjugate pair whose modulus is substantially smaller than that of the real root. Consequently, both processes are dominated by the positive real root, yielding slowly and monotonically decaying autocorrelation functions (not shown) and predominantly non-oscillatory dynamics (an application to cyclical behavior is presented in the final business-cycle example).\\

In contrast to the preceding example, the innovation sequence $\epsilon_t$ is not observed and must be recovered by inversion to obtain the Wold decompositions \[ x_{ti}=\sum_{k=0}^{\infty}\gamma_{ki}\epsilon_{t-k,i}, i=1,2. \] Because the dynamics are non-oscillatory, the MSE predictor tends to be strongly contemporaneously coupled, with a comparatively large CCF at $\delta=0$, even for fairly large forecast horizons. To attenuate this coupling, we employ the MSE DFP criterion (ref) and vary the decoupling parameter $\alpha_0$ to construct $h=5$-step-ahead MSE and DFP predictors (see Fig. (ref)). Under complete decoupling ($\alpha_0=0$), the solutions are

eqnarray*[eqnarray* omitted — 224 chars of source]

As noted, a limitation of the MSE-based DFP variant (ref) is that, in general, $\alpha_0$ lacks a correlation interpretation\footnote{That said, the optimization principle is simpler, yields a unique solution, and readily extends to nonstationary processes (not shown).} (except when $\alpha_0=0$). To address this, we can instead choose $\alpha_0$ based on a pre-specified time shift (advance) relative to the MSE predictor, as discussed later. Meanwhile, Table (ref) reports the effective CCF values at $\delta=0$ and $\delta=5$ implied by the selected $\alpha_0$ for both processes. The optimization pursues a demanding objective: minimize performance loss at $\delta=h$ while enforcing a controlled decline at $\delta=0$. This deliberate trade-off sets our method apart from classical forecasting approaches, which generally refrain from minimizing performance at zero lag—or at any lag (dual interpretation). \\

table[table omitted — 849 chars of source]

The contrast between the MSE (green) and DFP predictors for both processes in Fig.(ref) highlights distinctive properties of the DFP design: whereas the MSE predictors are smooth, strictly positive, and monotonically decaying—reflecting the influence of the dominant root of the characteristic polynomial—the DFP predictors become progressively less smooth, partially negative, and non-monotonic as decoupling increases. Specifically, as $\alpha_0$ decreases, $\lambda$ in (ref) becomes more negative, intensifying the cancellation between $\boldsymbol{\gamma}_0$ and $\boldsymbol{\gamma}_h$ in the DFP expression $\mathbf{{b}}=\boldsymbol{\gamma}_h+\lambda\boldsymbol{\gamma}_0$. The consequent down-weighting of the present relative to the future in $\mathbf{{b}}$ exposes dynamics otherwise obscured by the dominant root. Beyond the increasingly irregular patterns observed for the ARMA(3,2) case (right top panel), a damped cycle emerges in the AR(3) case (left top panel) despite all process roots being strictly positive—a configuration that typically implies noncyclical behavior (cyclical dynamics are examined in the final example of the next section).\\

We now select the hyperparameter $\alpha_0$ using a pre-specified time-shift (see Corollary (ref)) and compare the resulting procedure with the fully decoupled DFP; both are assessed against the MSE predictor. For illustration, we use the (above) AR(3) process \[ x_t=1.3x_{t-1}-0.46x_{t-2}+0.048x_{t-3}+\epsilon_t.\] We construct an $h=3$-step-ahead predictor and impose a relative time shift (advance) of $\tau=\tau_{b}(0)-\tau_h(0)=2$ with respect to the MSE predictor $\boldsymbol{\gamma}_{3}$ to obtain the new $\tau$-shifted DFP. While the particular forecast horizons used in this example are somewhat arbitrary, the conclusions carry over to any horizon. In particular, the resulting DFP predictors (based on $h=3$) largely preserve the left shift—even relative to the MSE predictor $\boldsymbol{\gamma}_{\tilde{h}}$—for arbitrarily long horizons $\tilde{h}$. See Table (ref) for an illustration with $\tilde{h}=20$.\\

Figure (ref) (top left) displays the finite-length MA inversion $\boldsymbol{\gamma}_0$, i.e., the truncated Wold decomposition of the AR(3), (black) alongside with the MSE predictor $\boldsymbol{\gamma}_{3}$ (green), the $\tau$-shifted DFP (blue), and the fully decoupled DFP (red). For ease of visual inspection, all designs are calibrated to unit-length; hence the Wold decomposition (black) does not start with a one. The associated CCFs and sample realizations are presented in the top-right and bottom panels, respectively. Table (ref) lists the time-shifts $\tau(\omega)$ and transfer functions $\Gamma(\omega)$ at $\omega=0$, along with the hyperparameters $\lambda$ and $\alpha_0$ ($\Gamma(0)$ is based on predictors scaled to unit-length). For comparison, we also evaluate an additional 20-step-ahead MSE predictor to confirm that increasing the forecast horizon only marginally changes the time shift. Hence, the `hyperparameter' $h$ (the forecast horizon) cannot be used to address a left shift (lead) of the MSE predictor. The fully decoupled DFP is obtained by setting $\alpha_0=0$, with $\lambda$ computed from (ref). The $\tau$-shifted DFP uses $\tau=2$, with $\lambda$ and $\alpha_0$ given by (ref) and (ref), respectively. \\

Because the fully decoupled DFP undergoes phase reversal at zero frequency ($\Gamma(0)<0$), its time shift is not well-defined and is omitted from the table.\footnote{In particular, the formula for $\tau(0)$ returns the shift of the sign-reversed predictor.} While the $\tau$-shifted DFP advances trends by $\tau=2$ relative to the MSE predictor, the fully decoupled version does not preserve the direction of a linear trend.\\

This inversion also reverses any nonzero mean. Yet, even as $\lambda$ becomes more negative—and this potentially undesirable effect becomes more pronounced—Figure(ref) (bottom panel) shows that the fully decoupled predictor still extends lead time for our zero-mean AR(3) process. Because the DFP criterion limits this sign reversal to specific frequencies, we accept the trend/mean inversion as the cost of gaining increased phase excess, i.e., lead time, under zero-mean assumptions. Pushing $\lambda$ too negative, however, risks violating the positivity condition, which mandates a strictly positive target correlation, see Section (ref). While our proposed predictors satisfy this condition (Fig. (ref), middle-right), the plot indicates that decreasing $\lambda$ past the point of full decoupling risks bordering on, or effectively violating, the assumption. Thus, in this example, the fully decoupled design is near-maximal for positivity-preserving lead time.\\ Broadly speaking, generating a lead in stationary mean-reverting processes requires key components to flip signs early to anticipate the coming reversion. The DFP criterion facilitates this by optimally weighting these components—specifically the low-frequency ones in our example—demonstrating efficacy even in non-periodic contexts. However, in such challenging scenarios, the bounds on $\lambda$ must be considered carefully if strict positivity is deemed relevant.\\

Enforcing accurate tracking of a nonzero mean (or linear trend)—i.e., not merely avoiding mean inversion but matching the level—would require adding constraints to the MSE-DFP problem (ref), thereby trading off lead (for a given target correlation) or target correlation (for a given lead).\footnote{Exact mean tracking requires $\Gamma_b(0)=\Gamma_h(0)$ (first-order); exact trend tracking additionally requires $\dot{\Gamma}_b(0)=\dot{\Gamma}_h(0)$ (second-order); see McElroy and Wildi (2016).}. Meanwhile, the unit-length DFP (ref), which optimizes correlation, ignores the mean entirely and cannot recover a fixed nonzero level without an ex post adjustment. \\ To handle (stationary) nonzero-mean series, we propose a two-step procedure: first maximize lead for a given target correlation (or vice versa), even if this entails mean inversion; then apply a static level correction. See Heinisch et al. (2026) for an application to multi-step GDP forecasting.\\

table[table omitted — 1,197 chars of source]
figure[figure omitted — 574 chars of source]
figure[figure omitted — 623 chars of source]

Application of PCS to Real Time Business-Cycle Analysis

figure[figure omitted — 256 chars of source]

Business cycle analysis characterizes fluctuations in macroeconomic activity either as deviations from a smooth potential-growth path (classical cycle) or via movements in growth rates (growth-cycle). Both perspectives require a reliable estimate of the underlying trend. We extract the trend (growth) component of quarterly U.S. real GDP using the Hodrick–Prescott (HP) filter (Hodrick and Prescott, 1997), with the conventional quarterly smoothing parameter $\lambda_{HP} = 1600$. The data, retrieved from the FRED database (https://fred.stlouisfed.org/), are displayed in Fig. (ref). The sample covers the last three recession episodes, spanning 1992-01-01 to 2024-04-01. \\

The HP filter is intrinsically acausal, two-sided, and bi-infinite. In real time, however, dependence on future observations is infeasible, complicating estimation. To accommodate this constraint, the HP filter admits a finite, one-sided approximation for nowcasting that tracks the two-sided design optimally under suitable data-generating assumptions (McElroy, 2008; Cornea-Madeira, 2017). Building on this framework, Wildi (2024) develops a customized concurrent (real-time) HP filter that explicitly trades off accuracy and smoothness to improve nowcast reliability. We extend this approach by applying the PCS principle to induce a predictive lead relative to the MSE benchmark. \\

We construct a leading indicator using the PCS variant in (ref), targeting a one-year lead ($h=4$) and specifying $\boldsymbol{\gamma}=\boldsymbol{\gamma}_0$ as the (causal, finite-length) HP(1600) trend nowcast applied to the log-differenced series (Fig. (ref), bottom-left panel). Note that this leading- indicator PCS is not a classic forecast, since the target is the nowcast, albeit with its CCF shifted. \\ We evaluate the PCS objective over a range of hyperparameter values $\beta_h=\beta_{4}$ and compare the resulting filters with the classical one-year MSE predictor under a white-noise assumption (justified by the bottom-right panel in Fig. (ref)). All predictors are constructed with length $L=50$.\\

Figure (ref) compares the MSE predictor (green) with (leading-indicator) PCS-based predictors for several values of $\beta_{4}$ (top left panel); all designs are normalized to unit-norm. As intended, a CCF peak shift to $\delta=4$ is achieved at $\beta_{4}=0$, with predictor \[ \mathbf{{b}}:=\boldsymbol{\gamma}-8.81(\mathbf{F}^{3}-\mathbf{F}^4)\boldsymbol{\gamma} \] (violet line), where $\boldsymbol{\gamma}=\boldsymbol{\gamma}_0$ is the indicator underlying the leading-design construction (HP nowcast). \\ The cross-correlation functions (top right panel) indicate that, as $\beta_{4}$ approaches zero, the peak moves toward $\delta=h=4$ (violet line). The standardized filter outputs (bottom panel) likewise show that the predictive lead of the PCS predictors increases as $\beta_{4}$ decreases.

figure[figure omitted — 460 chars of source]

Conclusion

We address the inherent accuracy–timeliness conflict in prediction by embedding it in an objective–constraint optimization framework. A formal definition of predictor `lead' maps this dilemma to an explicit MSE–lead trade-off.\\

We propose new look-ahead designs: i) unit-length and MSE-DFP, which decouple the forecast from the nowcast, and ii) unit-length, MSE, and leading-indicator PCS, which shift the CCF peak toward the forecast horizon. Closed-form solutions show that these designs are generally geometrically distinct, lying in different subspaces of the predictor space. We also derive their distributions, linking them to standard MSE-predictor theory through appropriate transformations. \\

For clarity, we focus on stationary univariate processes with full-rank predictor spaces, though the framework extends to nonstationary or multivariate settings and to singular (rank-deficient) spaces. We illustrate the approach in time-series forecasting and real-time signal extraction, including leading-indicator design. In each application, a single hyperparameter controls the accuracy–timeliness trade-off and admits clear statistical and geometric interpretations, tracing a new accuracy–timeliness efficient frontier. Whereas the MSE predictor occupies a single spot on this curve, our approaches span the entire frontier.\\

Future work could combine this framework with Wildi’s recent results (2024–2026) to develop a unified `prediction trilemma' that jointly balances accuracy, timeliness, and smoothness.