Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
96,159 characters · 20 sections · 69 citation commands
Who Saw It Coming? Historical Experience and the 2021 Inflation Forecast Failure
\noindentKeywords: Inflation forecasting, regime change, historical analogy, experience-based learning, expectations anchoring, large language models
\noindentJEL Classification: C22, C53, D84, E31, E37
\thispagestyle{empty} \setcounter{page}{1}
The 2021 U.S.\ inflation surge was not anticipated by the vast majority of forecasters. Econometric models, machine learning algorithms, professional surveys, and central bank projections all predicted moderate inflation for the year ahead, while the realized average reached almost 7 percent. The consequences of missing such a regime change extend beyond forecast accuracy: delayed recognition of a supply-driven surge postpones the monetary response and requires steeper corrective action later. Recognizing a regime change promptly is therefore a first-order policy problem, but it is precisely what standard forecasting tools are not designed to do, since adaptive methods detect departures from the prevailing regime only after observing enough data from the new one.
The cost of this delayed recognition was substantial. GiannonePrimiceri2024 show that demand forces were the primary driver of the post-COVID inflation surge. Figure (ref) illustrates this using a bivariate SVAR estimated on pre-pandemic data: setting demand shocks to zero from 2021:Q2 to 2021:Q4 reduces peak CPI inflation by approximately 2.5 percentage points. This counterfactual neutralizes only the demand component and does not account for the additional moderating effect that better-calibrated expectations would have had on the pass-through of supply shocks to prices and wages. GiannonePrimiceri2024 estimate that the Fed's accommodative stance contributed roughly 3 percentage points to peak inflation; CominJohnsonJones2023 and GagliardoneGertler2024 report comparable estimates. The American Rescue Plan of March 2021 was calibrated under the assumption that inflation would remain low. Younger households remained anchored to Great Moderation dynamics throughout 2021, while experienced cohorts adjusted earlier WeberEtAl2024. At the firm level, YotzovBloomBunnMizenThwaites2024 estimate a 30% pass-through from CPI changes to firm own-price expectations, a channel that is active only when inflation is elevated and media attention is high. Earlier recognition of the regime change would have moderated demand through monetary tightening, fiscal restraint, and faster expectation adjustment, while also reducing the pass-through of supply shocks to firm pricing decisions.
Figure (ref) previews the main result. Panel (a) shows U.S.\ CPI inflation from 1960 to 2024: the 1970s oil-shock episode and the 2021 surge, separated by nearly four decades of stable and predictable inflation during the Great Moderation. Panel (b) zooms in on 2021. The ARMA(1,1) estimated on the full sample produces a flat forecast around 3.5 percent, missing the surge entirely; the shaded area between this forecast and the realized path is the cost of anchoring to Great Moderation dynamics. The same model re-estimated on 1970s data tracks the realized inflation far more closely. Household expectations mirror this pattern: SCE respondents over 60, whose lifetime includes the 1970s, reported expectations near 6 percent by mid-year, while respondents under 40 remained below 5 percent throughout.
I show that, for the class of models considered, forecast errors are primarily driven by sample composition rather than functional form misspecification. Estimation samples dominated by the Great Moderation underweight rare but economically relevant regimes, so that full-sample estimators converge to normal-regime parameters. The same mechanism operates on the expectations side: agents whose entire experience falls within the Great Moderation form priors that are well anchored to a low-inflation regime. Such anchoring is ordinarily a sign of credible monetary policy, but it becomes a source of forecast bias when the regime changes, because the experiential prior actively resists the inflationary signal rather than merely lacking information about it. Historically informed adjustments, based on past episodes with similar macroeconomic conditions, can generate substantial improvements in forecast accuracy when such regimes are identifiable in real time. These results identify the source of the failure and show that a simple, real-time implementable correction was sufficient to close most of the forecast gap.
The analysis proceeds in three parts. The first documents the forecasting failure across econometric models, machine learning methods, and institutional projections, and formalizes the source of the bias. Age-cohort data from the NY Fed Survey of Consumer Expectations show that respondents over 60, whose lifetime includes the 1970s, reported inflation expectations near 6 percent by mid-2021, while respondents under 40 remained below 5 percent throughout the year. FOMC statement sentiment, Wall Street Journal coverage, and Google Trends data all confirm that inflation attention remained low through the first months of 2021, consistent with anchored priors resisting the early inflationary signal.
The second shows that two simple adjustments to an ARMA benchmark, an intercept correction borrowing 1978--1979 forecast errors and a similarity re-estimation on 1973--1980 data, substantially close the forecast gap. The intercept correction yields 6.83 percent versus the realized 6.93. A kernel-weighted variant using oil price growth as a state variable confirms this result through a data-driven criterion. The 1970s analogy rests on real-time observable signals: by December 2020, supply-chain pressure had reached unprecedented levels, commodity prices were rising, and monetary policy was accommodative. These findings extend beyond headline CPI: the same framework applied to eight additional U.S.\ price indices confirms that the historically informed methods dominate unadjusted benchmarks, including for volatile producer price series where conventional methods miss by an order of magnitude.
The third part investigates the experience channel through a controlled experiment with two LLMs (Claude and ChatGPT), each conditioned on three professional economist personas (experienced, young, and neutral) at four data vintages. The object of interest is the persona gap $\Delta_t$, the difference in forecasts between the experienced and young personas. Under a common training leakage assumption, $\Delta_t$ is identified from the LLM outputs regardless of the training corpus. The experience-based learning model of MalmendierNagel2016 yields two testable implications: $\Delta_t > 0$ at early vintages and a narrowing of $\Delta_t$ as data accumulate. Both patterns are observed. The similarity and kernel-weighted methods and the MN experienced persona produce forecasts of the same functional form, iteration of a perceived AR(1), and differ only in how they weight historical observations, providing a formal link between the statistical and experiential approaches.
The paper makes three contributions. First, it identifies sample composition as the primary source of the 2021 forecast failure and shows that the bias can be corrected using economically motivated historical analogies. Second, it documents that the same experience channel operates across statistical methods, household expectations, and large language models: the source of the prior mattered more than the sophistication of the model. Third, it offers an interpretation of the forecast failure as the cost of expectations anchoring: full-sample estimation and experience-based priors formed entirely within the Great Moderation both assign negligible weight to high-inflation dynamics, producing the same downward bias when a supply-shock regime materializes.
\paragraph{Related literature.} A large body of work has sought to decompose the 2021--2024 inflation surge into supply and demand components.\footnote{See diGiovanni2022 and Shapiro2022 on the supply side, and BenignoEggertsson2023 on the demand side.} An emerging consensus is that supply disruptions initiated the episode while demand pressures, amplified by fiscal stimulus and accommodative monetary policy, sustained it BlanchardBernanke2023, BallLeighMishra2022, GiannonePrimiceri2024. The present paper does not attempt such a decomposition; it takes the mixed supply-demand nature of the episode as given and asks how a forecaster could have exploited historical analogies to improve real-time predictions.
A large literature studies forecasting under structural instability. Intercept corrections ClementsHendry2006, window selection and observation weighting PesaranTimmermann2007, PTP2013, DendramisKapetaniosMarcellino2020, and kernel-weighted estimation LPU2022 all address the problem of parameter shifts in the estimation sample. The present paper uses these tools for a different purpose: not to adapt to a break that has already occurred in the sample, but to borrow information from a past regime that resembles the one anticipated at the forecast origin. The distinction matters because the 2021 regime change occurs in the forecast period, not in the estimation window, so that standard break-detection methods are uninformative. In a related direction, GouletCouloumeGobelKlieber2024 develop tools to decompose ML forecasts into contributions from individual training observations. Their approach identifies which historical data points matter most for a given prediction; the present paper imposes this selection directly by targeting a specific historical episode.
The paper also relates to the growing literature on expectations and forecasting during the post-COVID episode. Reis2026 shows that a Phillips curve augmented with household expectations predicts the 2021--2024 inflation path well, and BriandMarcellinoStevanovic2025 show that media-based attention measures improve inflation forecasts during high-inflation episodes. Both exercises use a recursive design in which the information set is updated each period. The present paper adopts a fixed-origin, multi-step design closer to ForoniMarcellinoStevanovic2022: the information set is frozen at December 2020 and forecasts cover horizons 1 through 12 without updating, a harder exercise, but arguably more relevant for real-time decision-making when the key question is the full trajectory over the coming year. On the expectations side, HajdiniKurmann2026 show that regime shifts under rational expectations generate predictable forecast errors whose sign depends on the persistence of the realized regime relative to agents' expectations. Their model accounts for the 1970s forecast errors through time-varying monetary policy credibility, but rejects for the post-2020 episode. The present paper takes a different approach: rather than testing whether forecast errors are consistent with rational expectations, it asks whether the forecasting failure can be traced to sample composition and whether historically informed estimation can close the gap.
Finally, the LLM experiment contributes to a nascent literature on LLM-based economic forecasting fariae2024ai, carriero2024macro, zarifhonarvar2025inflation, LundgaardHansenetal2025. The persona-based design addresses the look-ahead bias concern raised by fariae2024ai and LundgaardHansenetal2025 through a common training leakage assumption: because the training corpus, data, and instructions are identical across personas, any contamination from post-2021 outcomes shifts all persona outputs by the same amount and cancels in the persona gap. This complements the no-leakage and validation-based strategies formalized in LudwigMullainathanRambachan2025.
The rest of the paper is organized as follows. Section (ref) documents the forecasting failure across surveys, institutional projections, and household expectations. Section (ref) develops the historically informed approaches and presents the main empirical results. Section (ref) reports the LLM experiment and its formal connection to experience-based learning. Section (ref) concludes.
This section documents the breadth of the forecasting failure across professional surveys, institutional projections, and household expectations.
Table (ref) documents the forecasting failure across several survey-based and institutional forecasts. Because these sources differ in target definition (CPI inflation vs.\ unit costs), updating frequency (quarterly survey rounds vs.\ daily nowcasts), and horizon, the table is not a controlled comparison; its purpose is to illustrate the breadth of the miss across the forecasting landscape. The first column reports the realized annualized monthly CPI inflation, which averaged 6.93 percent in 2021; the remaining columns show that no forecasting source came close to this figure. The Survey of Professional Forecasters averages 3.89 percent, the Blue Chip consensus 3.37 percent, the CBO 2.65 percent, and the FOMC projections 3.83%. Even the Cleveland Fed's inflation nowcast, which is among the most timely available, updated daily using high-frequency data on gasoline and oil prices alongside monthly CPI and PCE releases, averages only 4.49 percent.\footnote{The Cleveland Fed nowcasting model combines four components: core inflation extrapolated from recent trends, food price inflation, and gasoline price inflation derived from daily Brent crude and weekly retail gasoline prices, which are then integrated into a final CPI or PCE nowcast. See \url{https://www.clevelandfed.org/indicators-and-data/inflation-nowcasting}.} The FOMC, despite its comprehensive access to macroeconomic information and internal modeling resources, projected only 3.83% on average, roughly half the realized outcome. Firms fared no better: the Atlanta Fed's Business Inflation Expectations survey averages only 2.82 percent for the year.\footnote{The Atlanta Fed survey asks firms about expected changes to unit costs over the next 12 months rather than expected inflation, making it not directly comparable to the household surveys discussed in Section (ref).} These nowcasts incorporate progressively more information as the year unfolds; by the fourth quarter, forecasters have observed most of the 2021 inflation path, yet the annual averages still fall well short of the realized outcome. The failure across this diverse set of forecasters suggests that the challenge in 2021 was not confined to any particular model or institution, but reflected a broader difficulty in recognizing that the economy had entered a regime poorly represented in recent experience.
Table (ref) also reveals an important temporal pattern. Most surveys and institutions began revising their projections upward only during the second half of 2021, well after the inflation surge was already under way, consistent with gradual learning as new data arrived. At the same time, forecaster disagreement rose sharply: the SPF interquartile range widened from 1.39 in 2021Q1 to 2.27 in 2021Q3, its highest level in the sample, before narrowing to 1.25 in 2021Q4 as the consensus consolidated around elevated inflation. The Blue Chip dispersion follows a similar pattern, peaking at 1.5 in June 2021 when the first large CPI surprises arrived, then compressing steadily to 0.2 by December. The simultaneous rise in forecast levels and forecast disagreement is characteristic of signal extraction in a noisier environment: as the inflationary signal strengthened, forecasters updated at different speeds depending on how much weight they placed on the new information relative to their prior beliefs.
Figure (ref) provides corroborating evidence. Panel (a) shows that the FOMC statement sentiment index of Gardner2021 remained dominated by labor market concerns through the first months of 2021. Inflation sentiment begins to rise around May but does not dominate until early 2022. Panel (b) tells a consistent story: the Google Trends index for “inflation” and the Wall Street Journal attention measure of BriandMarcellinoStevanovic2025 both remained flat through the first four months of 2021, then rose sharply after the May CPI release. In both cases, inflation attention lagged the inflationary signal by several months.
The forecasting failure therefore did not stem from a lack of information: the relevant signals were available in real time. What was missing was the capacity to recognize that the economy had entered a different regime before it materialized in the data. If the information existed but was not exploited, the natural question becomes: who did recognize the regime change?
Figure (ref) decomposes one-year-ahead inflation expectations from the NY Fed's Survey of Consumer Expectations by age cohort. Panels (a) and (b) show that respondents over 60, the cohort whose lifetime experience includes the high-inflation years of the 1970s and early 1980s, consistently reported higher expectations than respondents under 40. The expectation gap, which had fluctuated around 0.5 percentage points before 2020, widened sharply in early 2021 as the over-60 cohort revised upward faster than younger respondents, peaking at nearly 2 percentage points by mid-year. The pre-existing level difference is itself a prediction of the experience-based learning model: agents whose estimation sample includes the 1970s carry a permanently higher perceived mean, which positions them closer to the realized outcome when a supply shock materializes. In February 2021, before any large CPI surprise had materialized, the over-60 median was already 4.0 percent compared to 3.0 for the under-40 group. The gap then gradually narrows through 2022--2023 as all age groups converge toward elevated expectations, and by 2024 it has essentially vanished, consistent with the experience-based learning framework of MalmendierNagel2016: having now lived through their own supply-driven inflation episode, younger cohorts acquired the experiential prior that older respondents already possessed from the 1970s.
The experience channel has a natural counterpart: expectations anchoring. The under-40 cohort's low expectations in early 2021 did not reflect a mere absence of information about high-inflation regimes. Rather, these expectations were the rational output of an experiential prior built entirely on the Great Moderation, a period in which inflation targeting was credible and deviations from the target were small and transitory. In the MalmendierNagel2016 framework, this cohort's lifetime-weighted estimate of the inflation process concentrates on low-volatility, low-mean parameters. The resulting anchoring is well calibrated in normal times, but fragile to regime change: when a supply shock materializes, the prior assigns high probability to a transitory deviation and low probability to a persistent shift, producing forecasts consistent with the “transitory inflation” view that prevailed among professional forecasters and policymakers through mid-2021. Reis2022anchor documents that long-run inflation expectations remained anchored throughout the episode, a finding consistent with the present argument. If this interpretation is correct, then the same anchoring that preserved long-run credibility may have contributed to the delayed recognition of the regime change at shorter horizons. HajdiniKurmann2026 formalize a closely related mechanism: under rational expectations with Markov-switching regimes, agents who assign high persistence to a low-inflation regime systematically under-predict inflation when a high-inflation regime materializes. Their framework shows that this under-prediction is not a departure from rationality but a direct consequence of the perceived regime transition probabilities. The experience-based learning model adds a layer to this result: the transition probabilities themselves are shaped by lifetime experience, so that agents whose sample contains only the Great Moderation perceive the low-inflation regime as near-absorbing.
Panel (c) of Figure (ref) reports inflation uncertainty by age group. This measure is constructed from each respondent's individual density forecast: the SCE asks respondents to assign probabilities to predefined intervals of future inflation, and a parametric density is fitted to these probabilities. The reported uncertainty is the spread of this individual density, capturing how confident each respondent is in their own forecast, rather than the cross-sectional disagreement among respondents.\footnote{See Armantier2017 for a detailed description of the SCE methodology and the construction of individual density forecasts.} Throughout the sample, respondents over 60 report lower uncertainty than younger cohorts, typically 1 to 2 percentage points below the under-40 group. This pattern persists during the 2021 surge: the over-60 cohort maintains the lowest uncertainty even as it reports the highest point expectations. The combination of higher expectations and tighter individual density forecasts is consistent with more precise signal extraction by cohorts whose experiential prior encompasses a high-inflation regime.
This pattern is consistent with the experience-based learning framework of MalmendierNagel2016 and PedemonteTomaVerdugo2025: agents who have lived through high inflation overweight that experience in forming expectations. The experience channel documented here for U.S.\ age cohorts has a cross-country counterpart. WeberEtAl2024 show through randomized control trials that households in high-inflation countries are already attentive to inflation and do not respond to information treatments, while those in low-inflation countries do, confirming that experience shapes attention. The genesis of the present paper illustrates the point.
Having experienced the Yugoslav hyperinflation firsthand, and having just applied historical analogies to forecast the COVID recession in ForoniMarcellinoStevanovic2022, I used the same approach in May 2021 with the 1970s as the reference episode, producing forecasts far above any available at the time. Whether this reflects a well-calibrated prior or coincidence is what the rest of the paper investigates. SalleGorodnichenkoCoibion2025 provide direct evidence on the causal channel: individuals who recall having lived through past inflationary episodes report higher inflation expectations and lower uncertainty. Even a simulated experience of historical inflation dynamics in a laboratory setting creates a “pseudo-lifetime experience” that durably shifts participants' beliefs, confirming that the link between memory and expectations is causal rather than merely correlational.
Together, these findings suggest that if lived experience improves household inflation expectations, a forecaster should be able to exploit the same logic systematically, borrowing information from the most relevant historical precedent rather than treating all past observations equally.
This section develops the historically informed forecasting approaches and presents the main empirical results. The baseline predictive model is an ARMA(1,1) for annualized monthly CPI inflation:
where $y_t = 1200(\log(CPI_t)-\log(CPI_{t-1}))$.
The choice of an ARMA(1,1) specification is supported by several strands of the literature. Lutkepohl1987 shows that marginalization of a finite-order VAR generally yields an ARMA process for each individual series, and DufourStevanovic2013 extend this result to factor models, showing that the implied univariate process for any observable, including inflation, is an ARMA rather than a pure AR. Empirically, NgPerron2001 document the presence of a large moving-average root in U.S. inflation, while StockWatson2007 show that the MA component has grown in importance since 1984. Finally, ForoniMarcelliniStevanovic2019 confirm that including an MA term significantly improves inflation forecasts in a mixed-frequency setting. In a comprehensive out-of-sample forecasting exercise covering a large number of models and a very long evaluation period, KotchoniLerouxStevanovic2019 find that the ARMA(1,1) is among the best-performing specifications for U.S.\ inflation forecasting.
The key premise is that the ARMA benchmark can be improved by drawing on historical analogues: past episodes whose macroeconomic characteristics resemble those of the current environment. Rather than treating the entire postwar sample symmetrically, the historically informed approaches developed below selectively borrow information from periods of supply-driven inflation, most notably the 1970s oil-shock era, to adjust the baseline forecasts.
To see why full-sample estimation produces biased forecasts when a rare regime occurs, consider the ARMA(1,1) within each regime $S \in \{N, C\}$:
where regime $N$ denotes the stable, low-inflation dynamics of the Great Moderation and regime $C$ denotes a supply-disruption episode with $\alpha_C > \alpha_N$. When the sample contains $T_N \gg T_C$ observations, the full-sample MLE is dominated by regime-$N$ parameters:
and the $h$-step-ahead forecast at a regime-$C$ origin is biased downward:
The same structure applies to expectations formation. An agent whose lifetime falls entirely within regime $N$ computes a lifetime-weighted mean that converges to $\alpha_N$ by the same averaging logic as (ref). The analogy suggests that anchored expectations and full-sample estimation share a common structure: both assign negligible weight to regime-$C$ parameters when regime-$C$ observations are rare in the relevant sample, whether that sample is a statistical estimation window or a lifetime of experience.
An important feature of the 2021 problem is that the regime change occurs in the forecast period, not in the estimation sample: as of December 2020, the most recent observations still belong to regime $N$. Standard break tests are uninformative: they detect breaks within the observed sample, not breaks that have not yet produced data. Adaptive methods face the same limitation: Markov-switching models assign negligible probability to regime $C$ when recent data are from regime $N$, and TVP models that update through the anti-inflationary 2020 observations are pulled in the wrong direction. More generally, any model that weights historical observations uniformly inherits this bias. For the model classes examined in this paper, the issue is not functional-form misspecification but sample composition.
The historically informed approaches address this bias by reweighting the sample toward regime-$C$ observations. The intercept correction shifts the forecast by the bias observed during a reference episode; the similarity approach re-estimates all parameters on a regime-$C$ subsample; the kernel-weighted estimator uses a continuous weighting scheme that nests the full-sample estimator and the similarity approach as limiting cases. All three require the forecaster to identify, on economic grounds, that the current environment resembles a past supply-disruption regime.
Section (ref) established that full-sample estimation produces biased forecasts when a rare regime occurs, and that the bias can be corrected by reweighting the sample toward regime-$C$ observations. The practical question is which historical episode to use as the reference. The relevant comparison for the post-COVID inflation surge is not the low-and-stable inflation era of the Great Moderation, but rather the earlier period shaped by the first oil shock. Figure (ref) illustrates the parallel.
The first oil crisis shares several key features with the post-pandemic environment: a major adverse supply disruption (the OPEC embargo in 1973, global supply-chain bottlenecks in 2020--2021), rapidly rising commodity and energy prices, accommodative monetary policy at the onset, and substantial uncertainty about the persistence of the inflationary pressures. Figure (ref) makes this parallel visible by plotting annualized monthly CPI inflation alongside the WTI oil price level for each period. Panel (b) additionally includes the Global Supply Chain Pressure Index (GSCPI) of the Federal Reserve Bank of New York. In both episodes, the oil price rises sharply before or concurrently with the inflation acceleration, providing a real-time signal that was available to a forecaster willing to draw the historical analogy. Among the candidate historical episodes, the 1970s stand out as the most informative reference, the only one combining supply-side origins with a prolonged inflationary dynamic.\footnote{The Korean War inflation spike of 1950--1951 was sharp but short-lived and driven primarily by a demand surge from military mobilization. The 2008 financial crisis was a demand-driven contraction in which inflation fell rather than rose. Neither episode shares the supply-disruption-plus-persistence configuration of the 1970s and the post-COVID period.}
The 1970s analogy assumes that the post-COVID inflation surge was driven by familiar macroeconomic forces operating with unusual intensity, not by fundamentally new mechanisms. StockWatson2025 provide direct support for this view, showing that the dynamics of the COVID period can be explained by a single aggregate shock transmitted through standard channels. Similarly, MoranStevanovicSurprenant2026 find that a COVID-specific factor adds no explanatory power for the 2021--2023 inflation surge beyond what standard macroeconomic shocks, notably accommodative monetary policy, already capture. If the inflationary forces at work were not new but rather a recurrence of a known configuration, then borrowing parameters from the most relevant historical precedent is a well-motivated forecasting strategy.
The post-COVID episode is not a pure replay of the 1970s. GiannonePrimiceri2024 show that the post-pandemic surge reflects both supply disruptions and a demand rebound that outpaced a still-constrained supply side. If anything, this makes the 1970s analogy conservative: a forecasting adjustment based on the 1970s captures the supply-driven component while potentially underestimating the additional contribution of demand.
The choice of the 1970s as a reference episode connects directly to the evidence on household expectations presented in Section (ref). SCE respondents over 60, the cohort whose lifetime encompasses the oil-shock era, reported elevated inflation expectations from the outset of the 2021 episode, while younger cohorts remained anchored to the Great Moderation. The historically informed approaches can be viewed as a statistical counterpart to this behavior: they selectively draw on the most relevant historical precedent rather than treating all past observations equally.
An important question is how a forecaster would identify the relevant historical analogy in real time. In the present case, the selection was grounded in signals observable as of December 2020: the GSCPI had reached 4.0 standard deviations above its historical mean, oil prices were recovering at a rate comparable to the early stages of the 1973 OPEC embargo, and monetary and fiscal policy were both exceptionally accommodative. These conditions jointly pointed to a supply-disruption configuration for which the 1970s provided the closest, and essentially the only, precedent in the postwar U.S.\ sample. The kernel-weighted estimator partially automates this step by using the oil price growth rate as a state variable, allowing the data to determine the relevant weighting without manual window selection. The approach does require the forecaster to identify the nature of the shock, but this requirement is shared by any conditional forecasting method. The sensitivity analysis of Appendix (ref) shows that the gains are robust to a broad range of window choices and reference periods.
The ARMA(1,1) is estimated by maximum likelihood on monthly data from 1960M01 to 2020M12 using the January 2021 vintage of the FRED-MD database. Conditional on the December 2020 observation, unadjusted forecasts are constructed iteratively up to 12 months ahead:
with $\widehat{y}_{T+1|T} = \hat\alpha + \hat\rho\, y_T + \hat\theta\, \hat\varepsilon_T$, where $\hat\varepsilon_T$ is the last in-sample residual filtered recursively through (ref). The moving-average term contributes to the one-step-ahead forecast but vanishes at longer horizons because future residuals are unobservable and set to zero. All iterative forecast equations below follow the same convention: the MA component enters at $h=1$ and the recursion is purely autoregressive for $h \geq 2$.
\paragraph{Intercept correction.} Let $e_h^{ref} = y_{\tau+h} - \widehat{y}_{\tau+h|\tau}$ denote the out-of-sample forecast error at horizon $h$ from a reference historical episode, where $\tau$ marks the forecast origin of that episode and forecasts $\widehat{y}_{\tau+h|\tau}$ are produced from an ARMA(1,1) estimated on data up to $\tau$. The intercept-corrected forecast is:
This amounts to shifting the conditional mean of the forecast distribution by the bias observed during the reference regime, while leaving all autoregressive and moving-average parameters unchanged. The reference episode used here is 1978M02--1979M01, chosen because the inflation dynamics of that period, moderate but rising readings against a backdrop of supply pressures and energy price increases, closely resemble the configuration observed in early 2021. This follows the intercept correction framework of ClementsHendry1996, who show that forecast failure under structural breaks is typically driven by shifts in deterministic components, which a correction term applied to the intercept can offset DAgostinoGambettiGiannone2013.
\paragraph{Robust intercept correction.} The intercept correction described above relies on a single reference window (1978M02--1979M01), which involves an element of arbitrariness. A more robust variant estimates the ARMA(1,1) once on the pre-shock sample 1960M02--1973M01 and holds its parameters fixed. For each forecast origin $\tau$ in the window 1973M01--1980M12, the fixed-parameter model is used to produce $h$-step-ahead forecasts, and the forecast error $e_h^{\tau} = y_{\tau+h} - \widehat{y}_{\tau+h|\tau}$ is recorded. The robust correction is the average error across all origins:
and the adjusted forecast is $\widetilde{y}_{T+h|T}^{IC\text{-}R} = \widehat{y}_{T+h|T} + \bar{e}_h$. Because the ARMA parameters are estimated before the inflationary regime begins, the forecast errors $e_h^{\tau}$ measure the systematic bias of a “normal-times” model when applied to a supply-shock environment. Averaging over 96 origins smooths out the idiosyncratic month-to-month variation that makes the single-window correction sensitive to the exact choice of $\tau$.
\paragraph{Similarity approach.} The second approach re-estimates the ARMA(1,1) on a subsample more representative of an inflationary regime associated with large supply disturbances:
where $(\hat\alpha^{SIM}, \hat\rho^{SIM}, \hat\theta^{SIM})$ are estimated on the subsample 1973M01--1980M12 and the $h=1$ forecast includes the MA term as in (ref), with the residual $\hat\varepsilon_T$ obtained by filtering the full history $y_1, \ldots, y_T$ through the subsample parameters. Using the December 2020 observation as the initial condition, the model generates 12-month-ahead forecasts from this historically tailored parameterization. Unlike the intercept correction, this approach allows all model parameters to differ from their full-sample estimates, effectively treating the inflationary subsample as a distinct data-generating process. This is closely related to the idea that, under structural breaks, forecast accuracy can be improved by restricting the estimation window to observations drawn from the current regime PesaranTimmermann2007. The similarity approach is related to the $k$-nearest neighbor forecasting framework of DendramisKapetaniosMarcellino2020, who weight observations by proximity to the current state.
\paragraph{Kernel-weighted similarity estimation.}
The intercept correction and similarity approaches both require the forecaster to select a discrete reference window. A natural generalization replaces the binary inclusion rule with a continuous weighting scheme. The kernel-weighted ARMA(1,1) parameters are obtained by maximizing the weighted log-likelihood
where $\ell_t$ denotes the Gaussian log-likelihood contribution of observation $t$ and the weight is $w_t = K\!\left((z_t - z_T)/b\right) / \sum_{s} K\!\left((z_s - z_T)/b\right)$, with $K(u) = \exp(-u^2/2)$ a Gaussian kernel and $b > 0$ a bandwidth parameter. The state variable $z_t$ is the annualized monthly growth rate of the WTI crude oil price. Observations from periods with oil price dynamics similar to December 2020 receive high weight, including the 1970s oil shocks, without requiring manual window selection.\footnote{The bandwidth is selected by weighted cross-validation, with a minimum effective sample size of 30 PTP2013, LPU2022. The growth rate is preferred to the log-level because the oil price at end-2020 was moderate in absolute terms, whereas the rate of increase was comparable to the 1973--74 OPEC embargo.}
The kernel-weighted forecasts are constructed iteratively from the December 2020 initial condition, following the same convention as (ref):
with the MA term included at $h=1$. Proximity is defined in the space of economic conditions rather than calendar time, so that the 1970s observations receive high weight because oil price dynamics at those dates resemble those at the forecast origin.
All three approaches follow the logic of ForoniMarcellinoStevanovic2022: when the economy is hit by an unusual shock, forecasts can be improved by borrowing information from historically similar episodes. The regime classification is imposed by the forecaster rather than learned from observations that have not yet materialized.
A natural objection is that adaptive or data-rich models might have captured the regime change without manual intervention. I compare the historically informed forecasts against two classes of competing models.
\paragraph{Markov-switching AR(1).} A two-regime Markov-switching autoregressive model Hamilton1989:
where $s_t$ follows a first-order Markov chain with transition probabilities $p_{ij} = \Pr(s_t = j \mid s_{t-1} = i)$. The intercept, autoregressive coefficient, and innovation variance all switch between a low-inflation regime ($s_t = 1$) and a high-inflation regime ($s_t = 2$). Forecasts are constructed by integrating over the filtered regime probabilities at the forecast origin. The model detects regime changes from the data, but it requires observations from the new regime before updating: at the onset of the inflation surge, the filtered probability assigned to the high-inflation state was negligible.
\paragraph{TVP-AR(4).} A time-varying parameter autoregressive model estimated via Kalman filter, following the framework of Hall2026:
where $\boldsymbol{x}_t = (1, y_{t-1}, \ldots, y_{t-4})'$ and $\boldsymbol{\eta}_t \sim \mathcal{N}(\boldsymbol{0}, Q_t)$. The $Q$-matrix governs the degree of parameter variation and is activated at the suspected break date (2020M01) with $Q = 0.01 \cdot I_5$ and set to zero elsewhere. Hall2026 show that this approach outperforms both rolling windows and recursive OLS when the forecaster has prior information about a break location, the situation confronting an inflation forecaster in 2020--2021. Unlike the intercept correction, the TVP parameters adjust through the Kalman filter as new data arrive rather than being imposed ex ante from a historical analogy.
To assess whether data-rich methods could have detected the inflation regime change, I construct a real-time forecasting exercise using the FRED-MD macroeconomic database McCrackenNg2016, which provides a balanced panel of approximately 128 monthly U.S.\ macroeconomic and financial series with real-time vintages. The predictor set includes four lags of all FRED-MD series as well as $K = 8$ principal components extracted from the stationarity-transformed panel and their four lags, together with four lags of the target variable.
Five models are estimated on this predictor set. The first three (Ridge regression, LASSO, and Elastic Net) are penalized linear regressions. The penalization addresses the high dimensionality of the predictor set induced by the inclusion of the full FRED-MD panel alongside its principal components. The remaining two models (Random Forest and a feedforward neural network) are nonlinear and can capture interactions and threshold effects that the linear models miss. GouletCoulombeLerouxStevanovic2022 show that these nonlinear methods deliver the largest forecasting gains in macroeconomic applications, and GouletCoulombeMarcellinoStevanovic2021 document that they were particularly useful at the onset of the COVID-19 recession.
Table (ref) summarizes the hyperparameter grids for each model. Tunable hyperparameters are selected by five-fold cross-validation with expanding-window splits that respect the temporal ordering of the data. Forecasts are produced using a direct approach: for each horizon $h = 1, \ldots, 12$, a separate model is estimated with the target $y_{t+h} = f(\mathbf{Z}_t) + \varepsilon_{t+h}$, where $\mathbf{Z}_t$ collects the predictors described above.
Table (ref) reports the realized path of CPI inflation in 2021 alongside the forecasts produced by each method. Every historically informed approach substantially outperforms the full-sample ARMA benchmark (3.52 percent). The robust intercept correction yields 5.00 percent, the similarity approach 6.19, the single-origin intercept correction 6.83, and the kernel-weighted ARMA 7.82, against a realized average of 6.93. The robust IC is the most conservative because averaging over the full similarity window includes periods in which inflation fell, so that positive and negative forecast errors partially cancel. The kernel-weighted estimate exceeds the others because the selected bandwidth ($b^* = 23.18$) concentrates weight on episodes of rapidly rising oil prices (see the sensitivity analysis in Appendix (ref)).
The evaluation focuses on the annual average because the policy-relevant question in early 2021 was whether the year as a whole would be a high-inflation year, not the precise month-by-month path. The monthly forecasts in Table (ref) exhibit sizeable month-to-month errors that partially offset in the annual average, a feature common to all fixed-origin, multi-step forecasting exercises. The historically informed methods improve accuracy precisely at the annual-average horizon that matters for the regime-recognition question: would a forecaster standing in December 2020 have classified 2021 as a supply-shock year rather than a continuation of the low-inflation regime?
Table (ref) confirms that the subsample and kernel-weighted parameters differ materially from the full-sample estimates: the intercept nearly doubles (from 0.49 to 0.83--0.98) and persistence increases ($\hat{\rho}$ rises from 0.86 to 0.90--0.93), reflecting the inflationary dynamics of oil-shock episodes. This result confirms that the 1970s analogy, which motivates the intercept correction and similarity approaches, can be recovered by a purely data-driven criterion, provided the state variable captures the dynamics rather than the level of the supply shock.
The adaptive models perform no better, and in some cases worse, than the ARMA benchmark, for the reasons anticipated in Section (ref). The MS-AR(1) produces an average of 3.29 percent: as of December 2020, the filtered probability of a high-inflation state is negligible, and the model cannot signal a regime it has not yet observed. The TVP-AR(4), with the break date set at 2020M01 and $Q = 0.01$, produces only 1.73 percent, substantially worse than the unadjusted ARMA, because the Kalman filter absorbs the deflationary readings of 2020 and adjusts parameters in the wrong direction.\footnote{In a one-step-ahead recursive exercise where the Kalman filter is updated with each realized observation during 2021, the TVP-AR(4) with $Q = 0.01$ produces an average of 4.88 percent, an improvement over the fixed-origin forecast but still well below the historically informed approaches.}
The ML models produce an average forecast of only 3.00 percent, below even the parsimonious ARMA(1,1), despite exploiting 128 predictors from the FRED-MD panel, cross-validated hyperparameters, and both linear and nonlinear specifications GouletCoulombeLerouxStevanovic2022, GouletCoulombeMarcellinoStevanovic2021. Combined with the failure of the Markov-switching and TVP models, of the professional surveys (SPF, Blue Chip), and of institutional projections (FOMC, CBO, Cleveland Fed), this evidence suggests that the forecasting failure of 2021 was not caused by using the wrong model, but by estimating any model on a sample in which the relevant regime is underrepresented.
The gains are not specific to the ARMA(1,1) specification: Table (ref) in Appendix (ref) shows that the intercept correction and similarity approach produce comparable improvements when applied to AR(1), AR(4), and ARMA(2,1) base models.
The fixed-origin results of the previous subsection condition on information available in December 2020. A natural question is whether data-driven methods can close the gap once the inflation surge becomes visible in the data. To investigate this, Figure (ref) reports an expanding-window exercise in which each model is re-estimated on successive real-time vintages of the FRED-MD database, from January to September 2021, and used to forecast the remaining months of the year.
The ARMA(1,1) forecasts (panel a) update gradually but never exceed approximately 5 percent even by the September vintage. The kernel-weighted ARMA (panel b) produces elevated forecasts from the January vintage but is unstable across subsequent vintages because the oil price growth rate fluctuates sharply from month to month. The MS-AR (panel c) hovers around 3--4 percent across all vintages, while the TVP-AR reacts aggressively at short horizons from April onward before rapidly mean-reverting. The ML models (panel d) are initially anchored near 3 percent, then react at short horizons from the May vintage onward as the CPI surprises enter the estimation sample, but still project a rapid return toward pre-pandemic levels.
Across all panels, data-driven methods either fail to anticipate the regime change or cannot sustain the signal beyond the short horizon. This mirrors the gradual adjustment observed in the survey and institutional forecasts documented in Section (ref). In contrast, the historically informed approaches produce elevated forecasts from the outset, because the 1970s analogy is imposed before any 2021 data are observed.
These results are robust to the choice of reference window (Appendix (ref)). For the similarity approach, 34 out of 56 window combinations produce average 2021 forecasts between 5.5 and 8.5 percent, all above the ARMA benchmark. The intercept correction is more sensitive to the specific origin, but origins spanning 1977M09 to 1979M04 consistently yield forecasts between 6 and 8 percent. A historical out-of-sample exercise applying the same methodology to the 1979 inflation surge produces a forecast within 0.5 percentage points of the realized outcome.
The results documented above are based on headline CPI inflation. To assess whether the approach extends to other price measures, we apply the same framework, with identical hyperparameters, estimation windows, and forecast origin, to the PCE price index (PCEPI), core CPI (CPILFESL), and six PPI series spanning the price chain from raw materials to finished consumer goods.
Table (ref) reports the results. Across all eight price indices, the forecast closest to the realized value is produced by either the similarity approach (five series) or the kernel-weighted ARMA (three series). The unadjusted ARMA, TVP-AR, and MS-AR never produce the best forecast. For consumer prices, the similarity approach closely tracks the realized values for both PCEPI and core CPI. For producer prices, the gains are larger: the ARMA benchmark misses by an order of magnitude for the most volatile series, while the historically informed methods remain within range. The ML models, available for five of the eight series, confirm the pattern observed for headline CPI: they produce forecasts comparable to or below the unadjusted ARMA and never approach the historically informed methods. Detailed monthly forecasts for each series are reported in Appendix (ref).
These results confirm that the gains from historically informed methods extend across the price chain. For the remainder of the paper, we focus on headline CPI, the most widely monitored measure and the object of the survey and institutional forecasts documented in Section (ref).
Sections (ref) and (ref) showed that experienced agents and historically informed methods both outperform their counterparts anchored to the Great Moderation, but the evidence is observational: the SCE age decomposition does not control for confounders, and the statistical methods do not involve expectations. This section uses large language models to test the experience channel in a controlled setting. By assigning an “experienced” persona (with professional memory of the 1970s) and a “young” persona (shaped by the Great Moderation) to each LLM, while holding the data, instructions, and macroeconomic context fixed, the exercise isolates the effect of the experiential prior. fariae2024ai show that LLMs generate inflation forecasts that rival the SPF, and LundgaardHansenetal2025 construct synthetic forecaster personas that often outperform human experts, motivating the persona-based design adopted here.
I adopt the notation of LudwigMullainathanRambachan2025. Let $\Sigma$ denote a finite alphabet and $\Sigma^{*}$ the set of all finite-length strings over $\Sigma$. A large language model is a mapping $\hat{m}(\cdot\,;\mathcal{C}):\Sigma^{*}\to\Sigma^{*}$ from input strings (prompts) to output strings (responses), parameterized by its training corpus $\mathcal{C}$. The researcher constructs a prompt $r \in \Sigma^{*}$ and obtains the LLM output $\hat{m}(r;\mathcal{C}) \in \Sigma^{*}$.
In the present application, each prompt $r$ encodes three objects: a data vintage $\mathcal{D}_t$ (a subset of the FRED-MD release available at date $t$), a set of forecasting instructions $\mathcal{I}$ (identical across personas), and a persona $\mathcal{P}_i \in \{\mathcal{P}_E,\mathcal{P}_Y,\mathcal{P}_0\}$. The prompt is constructed as
where $\rho:\Sigma^{*}\times\Sigma^{*}\times\Sigma^{*}\to\Sigma^{*}$ is the prompt-construction function that concatenates the three inputs into a single string. The LLM output is a vector of monthly inflation forecasts
Because repeated runs of the same prompt do not return identical outputs, $\hat{\boldsymbol{\pi}}^{\,i}_t$ is treated as a random vector. Define its conditional mean as
The expectation in (ref) conditions on the two objects that vary across the experiment, the data vintage $\mathcal{D}_t$ and the persona $\mathcal{P}_i$, and marginalizes over the training corpus $\mathcal{C}$, which is fixed but unknown to the researcher.
I construct three forecaster personas for each LLM that receive identical data and instructions but differ in their assigned professional history:
The exercise is repeated at four data vintages (January, April, July, and October 2021), each providing the corresponding FRED-MD release.\footnote{For instance, the April 2021 vintage of FRED-MD contains CPI data through March 2021, so annualized inflation for January--March is computed directly from the data, while April--December must be forecast.} For each of the $3 \times 4 = 12$ persona--vintage combinations, five independent runs are conducted per LLM, yielding $5 \times 12 \times 2 = 120$ forecast paths. The Claude model used is Claude Opus 4 (Anthropic, 2025 vintage); the ChatGPT model is ChatGPT-4o with a Pro subscription (OpenAI, March 2026 vintage).
In terms of the notation introduced above, this design generates 120 output vectors $\hat{\boldsymbol{\pi}}^{\,i}_t = \hat{m}(r^i_t;\mathcal{C})$, one for each run. Because the data vintage $\mathcal{D}_t$ and the instructions $\mathcal{I}$ are held fixed across personas, the only source of cross-persona variation in the prompt $r^i_t = \rho(\mathcal{D}_t, \mathcal{I}, \mathcal{P}_i)$ is the third argument $\mathcal{P}_i$. For a given LLM, the sample average across the five runs estimates the corpus-conditional expectation, which includes a training leakage component addressed in Section (ref).
The personas are defined as professional economists rather than consumers. The SCE evidence in Section (ref) already documents the experience channel among households; the LLM experiment targets professional forecasters, for whom no comparable age-stratified survey exists. LLMs are trained predominantly on professional and academic text, making them better suited to simulate professional reasoning. The forecasting instructions are deliberately generic to avoid triggering differential recall from the training corpus.
The design is motivated by the experience-based learning framework of MalmendierNagel2016. In their model, each agent $i$ recursively estimates a perceived law of motion for inflation using only the data observed during their professional lifetime starting at date $s_i$. The two personas weight the same historical inflation data differently: the experienced agent's estimate incorporates the high-inflation episodes of the 1970s, while the young agent's is shaped predominantly by the low and stable inflation of the Great Moderation.
The object of interest is the persona gap:
the difference in conditional mean forecasts between the experienced and young personas at vintage $t$, where $t$ indexes the information vintage (January, April, July, or October 2021), not calendar time. If the persona assignment has no effect on the LLM output, $\Delta_t = 0$. A positive $\Delta_t$ indicates that the experienced persona produces higher inflation forecasts, consistent with the experience channel: an agent whose professional memory includes the 1970s perceives inflation as more persistent and mean-reverting to a higher level than an agent shaped by the Great Moderation.
A fundamental concern with retrospective LLM forecasting exercises is look-ahead bias: both models were trained on data that includes the realized 2021 inflation outcomes and the extensive subsequent commentary crane2025total, alam2026chatmacro. In the language of LudwigMullainathanRambachan2025, this is training leakage: the LLM's training corpus $\mathcal{C}$ overlaps with the researcher's evaluation sample.
To formalize this concern, recall that the conditional mean $\mu^{i}_t = \mathbb{E}[\hat{\boldsymbol{\pi}}^{\,i}_t \mid \mathcal{D}_t, \mathcal{P}_i]$ marginalizes over the training corpus $\mathcal{C}$: it averages over all corpora that could have produced the LLM. In practice, however, the researcher faces a single realized corpus $\mathcal{C}$, one that almost certainly contains the 2021 inflation outcomes. Conditioning on this realized $\mathcal{C}$ may shift the LLM's expected output relative to the unconditional mean. Define the training leakage bias for persona $i$ at vintage $t$ as
The first term is the expected output of the LLM that the researcher actually uses, trained on the specific corpus $\mathcal{C}$. The second term is what that expectation would be if the corpus were unknown. When $\lambda^{i}_t = 0$, the corpus is uninformative about the LLM output given the prompt. While this holds in expectation over possible corpora (by iterated expectations), it need not hold for the realized $\mathcal{C}$. Since the realized corpus contains the 2021 inflation data, $\lambda^{i}_t \neq 0$ for the individual persona outputs.
Assumption (ref) permits arbitrary contamination ($\lambda_t$ may be large) but requires that it be common across personas. Three observations motivate this restriction. The corpus $\mathcal{C}$ is a property of the LLM, not of the prompt: it is identical across personas. The data vintage $\mathcal{D}_t$ and instructions $\mathcal{I}$, which constitute the vast majority of the prompt, are held fixed by design. And the persona description $\mathcal{P}_i$ is a short biographical paragraph referring only to pre-2020 events.
The threat to Assumption (ref) is not training leakage per se, but a treatment--contamination interaction: if the persona description triggered differential recall from the corpus, for instance by activating persona-specific subsets of training documents that also contain the realized 2021 outcomes, then $\lambda^{E}_t \neq \lambda^{Y}_t$ and persona differencing would not eliminate the bias. The experienced persona prompt explicitly mentions the OPEC embargo, the stagflation of 1974--1980, and the Volcker disinflation, terms that also appear frequently in ex-post analyses of the 2021 inflation episode. If these keywords activate training documents that discuss both the 1970s and the 2021 outcome, the leakage could be persona-specific. Three considerations mitigate this concern. First, the keywords in the experienced persona describe pre-2020 events that appear in textbooks and historical surveys regardless of their connection to 2021; the young persona prompt similarly invokes the “Great Moderation” and the “2008 financial crisis,” which also feature in post-2021 commentary. Second, the persona gap narrows across vintages (from 1.70 pp in April to 0.38 pp in October), a pattern that would be difficult to generate through differential recall, since the contamination from the realized outcome is fixed across vintages while the LLM's information set is expanding. Third, the neutral persona, which contains no experiential keywords, produces forecasts between the experienced and young personas at every vintage, consistent with the interpretation that the persona assignment induces genuine variation in the weighting of historical information rather than differential access to future outcomes.
Proposition (ref) complements the two strategies formalized in LudwigMullainathanRambachan2025: enforcing a no-leakage condition ($\lambda^{i}_t = 0$) for prediction problems, or collecting a validation sample to debias estimates. The leakage in the present application also operates through a different channel than theirs: not through overlap between the prompts and the training data, but through $\mathcal{C}$ containing the realized outcomes the LLM is asked to forecast. Persona differencing is a third possibility that exploits the experimental structure of the prompt to eliminate the bias without restricting the level of contamination and without a validation sample.
The neutral persona ($\mathcal{P}_0$) is not part of Proposition (ref) and is not needed for identification, but it provides three pieces of evidence against prompt priming, as documented in the results below. First, the ordering $\mu^{Y}_t \le \mu^{0}_t \le \mu^{E}_t$ holds at every vintage, even though the neutral prompt contains no experiential keywords. If the persona gap were driven by keyword priming, the neutral persona should produce forecasts comparable to one of the two treatments or erratic across vintages; instead, its forecasts fall consistently between them. Second, the neutral persona produces forecasts above the young persona at every vintage. Since the neutral prompt contains no hawkish content, this implies that the young persona's low forecasts reflect active anchoring to the Great Moderation induced by its biographical conditioning, not the mere absence of priming. Third, the convergence of all three personas toward the realized outcome as vintages accumulate is consistent with Bayesian updating on incoming data, a pattern that fixed prompt priming cannot generate.
To derive testable implications from Proposition (ref), the conditional mean is decomposed into a common component and a persona-dependent component:
where $c_t$ captures all elements common to both personas at vintage $t$, $\pi^{\mathrm{exp},i}_t$ is the experience-based forecast component for persona $i$, and $\phi_t$ is a scalar loading. The experience-based learning model of MalmendierNagel2016 provides a natural parameterization of $\pi^{\mathrm{exp},i}_t$. In their framework, each agent $i$ recursively estimates a perceived AR(1) law of motion for inflation,
with age-dependent updating:
where $b_{\tau,i}=(\alpha_{\tau,i},\beta_{\tau,i})'$ and the gain depends on the agent's age:
Here $s_i$ denotes the career start date implied by persona $i$. Let $\bar{\pi}_{t,i} \equiv \alpha_{t,i}/(1-\beta_{t,i})$ denote persona $i$'s perceived long-run mean of inflation, where $b_{t,i} = (\alpha_{t,i},\beta_{t,i})'$ are the coefficients estimated recursively up to vintage $t$. The $h$-step-ahead forecast from the perceived AR(1) is
and the average annual forecast used in the experiment is $\pi^{\mathrm{exp},i}_t = H^{-1}\sum_{h=1}^{H}\pi^{\mathrm{exp},i}_{t+h|t}$.
Two implications follow from the learning mechanism in (ref)--(ref), the forecast formula (ref), and Proposition (ref). First, because the experienced persona's estimation sample includes the high-inflation 1970s while the young persona's does not, the experienced agent estimates both a higher perceived mean ($\bar{\pi}_{t,E} > \bar{\pi}_{t,Y}$) and a higher persistence ($\beta_{t,E} > \beta_{t,Y}$), so $\pi^{\mathrm{exp},E}_{t} > \pi^{\mathrm{exp},Y}_{t}$ and the persona gap $\Delta_t$ in (ref) is positive at early vintages. Second, after positive inflation surprises in 2021, the young persona updates more strongly because its gain $\gamma_{\tau,Y}$ is larger (shorter career implies larger $\gamma$), so $\Delta_t$ narrows across vintages.
Equation (ref) also connects the LLM experiment to the historically informed methods of Section (ref). The similarity approach, the kernel-weighted estimator, and the MN persona all produce forecasts of the same functional form, iteration of a perceived AR(1) (the MA component of the ARMA(1,1) enters only at $h=1$ and does not affect the multi-step recursion), and differ only in how they weight historical observations when estimating $(\alpha,\beta)$. The similarity approach uses hard truncation to a reference window; the kernel-weighted estimator uses smooth weights $w_\tau \propto K\!\left((z_\tau - z_t)/b\right)$ based on a state variable; the MN experienced persona uses age-dependent weights determined by the gain $\gamma_{\tau,i}$. When the forecast origin resembles the 1970s, all three weighting schemes concentrate mass on the supply-shock episodes, producing higher values of $\bar{\pi}_w$ and $\beta_w$ than the full-sample benchmark. The experience channel formalized by MalmendierNagel2016 is thus one particular weighting scheme, age-dependent, that achieves for the experienced persona what the historically informed methods achieve by design.
Panel (a) of Figure (ref) plots the average 2021 forecast by persona across the four information vintages. Consistent with the first implication of the MN framework, the experienced persona predicts higher inflation than the young persona at every vintage. The persona gap $\hat{\Delta}_t$ is 1.45 pp at the January vintage (experienced: 3.87%, young: 2.42%) and peaks at 1.70 pp at the April vintage (experienced: 5.33%, young: 3.63%), comparable in magnitude to the age-based expectation gap in the SCE documented in Figure (ref). As additional 2021 inflation data are incorporated, the gap narrows to 0.38 pp by October (experienced: 6.08%, young: 5.70%), consistent with the second implication (faster updating by the young persona due to a larger gain $\gamma_{\tau,Y}$). Panel (b) shows that the pattern is stable across LLMs: the gap is 1.56 pp (Claude) and 1.35 pp (ChatGPT) at the January vintage, and both converge toward zero by October. The neutral persona satisfies $\mu^{Y}_t \le \mu^{0}_t \le \mu^{E}_t$ at every vintage.\footnote{Because the 10 runs per persona--vintage cell (5 per LLM) are stochastic draws from the same foundation model and prompt, they do not constitute independent experimental units in the classical sense. The persona gap is therefore reported as a descriptive summary of the LLM outputs rather than as a formal test statistic.}
Figure (ref) decomposes the aggregate gap into monthly forecast profiles. At each vintage, the young persona predicts rapid mean-reversion toward 2--3 percent by year-end, while the experienced persona maintains forecasts in the 4--6 percent range. The neutral persona traces a path between the two. Under (ref) and (ref), this pattern reflects differences in the perceived mean $\bar{\pi}_{t,i}$ and persistence $\beta_{t,i}$ across personas.
The contrast is sharpest at the April and July vintages, when incoming data are high but the forecast horizon remains long. By October, the gap narrows but does not vanish: the experienced persona forecasts October--December at 5.9, 5.5, and 5.2 percent, while the young persona projects 4.6, 3.9, and 3.5 percent.
The experienced persona's January forecast (3.87 percent) is comparable to the SPF (3.89 percent) and FOMC projections (3.83 percent) from Table (ref). By July, it reaches 6.58 percent, approaching the intercept correction (6.83 percent) and similarity approach (6.19 percent) from Table (ref). The young persona's trajectory tracks the survey consensus instead: its January forecast (2.42 percent) is close to the CBO (2.65 percent) and Atlanta Fed (2.82 percent) projections. By Proposition (ref), this heterogeneity is not attributable to differential training leakage.
Beyond the numerical forecasts, the written reasoning accompanying each forecast provides descriptive evidence on how the persona shapes the economic narrative. Figure (ref) reports four text-based indicators constructed from the forecast justifications produced across all experimental runs (Appendix (ref) provides the methodology). These indicators are not derived from the formal framework of Section (ref), which concerns the conditional mean of the numerical forecasts; they characterize properties of the text that fall outside the scope of Proposition (ref).
Panel (a) shows that the experienced persona's net sentiment turns progressively hawkish (from $-2.8$ at the January vintage to $+4.7$ by July), while the young persona remains dovish throughout ($-10$ to $-4$). Panel (b) reveals that the young persona shows rising net uncertainty at the April and July vintages ($\sim$22 per thousand words), while the experienced persona maintains stable net uncertainty ($\sim$16--17) across vintages. Panel (c) shows that the experienced persona becomes increasingly backward-looking (net temporal orientation declining from $+19$ to $+14$), while the young persona maintains a forward orientation ($\sim$20--23). Panel (d) shows that the experienced persona invokes the 1970s two to three times more frequently than the young persona at every vintage.
Table (ref) illustrates how the two personas interpret identical macroeconomic signals, based on the LLMs' own methodological descriptions (Appendix (ref)).
The experienced persona maps supply disruptions and monetary accommodation into a more persistent inflation process, whereas the young persona maps the same signals into a more transitory path with faster mean reversion, consistent with the differences in perceived mean and persistence implied by equation (ref).
This paper shows that the 2021 inflation forecasting failure was primarily driven by sample composition rather than functional-form misspecification. Full-sample estimators dominated by the Great Moderation underweight supply-shock regimes and produce systematically low forecasts when the economy enters such a regime. Three simple adjustments, an intercept correction, a similarity-based re-estimation, and a kernel-weighted estimator, substantially close the forecast gap. The robust IC yields 5.00 percent, the similarity approach 6.19, and the single-origin IC 6.83 versus the realized 6.93, all large improvements over the ARMA benchmark of 3.52. The kernel-weighted approach (7.82) confirms the 1970s analogy through a data-driven criterion, and the sensitivity analysis shows that these results are robust to the choice of reference window. The same framework applied to eight additional price indices confirms that the mechanism extends across the entire U.S.\ price chain. Adaptive methods can perform worse than unadjusted benchmarks when a transitory shock precedes a persistent regime change: the TVP-AR(4) yields only 1.73 percent because the Kalman filter absorbs the deflationary 2020 readings and adjusts parameters in the wrong direction.
The evidence on household expectations reinforces this interpretation: respondents over 60 reported higher inflation expectations than younger cohorts throughout 2021, consistent with the experience-based learning framework of MalmendierNagel2016. The LLM experiment extends this evidence to professional forecasters. Under a common training leakage assumption, the experienced persona produces higher forecasts than the young persona at every vintage, and the gap narrows as data accumulate. The similarity approach, the kernel-weighted estimator, and the experienced persona all produce forecasts of the same functional form, iteration of a perceived AR(1), and differ only in how they weight historical observations.
The historically informed approaches succeed because the post-COVID episode did, in fact, resemble the 1970s. Their value lies in providing a disciplined device for encoding a specific economic hypothesis into the forecast. Developing a systematic framework for selecting the relevant historical precedent remains an important direction for future work.