EconBase
← Back to paper

A Benchmark Model for Fixed-Target Arctic Sea Ice Forecasting

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

24,133 characters · 8 sections · 11 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Benchmark Model for Fixed-Target Arctic Sea Ice Forecasting

spacing{1} Abstract: We propose a reduced-form benchmark predictive model (BPM) for fixed-target forecasting of Arctic sea ice extent, and we provide a case study of its real-time performance for target date September 2020. We visually detail the evolution of the statistically-optimal point, interval, and density forecasts as time passes, new information arrives, and the end of September approaches. Comparison to the BPM may prove useful for evaluating and selecting among various more sophisticated dynamical sea ice models, which are widely used to quantify the likely future evolution of Arctic conditions and their two-way interaction with economic activity. \thispagestyle{empty} Replication files (data, R code, matlab code, etc.): Available in the ancillary materials repository at \url{https://arxiv.org/abs/2101.10359}. {\bf Acknowledgments}: For helpful communications we thank, without implicating, Uma Bhatt, Philippe Goulet Coulombe, Walt Meier, Glenn Rudebusch, and Boyuan Zhang. {\bf Key words}: Climate forecasting, climate prediction, climate change, forecast evaluation { {\bf JEL codes}: Q54, C22, C51, C52, C53} {\bf Contact}: [email removed]

\setcounter{page}{1} \thispagestyle{empty}

Introduction

The Arctic is warming much faster than the rest of the planet, and it has emerged as a crucial focal point of climate change study. The path and pattern of Arctic sea ice diminution is of particular interest, and sea ice forecasting has received significant attention. From a real-time online perspective, there are two key forecast types: fixed-horizon (e.g., each month we might forecast one month ahead, month after month, ongoing) and fixed-target (e.g., each month we forecast a fixed future target month, month after month, ending when we arrive at the target month). In this paper we consider the fixed-target scenario, which has generated substantial interest in highlighting Arctic sea ice diminution both within years (as September 30 is approached, say) and across years (comparing the sequence of Septembers, say).\footnote{A third forecast type arises from an offline perspective -- the so-called extrapolation forecast, with a fixed origin and an expanding range of horizons, as with a forecast for every month from now until the end of the century.}

For example, each summer since 2008 the Sea Ice Prediction Network (SIPN) has sponsored the Sea Ice Outlook (SIO) competition for fixed-target prediction of September average daily Arctic sea ice extent.\footnote{See \url{https://www.arcus.org/sipn} for SIPN, and see \url{https://www.arcus.org/sipn/sea-ice-outlook} for SIO.} September extent forecasts are produced by many research groups mid-month in June, July, and August, and evaluated once September ends and the outcome is known. Insightful post-season SIO assessments have been produced annually (the most recent is SIOinterim2020), and similarly-insightful multi-year retrospective SIO assessments have been produced occasionally StroeveEtAl2014,HamiltonStroeve2016,Hamilton2020. Those assessments focus primarily on the forecasting skill of the SIO point-forecast ensembles.

In this paper we take an approach different from the SIO analyses, drilling very far down, focusing not on a point-forecast ensemble but rather on the point, interval, and density forecast paths for a single and very simple model (which we call the Benchmark Predictive Model, or BPM) in a single season (2020). The broad insights gained -- associated in particular with the evolution of forecast uncertainty from a simple yet sophisticated reduced-form sea ice forecasting model as time progresses and the target date is approached -- are of wide use. Indeed the BPM approach and results feature prominently in the “glide chart" climate model evaluation and comparison framework developed in DGGC2022, in which the BPM is used as the “naive" reference model in climate model skill scores.

We proceed as follows. In section (ref) we introduce the target-date forecasting framework and the BPM. In section (ref) we provide the 2020 case study. We consider forecasts made on the SIO dates, as well as a generalized set of forecasts made daily from June through September, and we pay particular attention to forecast uncertainty as the target date is approached. We conclude in section (ref).

A Benchmark Predictive Model for Arctic Sea Ice Extent

We consider target-date forecasting for September average daily sea ice extent, $SIE_{9}$, conditioning on the expanding historical sample as we move from June through the end of September. We forecast using a simple reduced-form model, which we call the “benchmark predictive model" (BPM), regressing September extent on four covariates:

equation[equation omitted — 123 chars of source]

where $SIE_p$ denotes average daily extent during period $p$ (hence, for example, $SIE_9$ denotes September extent), “$\rightarrow$" denotes “is regressed on", and the rest of the notation is obvious.

Approximately following SIO, we make $SIE_{9}$ forecasts on four days: 6/10, 7/10, 8/10 and 9/10.\footnote{We include a 9/10 forecast even though the 2020 SIO did not. The 9/10 forecast is of interest because September average extent is not known with certainty until the last day of September, well after 9/10. Indeed subsequent installments of the SIO will solicit 9/10 forecasts.} Immediately, the 6/10 regression used to produce the June forecast is $$ SIE_{9} ~{\rightarrow}~ c, ~ Time, ~ SIE_{5}, ~ SIE_{6/1\_ 6/10} ,~ SIE_{6/10}, $$ the 7/10 regression used to produce the July forecast is $$ SIE_{9} ~{\rightarrow}~ c, ~ Time, ~ SIE_{6}, ~ SIE_{7/1\_ 7/10} ,~ SIE_{7/10}, $$ the 8/10 regression used to produce the August forecast is $$ SIE_{9} ~{\rightarrow}~ c, ~ Time, ~ SIE_{7}, ~ SIE_{8/1\_ 8/10} ,~ SIE_{8/10}, $$ and the 9/10 regression used to produce the September forecast is $$ SIE_{9} ~{\rightarrow}~ c, ~ Time, ~ SIE_{8}, ~ SIE_{9/1\_ 9/10} ,~ SIE_{9/10}. $$

Perhaps surprisingly given their simplicity, the BPM forecasts are quite sophisticated in certain respects of relevance for forecasting Arctic sea ice:

enumerate• They capture low-frequency trend dynamics, via conditioning on $Time$. • They capture medium-frequency inertial (autoregressive) dynamics around trend, via conditioning on $SIE_{LastMonth}$. • They capture high-frequency dynamics by augmenting the conditioning on historical monthly information (via $SIE_{LastMonth}$) with potentially-invaluable recent daily information, via $SIE_{ThisMonthSoFar, t}$ and $SIE_{Today, t}$. • They readily enable probabilistic quantification of forecast uncertainty, which lets us move easily from point forecasts to interval and density forecasts. • They are based on a BPM estimated using direct rather than iterated projections.\footnote{One makes a multi-period “iterated" forecast with a one-period-ahead estimated model, iterated forward for the desired number of periods. In contrast, one makes a multi-period “direct" forecast with a horizon-specific multi-period-ahead estimated model.} Direct projections are theoretically superior under model misspecification (which is always the relevant case), because they directly minimize the relevant multi-step predictive loss.\footnote{See Ing2003, Theorem 4 and Corollary 3.} • They are easily made day-by-day, using model parameter estimates optimized day-by-day to the remaining predictive horizon, thanks to the BPM's simplicity. We exploit this fact below to make and examine not only the monthly SIO forecasts, but also 120 daily forecasts from June through September.

The BPM combination of trivial simplicity and subtle sophistication makes it an appropriate benchmark for skill score comparisons, as in DGGC2022. On the one hand, one would hope that a best-practice scientific model (e.g., a sophisticated structural climate model) should outperform the simple BPM, but on the other hand, it may not be easy!

Forecasting 2020 September Arctic Sea Ice Extent

Estimation

table[table omitted — 1,998 chars of source]

The left-hand-side variable of the BPM is September extent. September 2020 extent data were obviously unavailable on June 10, July 10, August 10, or September 10. Hence all estimation samples are 1979-2019, for a total of 41 annual observations.\footnote{Our daily extent measure is the National Snow and Ice Data Center (NSIDC) Sea Ice Index, Version 3 (\url{https://doi.org/10.7265/N5K072F8}), which uses the NASA team algorithm to convert microwave brightness readings into ice coverage Fettereretal2017. Until August 1986, data are reported only every other day, and we fill missing days with the average of the two adjacent days.}

Estimation results appear in the top and middle panels of Table (ref). Several points are worth noting. First, the negative linear trend becomes progressively less important as September approaches, whereas the positive autoregressive effect $SIE_{LastMonth}$ becomes progressively more important as September approaches. This is completely natural. The conditioning on May extent in the June 10 forecast, for example, is of little value for forecasting September extent, so the trend plays an important role. In contrast, moving to the end of the summer, the conditioning on August extent in the September 10 forecast is of great value for forecasting September extent, so the trend plays almost no role.

Second, $SIE_{ThisMonthSoFar}$ has a negative effect and $SIE_{Today}$ has a positive effect. Hence the estimates, and the forecasts that we construct from them, are influenced not just by $SIE_{Today}$, but also by $SIE_{Today}$ relative to $SIE_{ThisMonthSoFar}$.

Finally, adjusted R-squared ($\bar{R}^2$) naturally increases toward 1.0 as September approaches, because the value of the conditioning information ($SIE_{LastMonth}$, $SIE_{ThisMonthSoFar}$, $SIE_{Today}$) increases as September approaches. In parallel, the standard error of the regression ($\hat{\sigma}$) naturally decreases toward 0 as September approaches, again because the value of the conditioning information increases as September approaches.

Forecasting

To use an estimated forecasting model to make a point forecast, we simply insert the relevant right-hand-side variables, all of which are known at the time the forecast is made. For example, to form the July 10 forecast we evaluate the fitted July model at $Time{=}42$, $SIE_{LastMonth}{=}SIE_{6/2020}$, $SIE_{ThisMonthSoFar}{=} SIE_{7/1/2020\_ 7/10/2020}$, $SIE_{Today}{=} SIE_{7/10/2020}$.\footnote{There is typically a 1-day data availability lag, so we would actually insert $SIE_{LastMonth}{=}$ $SIE_{6/2020}$, $SIE_{ThisMonthSoFar}{=}$ $SIE_{7/1/2020\_ 7/9/2020}$, $SIE_{Today}{=}$ $SIE_{7/9/2020}$.} This point forecast is an estimate of the mean of $SIE_{7/2020}$ conditional on $Time{=}42$, \\ $SIE_{LastMonth}{=} SIE_{6/2020}$, $SIE_{ThisMonthSoFar}{=} SIE_{7/1/2020\_ 7/10/2020}$, and $SIE_{Today}{=} SIE_{7/10/2020}$. Hence we denote the point forecast by $\hat{\mu}$ in Table (ref).

Now consider interval forecasts (predictive intervals). Let us stay with the same July example. To make an interval forecast we need an estimate of the standard deviation of $SIE_{7/2020}$ conditional on the same covariates: $Time{=}42$, $SIE_{LastMonth}{=}SIE_{6/2020}$, \\ $SIE_{ThisMonthSoFar}{=}$ $SIE_{7/1/2020\_ 7/10/2020}$, and $SIE_{Today}{=}SIE_{7/10/2020}$. The standard error of the regression, denoted $\hat{\sigma}$ in Table (ref), is precisely such an estimate.\footnote{Note that $\hat{\sigma}$ measures true forecast uncertainty, which is a very different concept from the cross-section dispersion in the ensemble of forecasts, $\hat{d}$. We want $\hat{\sigma}$, and in general $\hat{\sigma} {\ne} \hat{d}$.} An interval forecast (ignoring parameter estimation uncertainty) is then $\hat{\mu} {\pm} 2 \hat{\sigma}$. If the regression disturbances are approximately Gaussian, then the $\hat{\mu} {\pm} 2 \hat{\sigma}$ interval is an approximate 95% predictive interval.\footnote{One could use simulation-based bootstrap procedures to accommodate parameter estimation uncertainty and/or non-Gaussian disturbances in forming interval forecasts, but we do not pursue that here.}

Finally, again ignoring parameter estimation uncertainty, consider density forecasts (predictive densities). If the regression disturbances are approximately Gaussian, then the full predictive density is approximately $N(\hat{\mu}, \hat{\sigma}^2)$.\footnote{As with the interval forecast case, bootstrap procedures could be used to accommodate parameter estimation uncertainty and/or non-Gaussian disturbances.}

Four Month-by-Month Predictive Densities

figure[figure omitted — 548 chars of source]

In Figure (ref) we show the four monthly predictive densities (June, July, August, and September) corresponding to our generalized SIO exercise that includes a September 10 forecast. The density locations (their means, the $\hat{\mu}$'s in Table (ref)) naturally evolve throughout the summer as the conditioning information evolves, but they eventually get closer to the end-of-September value. The density mean is above the realization in June, below in July, above again in August, and then almost spot-on in September.

Not unrelated, and importantly, the forecast uncertainty as captured by the predictive density dispersion ($\hat{\sigma}$ in Table (ref)) decreases monotonically moving through the summer: from Table (ref) it is $0.46$, $0.40$, $0.27$, $0.10$ for June, July, August, and September, respectively.

120 Day-by-Day Predictive Densities

figure[figure omitted — 851 chars of source]

There is nothing sacrosanct about the set of once-per-month SIO forecast dates examined thus far. Given the simplicity of our forecasting model and its estimation, we can examine many other dates. We simply generalize the BPM from

equation[equation omitted — 106 chars of source]

to

equation[equation omitted — 118 chars of source]

and the framework is otherwise unchanged.

In Figure (ref), we show predictive densities for the 120 days leading to the end of September, produced using 120 different estimated models. In the top panel we plot the entire sequence [-120, 0], and in the bottom panel we plot only [-120, -20] to enhance visualization detail. Throughout, the horizontal axis represents the number of days until the end of September, and the green line is the evolving point forecast (the mean of the predictive density). One can readily see the densities wandering left and right as new information arrives, but nevertheless eventually rising sharply and clustering tightly around the realized value as the end of September nears.

figure[figure omitted — 762 chars of source]

In Figure (ref) we reduce the predictive densities to predictive intervals. As the target date approaches, the interval forecast midpoint (the point forecast, $\hat{\mu}$) evolves as the conditioning information evolves, converging to the eventually-realized September value. Simultaneously the interval forecast width ($4 \hat{\sigma}$) also evolves as information accumulates, converging to zero by the target date.\footnote{Of course the densities of Figure (ref) and the intervals of Figure (ref) are isomorphic in a Gaussian environment -- if one knows the density, then one knows the interval, and conversely, so that nothing new is learned by reduction of densities to intervals. Nevertheless the sequence of intervals may be visually revealing in certain ways that the sequence of densities is not, more clearly emphasizing both the point forecast trajectory and its associated uncertainty, and hence serving as a complement rather than a substitute for the sequence of densities. Moreover, and importantly, the environment may not be Gaussian, in which case the $\pm 2 \hat{\sigma}$ intervals are still a useful and transparent quantification of forecast uncertainty, even if they lose their interpretation at 95% confidence intervals.}

Concluding Remarks

We earlier asserted that our benchmark predictive model (BPM) is quite sophisticated in certain respects. As it turned out, its performance in the 2020 Sea Ice Outlook competition was in the middle of the pack, a thoroughly respectable performance for a simple BPM. And the point, of course, is not that the BPM should dominate its competitors, but rather that it should serve as a simple benchmark against which allegedly more sophisticated competitors can be compared.

Following that path, one may use the BPM as the reference model in “skill score glide charts" for climate model evaluation and comparison, tracking relative forecasting performance of competitors vs BPM as time evolves and the target date is approached. Such competitor vs BPM skill score glide charts are proposed and explored in work in progress DGGC2022.

Skill score competitors may include more sophisticated reduced-form models, including, for example, models that:

enumerate• incorporate nonlinearity, whether parametrically (e.g., DRice), or nonparametrically as in a variety of statistical machine learning methods (e.g., hastie2009); • incorporate and forecast the entire daily sea ice extent history (note that we do not model the entire daily history -- we model the monthly history augmented with certain aspects of the very recent daily history); • drop the normality assumption for calculating predictive densities, instead using simulation-based bootstrap procedures to approximate them nonparametrically by sampling with replacement from regression residuals Efron79; • broaden the information set from univariate to multivariate, conditioning as well on natural covariates like sea ice thickness, surface air temperature, and radiative forcings, as for example in VARCTIC.

Alternatively, and of great interest, competitors may include large-scale structural dynamical climate models. That is, given a particular dynamical climate model, one could compare its “model-based theoretical Figure (ref)" to the “data-based BPM Figure (ref)" via skill score glide charts.

In any event, comparison to the BPM may prove useful for evaluating and selecting among various more sophisticated sea ice models -- whether reduced-form or structural -- which are widely used to quantify the likely future evolution of Arctic conditions and their two-way interaction with economic activity.

\addcontentsline{toc}{section}{References}