Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
24,133 characters · 8 sections · 11 citation commands
A Benchmark Model for Fixed-Target Arctic Sea Ice Forecasting
\setcounter{page}{1} \thispagestyle{empty}
The Arctic is warming much faster than the rest of the planet, and it has emerged as a crucial focal point of climate change study. The path and pattern of Arctic sea ice diminution is of particular interest, and sea ice forecasting has received significant attention. From a real-time online perspective, there are two key forecast types: fixed-horizon (e.g., each month we might forecast one month ahead, month after month, ongoing) and fixed-target (e.g., each month we forecast a fixed future target month, month after month, ending when we arrive at the target month). In this paper we consider the fixed-target scenario, which has generated substantial interest in highlighting Arctic sea ice diminution both within years (as September 30 is approached, say) and across years (comparing the sequence of Septembers, say).\footnote{A third forecast type arises from an offline perspective -- the so-called extrapolation forecast, with a fixed origin and an expanding range of horizons, as with a forecast for every month from now until the end of the century.}
For example, each summer since 2008 the Sea Ice Prediction Network (SIPN) has sponsored the Sea Ice Outlook (SIO) competition for fixed-target prediction of September average daily Arctic sea ice extent.\footnote{See \url{https://www.arcus.org/sipn} for SIPN, and see \url{https://www.arcus.org/sipn/sea-ice-outlook} for SIO.} September extent forecasts are produced by many research groups mid-month in June, July, and August, and evaluated once September ends and the outcome is known. Insightful post-season SIO assessments have been produced annually (the most recent is SIOinterim2020), and similarly-insightful multi-year retrospective SIO assessments have been produced occasionally StroeveEtAl2014,HamiltonStroeve2016,Hamilton2020. Those assessments focus primarily on the forecasting skill of the SIO point-forecast ensembles.
In this paper we take an approach different from the SIO analyses, drilling very far down, focusing not on a point-forecast ensemble but rather on the point, interval, and density forecast paths for a single and very simple model (which we call the Benchmark Predictive Model, or BPM) in a single season (2020). The broad insights gained -- associated in particular with the evolution of forecast uncertainty from a simple yet sophisticated reduced-form sea ice forecasting model as time progresses and the target date is approached -- are of wide use. Indeed the BPM approach and results feature prominently in the “glide chart" climate model evaluation and comparison framework developed in DGGC2022, in which the BPM is used as the “naive" reference model in climate model skill scores.
We proceed as follows. In section (ref) we introduce the target-date forecasting framework and the BPM. In section (ref) we provide the 2020 case study. We consider forecasts made on the SIO dates, as well as a generalized set of forecasts made daily from June through September, and we pay particular attention to forecast uncertainty as the target date is approached. We conclude in section (ref).
We consider target-date forecasting for September average daily sea ice extent, $SIE_{9}$, conditioning on the expanding historical sample as we move from June through the end of September. We forecast using a simple reduced-form model, which we call the “benchmark predictive model" (BPM), regressing September extent on four covariates:
where $SIE_p$ denotes average daily extent during period $p$ (hence, for example, $SIE_9$ denotes September extent), “$\rightarrow$" denotes “is regressed on", and the rest of the notation is obvious.
Approximately following SIO, we make $SIE_{9}$ forecasts on four days: 6/10, 7/10, 8/10 and 9/10.\footnote{We include a 9/10 forecast even though the 2020 SIO did not. The 9/10 forecast is of interest because September average extent is not known with certainty until the last day of September, well after 9/10. Indeed subsequent installments of the SIO will solicit 9/10 forecasts.} Immediately, the 6/10 regression used to produce the June forecast is $$ SIE_{9} ~{\rightarrow}~ c, ~ Time, ~ SIE_{5}, ~ SIE_{6/1\_ 6/10} ,~ SIE_{6/10}, $$ the 7/10 regression used to produce the July forecast is $$ SIE_{9} ~{\rightarrow}~ c, ~ Time, ~ SIE_{6}, ~ SIE_{7/1\_ 7/10} ,~ SIE_{7/10}, $$ the 8/10 regression used to produce the August forecast is $$ SIE_{9} ~{\rightarrow}~ c, ~ Time, ~ SIE_{7}, ~ SIE_{8/1\_ 8/10} ,~ SIE_{8/10}, $$ and the 9/10 regression used to produce the September forecast is $$ SIE_{9} ~{\rightarrow}~ c, ~ Time, ~ SIE_{8}, ~ SIE_{9/1\_ 9/10} ,~ SIE_{9/10}. $$
Perhaps surprisingly given their simplicity, the BPM forecasts are quite sophisticated in certain respects of relevance for forecasting Arctic sea ice:
The BPM combination of trivial simplicity and subtle sophistication makes it an appropriate benchmark for skill score comparisons, as in DGGC2022. On the one hand, one would hope that a best-practice scientific model (e.g., a sophisticated structural climate model) should outperform the simple BPM, but on the other hand, it may not be easy!
The left-hand-side variable of the BPM is September extent. September 2020 extent data were obviously unavailable on June 10, July 10, August 10, or September 10. Hence all estimation samples are 1979-2019, for a total of 41 annual observations.\footnote{Our daily extent measure is the National Snow and Ice Data Center (NSIDC) Sea Ice Index, Version 3 (\url{https://doi.org/10.7265/N5K072F8}), which uses the NASA team algorithm to convert microwave brightness readings into ice coverage Fettereretal2017. Until August 1986, data are reported only every other day, and we fill missing days with the average of the two adjacent days.}
Estimation results appear in the top and middle panels of Table (ref). Several points are worth noting. First, the negative linear trend becomes progressively less important as September approaches, whereas the positive autoregressive effect $SIE_{LastMonth}$ becomes progressively more important as September approaches. This is completely natural. The conditioning on May extent in the June 10 forecast, for example, is of little value for forecasting September extent, so the trend plays an important role. In contrast, moving to the end of the summer, the conditioning on August extent in the September 10 forecast is of great value for forecasting September extent, so the trend plays almost no role.
Second, $SIE_{ThisMonthSoFar}$ has a negative effect and $SIE_{Today}$ has a positive effect. Hence the estimates, and the forecasts that we construct from them, are influenced not just by $SIE_{Today}$, but also by $SIE_{Today}$ relative to $SIE_{ThisMonthSoFar}$.
Finally, adjusted R-squared ($\bar{R}^2$) naturally increases toward 1.0 as September approaches, because the value of the conditioning information ($SIE_{LastMonth}$, $SIE_{ThisMonthSoFar}$, $SIE_{Today}$) increases as September approaches. In parallel, the standard error of the regression ($\hat{\sigma}$) naturally decreases toward 0 as September approaches, again because the value of the conditioning information increases as September approaches.
To use an estimated forecasting model to make a point forecast, we simply insert the relevant right-hand-side variables, all of which are known at the time the forecast is made. For example, to form the July 10 forecast we evaluate the fitted July model at $Time{=}42$, $SIE_{LastMonth}{=}SIE_{6/2020}$, $SIE_{ThisMonthSoFar}{=} SIE_{7/1/2020\_ 7/10/2020}$, $SIE_{Today}{=} SIE_{7/10/2020}$.\footnote{There is typically a 1-day data availability lag, so we would actually insert $SIE_{LastMonth}{=}$ $SIE_{6/2020}$, $SIE_{ThisMonthSoFar}{=}$ $SIE_{7/1/2020\_ 7/9/2020}$, $SIE_{Today}{=}$ $SIE_{7/9/2020}$.} This point forecast is an estimate of the mean of $SIE_{7/2020}$ conditional on $Time{=}42$, \\ $SIE_{LastMonth}{=} SIE_{6/2020}$, $SIE_{ThisMonthSoFar}{=} SIE_{7/1/2020\_ 7/10/2020}$, and $SIE_{Today}{=} SIE_{7/10/2020}$. Hence we denote the point forecast by $\hat{\mu}$ in Table (ref).
Now consider interval forecasts (predictive intervals). Let us stay with the same July example. To make an interval forecast we need an estimate of the standard deviation of $SIE_{7/2020}$ conditional on the same covariates: $Time{=}42$, $SIE_{LastMonth}{=}SIE_{6/2020}$, \\ $SIE_{ThisMonthSoFar}{=}$ $SIE_{7/1/2020\_ 7/10/2020}$, and $SIE_{Today}{=}SIE_{7/10/2020}$. The standard error of the regression, denoted $\hat{\sigma}$ in Table (ref), is precisely such an estimate.\footnote{Note that $\hat{\sigma}$ measures true forecast uncertainty, which is a very different concept from the cross-section dispersion in the ensemble of forecasts, $\hat{d}$. We want $\hat{\sigma}$, and in general $\hat{\sigma} {\ne} \hat{d}$.} An interval forecast (ignoring parameter estimation uncertainty) is then $\hat{\mu} {\pm} 2 \hat{\sigma}$. If the regression disturbances are approximately Gaussian, then the $\hat{\mu} {\pm} 2 \hat{\sigma}$ interval is an approximate 95% predictive interval.\footnote{One could use simulation-based bootstrap procedures to accommodate parameter estimation uncertainty and/or non-Gaussian disturbances in forming interval forecasts, but we do not pursue that here.}
Finally, again ignoring parameter estimation uncertainty, consider density forecasts (predictive densities). If the regression disturbances are approximately Gaussian, then the full predictive density is approximately $N(\hat{\mu}, \hat{\sigma}^2)$.\footnote{As with the interval forecast case, bootstrap procedures could be used to accommodate parameter estimation uncertainty and/or non-Gaussian disturbances.}
In Figure (ref) we show the four monthly predictive densities (June, July, August, and September) corresponding to our generalized SIO exercise that includes a September 10 forecast. The density locations (their means, the $\hat{\mu}$'s in Table (ref)) naturally evolve throughout the summer as the conditioning information evolves, but they eventually get closer to the end-of-September value. The density mean is above the realization in June, below in July, above again in August, and then almost spot-on in September.
Not unrelated, and importantly, the forecast uncertainty as captured by the predictive density dispersion ($\hat{\sigma}$ in Table (ref)) decreases monotonically moving through the summer: from Table (ref) it is $0.46$, $0.40$, $0.27$, $0.10$ for June, July, August, and September, respectively.
There is nothing sacrosanct about the set of once-per-month SIO forecast dates examined thus far. Given the simplicity of our forecasting model and its estimation, we can examine many other dates. We simply generalize the BPM from
to
and the framework is otherwise unchanged.
In Figure (ref), we show predictive densities for the 120 days leading to the end of September, produced using 120 different estimated models. In the top panel we plot the entire sequence [-120, 0], and in the bottom panel we plot only [-120, -20] to enhance visualization detail. Throughout, the horizontal axis represents the number of days until the end of September, and the green line is the evolving point forecast (the mean of the predictive density). One can readily see the densities wandering left and right as new information arrives, but nevertheless eventually rising sharply and clustering tightly around the realized value as the end of September nears.
In Figure (ref) we reduce the predictive densities to predictive intervals. As the target date approaches, the interval forecast midpoint (the point forecast, $\hat{\mu}$) evolves as the conditioning information evolves, converging to the eventually-realized September value. Simultaneously the interval forecast width ($4 \hat{\sigma}$) also evolves as information accumulates, converging to zero by the target date.\footnote{Of course the densities of Figure (ref) and the intervals of Figure (ref) are isomorphic in a Gaussian environment -- if one knows the density, then one knows the interval, and conversely, so that nothing new is learned by reduction of densities to intervals. Nevertheless the sequence of intervals may be visually revealing in certain ways that the sequence of densities is not, more clearly emphasizing both the point forecast trajectory and its associated uncertainty, and hence serving as a complement rather than a substitute for the sequence of densities. Moreover, and importantly, the environment may not be Gaussian, in which case the $\pm 2 \hat{\sigma}$ intervals are still a useful and transparent quantification of forecast uncertainty, even if they lose their interpretation at 95% confidence intervals.}
We earlier asserted that our benchmark predictive model (BPM) is quite sophisticated in certain respects. As it turned out, its performance in the 2020 Sea Ice Outlook competition was in the middle of the pack, a thoroughly respectable performance for a simple BPM. And the point, of course, is not that the BPM should dominate its competitors, but rather that it should serve as a simple benchmark against which allegedly more sophisticated competitors can be compared.
Following that path, one may use the BPM as the reference model in “skill score glide charts" for climate model evaluation and comparison, tracking relative forecasting performance of competitors vs BPM as time evolves and the target date is approached. Such competitor vs BPM skill score glide charts are proposed and explored in work in progress DGGC2022.
Skill score competitors may include more sophisticated reduced-form models, including, for example, models that:
Alternatively, and of great interest, competitors may include large-scale structural dynamical climate models. That is, given a particular dynamical climate model, one could compare its “model-based theoretical Figure (ref)" to the “data-based BPM Figure (ref)" via skill score glide charts.
In any event, comparison to the BPM may prove useful for evaluating and selecting among various more sophisticated sea ice models -- whether reduced-form or structural -- which are widely used to quantify the likely future evolution of Arctic conditions and their two-way interaction with economic activity.
\addcontentsline{toc}{section}{References}