Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
78,343 characters · 12 sections · 43 citation commands
Time preference effects in forecasting
\baselineskip18pt \setcounter{totalnumber}{50} \setcounter{topnumber}{50} \setcounter{bottomnumber}{50} \abovedisplayskip1.5ex plus1ex minus1ex \belowdisplayskip1.5ex plus1ex minus1ex \abovedisplayshortskip1.5ex plus1ex minus1ex \belowdisplayshortskip1.5ex plus1ex minus1ex
\doublespacing
When collecting expectations and forecasts regarding disruptive future events, such as wars, technological breakthroughs, and volcanic eruptions, we would like to reward forecasters for their accurate assessment of the matter.\footnote{ The idea that we can create incentives for careful and honest assessments of probability judgments dates back at least to brier_verification_1950. He found that there exist incentive-compatible payment schemes that provide the most money to forecasters who honestly share their expectation regarding future outcomes.} However, it is difficult to incentivize forecasters for accurately sharing their expectation when we do not know when, or if, the event will occur. Specifically, it is common practice to evaluate forecasts after outcomes are known, which may distort incentives to report honestly if human forecasters discount potential rewards in the distant future relative to more immediate ones.
To illustrate, let us consider the question: When will there be a volcanic eruption, which ejects over 100 cubic kilometers of erupted material? An expert volcanic forecaster will have an expectation regarding this event, assigning each future point in time a probability of such an eruption occurring. Assume we wanted to elicit this expectation, e.g., using a survey or an interview. We would want to make sure that the forecaster reports their honest expectation. Ideally, the forecaster would even refine their expectation by engaging in research relevant to the prediction task. How should such a forecast be evaluated and rewarded? Evaluation refers to determining whether the forecast has been “good” in the most general sense. This requires either that the eruption has happened, upon which we can evaluate whether the forecast was indeed indicative of the eruption, or that a certain time has passed, say 20 years without any such eruption. The problem is that most forecasters will likely care less about evaluations in 20 years time, and more about those in the short-term future. Since any evaluation of their forecast in the short-term future can only be extraordinarily positive in the case of such an eruption (where the prediction will get attention and evaluation), the forecaster has an incentive to report a higher probability of an early eruption than they truly expect there to be.
Against this backdrop, we formally investigate the incentive to misreport one's expectation when forecasters discount the future and the predicted event can occur at arbitrary points in time, such that the timing of the reward is uncertain. Our first main contribution is to show that every incentive-compatible payment scheme that rewards accurate forecasts becomes incentive-incompatible with honest reporting when forecasters discount the future and---as is common and often unavoidable---are evaluated upon occurrence. Our second main contribution is to empirically investigate whether real human forecasters misreport their own expectations using data from a real-world forecasting tournament. As predicted by our theory, we find strong evidence that forecasters respond to incentives to misreport their expectation. In discussing four potential solutions to this problem, we find that it is inherently challenging to provide incentives for honest reports when the predicted event can occur at any future time.
The hypothesis that we investigate in this article may strike readers as an unintuitive one. Why would forecasters be misreporting their own expectation? There is a literature on probabilistic judgment in experimental economics, statistical decision theory and related fields that provides ample evidence of forecasters misreporting their expectation when given incentives to do so. For example, armantier_eliciting_2013 find in a series of experiments that participants are susceptible to opportunities to hedge or avert risk when making probabilistic judgments. Another reason to actively change the information reported via forecasts arises in competition with other forecasters ottaviani_strategy_2006. Forecasters may try to actively stand out from the crowd witkowski_incentive-compatible_2023 or not be the one with the worst prediction, i.e., to make predictions more similar to the crowd in spite of contradictory private information.
Despite this large amount of literature charness_experimental_2021, we are not aware of a single work that studies the issue of time preferences in forecasts. The only explicit mention of time preferences related to forecasts that we know of appears in chambers_dynamic_2021, who state that “time preferences [...] complicate the task of elicitation”. The main reason for this seems to be that the subtle---yet important---distinction between scheduled and unscheduled future events has not been made in the past. Historically speaking, the literature on forecasting and belief elicitation has mostly focused on events which are certain, or almost certain, to occur at one point in time or form a time series zellner_survey_2021. Henceforth, we refer to events that are unscheduled and can happen at multiple points in time as time-varying events, and others as time-fixed events, where resolution occurs at some scheduled (fixed) time in the future.
Time-fixed events are prevalent in macroeconomics (GDP growth next quarter, inflation one year ahead), meteorology (rainfall tomorrow, temperature next week) or finance (Value-at-Risk 10-days-ahead, interest rates in one month). All of these events have in common that they are scheduled, that is, their verification is fixed in time. However, a great number of events for which we would like to gather forecasts are not scheduled and can happen at arbitrary points in time. Examples include floods, volcanic eruptions, deaths, insolvency, technological breakthroughs, wars and whether a team will be eliminated in the playoffs. As vere-jones_forecasting_1995 puts it: “[...] forecasting earthquakes differs from most routine forecasting problems in that it deals not with a discrete or continuous time series, but with sudden random events. But, although problems of this kind may be relatively uncommon in forecasting practice, they are certainly not unknown. In essence, they fall into the same category as forecasting lifetimes, for which there exist substantial literatures in the engineering (reliability), actuarial, and medical contexts.”
Indeed, Survival Analysis---also known as reliability analysis, duration analysis or event history analysis---is an entire subfield which is predominantly concerned with the timing and occurrence probability of one-time events allison_event_2014. Whilst there is overlap with the contents of this article, survival analysis deals less with forecasting and more with inference. Usually, survival analysis is concerned with estimating the survival of a group of subjects (hence the name), e.g., estimating the life expectancy of a population. Furthermore, these estimates are often conditional on observed covariates, such as age and gender. Survival analysis may also include causal inference, such as inferring the effect of medication on life expectancy.
Although there is some relation of our work to survival analysis, the most closely related paper to ours in spirit is Lea17. Similar to us, they also take a decision-theoretic perspective on the task of forecasting---understanding the forecaster as a self-interested and rational reporter of private information. Specifically, Lea17 describe the Forecaster's Dilemma, where forecasts get outsized attention if, and only if, extreme events happen, such as after the 2008 financial crisis and the COVID-19 pandemic. This is equivalent to weighting the evaluation of forecasts conditional on specific outcomes, such as the occurrence of a crisis. GR11 show that any such weighting will inevitably distort optimal forecasts. In the context of Lea17, forecasters that always predict calamity will be perceived as more accurate than skilled forecasters. At their core, the Forecaster's Dilemma and the problem we investigate in this paper share the same underlying structure: in both cases, the evaluation of a forecast is not weighted equally across all possible real-world outcomes. Instead, certain outcomes cause the evaluation to matter more than others. In the Forecaster's Dilemma, this imbalance arises externally---extreme outcomes attract disproportionate public attention to the forecast. In the setting we describe, the imbalance arises internally---forecasters who discount the future place greater weight on more immediate evaluations, and thus, more immediate outcomes.
The remainder of this article proceeds as follows. Section (ref) investigates the incentive-compatibility of strictly proper scoring rules for binary probabilistic forecasts. Section (ref) is concerned with the incentive-compatibility of the Continuous Ranked Probability Score, an error measure for distributional forecasts. Section (ref) features the empirical analysis of forecasters in a real-world forecasting tournament. Section (ref) discusses the broader relevance of the results and outlines future research directions. Finally, Section (ref) concludes.
We investigate the incentive to misreport by employing an illustrative example and leave a more formal investigation to Section (ref). Let two players, Alice and Bob, compete against each other in a best-of-three series. The first player to achieve two victories wins the series. Alice has won the first matchup, and thus the standing is 1-0. Potential outcomes are displayed in Figure (ref). We are interested in the probability of Alice winning the entire series. We therefore turn to a forecaster and ask:
We denote the event of interest (that Alice wins the series) by $X \in \{0,1\}$. Here, $X=1$ corresponds to the case where Alice has won the series. Let $T \in \{2,3\}$ indicate when Alice has won the series; $T=2$ is the case where Alice has won in matchup 2, and $T=3$ is the case where Alice has won in matchup 3. The event of Alice winning the series is binary and the forecast is a probability $G\in[0,1]$. Throughout the paper, we use the terms report, forecast and prediction interchangeably.
In order to incentivize the forecaster to report honestly and accurately, we evaluate the forecast with a strictly proper scoring rule\footnote{Strictly proper scoring rules are a widely employed tool to evaluate forecasts. Such scoring rules are functions with the property that the expected score is minimized only by setting the reported variable $G$ equal to the expected outcome $\mathbb{E}[X]$ GR07. Therefore, if a forecaster issues a forecast $G$ and gets a fixed reward from which we subtract the score $S(G,X)$ when $X$ materializes, then---on average---she can do no better than to issue the true forecast $G=\mathbb{E}[X]$ to maximize her expected reward (or minimize the expected score). In this sense, the score incentivizes honest forecasts or, more lyrically, serves as a “truth serum”.} $S(G,X)$ and reward the forecaster proportionally to the negated score as soon as the outcome $X$ is known. The forecaster is a risk-neutral score minimizer and discounts the future.\footnote{That is, we ignore risk aversion throughout this analysis. We conjecture that the problem remains even without risk-neutrality. However, risk aversion is then an additional reason why the forecaster may report a dishonest forecast hossain_binarized_2013.} Without loss of generality, we assume that the forecaster discounts between matchups 2 and 3 at rate $r>0$; that is, a payoff realized after matchup 2 is valued $(1+r)$-times as much as the same payoff realized after matchup 3.
The forecaster now has an incentive to misreport their own true expectation $\mathbb{E}[X]$. This is because the forecaster gives extra weight ($1+r$) to the case where Alice wins the series in matchup 2, inflating the reported probability that Alice will win the series. Next, we show that the error-minimizing forecast $G$ is strictly greater than the forecasters true expectation $\mathbb{E}[X]$. We can integrate the discount rate $r$ directly into the forecaster's valuation of her expected score and call this a discounted scoring rule $S_d$:
We denote the true expectation of Alice winning the series as $\mathbb{E}[X] = \mathbb{P}\{T=2\} + \mathbb{P}\{T=3\}$, the sum of probabilities of Alice winning either in match 2 or 3. If the forecaster were to minimize the expected undiscounted score, then they could do no better than to issue $G=\mathbb{E}[X]$ (because this maximizes the expected payoff, $-\mathbb{E}[S(G,X)]$). However, because of the forecasters present bias, they minimize $\mathbb{E}[S_d(G,X)]$, issuing a different forecast.\footnote{Throughout the paper, we simply take “present bias” to mean that earlier payoffs are preferred to later payoffs of the same amount. However, there is a large literature in economics where the term is understood more narrowly OR15,Cha21. Specifically, in behavioral economics, present bias (or also: the immediacy effect) refers to the tendency to favor a smaller reward available immediately over a larger reward received later, while reversing this preference when both rewards are shifted equally into the future.} Specifically, by straightforward computations, the forecaster's expected score is
The added factor in (ref) is essentially an additional constant weight. This suffices to distort incentives for honest reporting as multiplying any strictly proper scoring rule with a constant factor conditional on outcomes, i.e., putting more weight on one outcome, leads to a predictably improper scoring rule. For the proof, we refer to Lemma 4 in lindley_scoring_1982, which---in the context of forecasting---is accessibly explained by parmigiani_decision_2009. The optimal forecast is no longer the true expectation.
We illustrate this result by demonstrating the improperness of the discounted quadratic scoring rule ($\operatorname{QSR}_d$), where the quadratic scoring rule (or also: Brier score)
is one specific strictly proper scoring rule. The expected score then is
From Lemma 4 of lindley_scoring_1982 we know that the relationship between the expectation $\mathbb{E}[X]$ and the reported forecast $G$ (that minimizes $\mathbb{E}[S_d(G,X)]$) is given by
Solving for $G$ and re-arranging gives that
It is obvious that the error-minimizing forecast $G$ is strictly greater than $\mathbb{E}[X]$ for $r>0$. If either time preferences $r$ or the possibility for the event to occur early $\mathbb{P}\{T=2\}$ is set to zero, then (ref) yields that the error-minimizing report is $G=\mathbb{E}[X]$.
Figure (ref) plots the optimal forecast on the horizontal axis and the true expectation on the vertical axis, setting $\mathbb{P}\{T=2\}= 0.2$. The blue lines represent the score-minimizing forecast for different $r>0$. The red line corresponds to the 45 degree diagonal, where $G=\mathbb{E}[X]$. We see that the reported probability increases with a rising discount factor.
Clearly, there is an incentive to report a probability that is too high. This is not a general result. For example, had the forecaster been asked to forecast if Bob will win the series, i.e., had been asked to predict the complement, they would have been incentivized to report a probability that is lower than truly expected, as Bob can only lose early. When the time-varying event can both occur early and be ruled out early, the incentive to misreport exists, but it is not obvious in which direction forecasts will be influenced. For example, imagine we collected forecasts on whether a certain person will succeed the CEO of a company at the end of her/his term. This can happen at any time, but it can also be ruled out by a premature death of the candidate, a lifelong jail sentence, the collapse of the company or the appointment of another CEO.
While Section (ref) investigates probabilistic forecasts for the occurrence of future events, we now investigate the incentive to misreport forecasts regarding the timing of future events. For technical ease, we treat time as continuous now. The event of interest, $X$, remains binary. The forecasting question thus is:
Denote by $T\in[0,\tau]$ the random point in time when $X$ occurs. Here, $T=\tau$ corresponds to the case where $X$ occurs at time $\tau$, or at some later time or, possibly even, never. Denote the true cumulative distribution function (cdf) of $T$ by $F(t)=\mathbb{P}\{T\leq t\}$. Then, $F(\tau-)=\lim_{t\uparrow \tau}F(t)$ denotes the probability that $X$ occurs during $[0,\tau)$, and $1-F(\tau-)$ denotes the probability that $X$ occurs in $\tau$, afterward or never. We denote a generic forecast for the cdf of $T$ by $G(\cdot)$. A popular score to rank different distributional forecasts for $T$ is the continuous ranked probability score (CRPS), defined as \[ \operatorname{CRPS}(G,T) = \int_{-\infty}^{\infty}\big[G(t) - \mathds{1}_{\{t\geq T\}}\big]^2\,\mathrm{d} t, \] where $G$ is a generic cdf forecast and $T$ is a verifying realization. The CRPS is essentially an extension of the quadratic scoring rule from (ref) to continuous variables; see MW76, Her00 and GR07 for more details. Therefore, this section extends the previous one. The CRPS is strictly proper (relative to the class of probability measures with finite first moments) in the sense that \[ \mathbb{E}\big[\operatorname{CRPS}(F,T)\big] < \mathbb{E}\big[\operatorname{CRPS}(G,T)\big]\quad\text{for all }G\neq F, \] where $F$ and $G$ possess finite first moments, and $T\sim F$ GR07.
In our setting with $T\in[0,\tau]$, the CRPS simplifies to \[ \operatorname{CRPS}(G,T) = \int_{0}^{\tau}\big[G(t) - \mathds{1}_{\{t\geq T\}}\big]^2\,\mathrm{d} t. \] Note that once $X$ is realized, say at $T=t$, then the reward $\operatorname{CRPS}(G,t)$ for issuing forecast $G$ can be computed immediately, such that there is no need to wait until the terminal time $\tau$. Therefore, the reward of $\operatorname{CRPS}(G,t)$ can be paid out at time $t$ (i.e., the point in time when $X$ occurs). As in Section (ref), if this occurs, present-biased forecasters may be inclined to discount later payoffs (where $t$ is larger) more heavily than earlier ones. Since present bias is one of the most robust features of human behavior, we model the payoff $\operatorname{CRPS}(G,t)$ as being discounted by a factor of $e^{-rt}$. The choice of the discount function is arbitrary and will be lifted later (see Remark (ref)). Here, $r>0$ corresponds to a continuously compounded discount rate, with larger values of $r$ implying higher discounting of later rewards. In light of these considerations, forecasters may not actually minimize the expected CRPS, but (implicitly or explicitly) minimize the expected discounted CRPS
When this happens, the optimal forecast no longer equals the true report, as we show next.
We now graphically illustrate how the optimal forecast shifts under discounting for various intermediate $r\in(0,\infty)$. To do so, we assume that $T\sim\mathcal{U}[0,\tau]$, i.e., $T$ follows a uniform distribution supported on $[0,\tau]$. The true cdf of $T$ is then given by $F(t)=t/\tau$ for $t\in[0,\tau]$ (see the black line in Figure (ref)). It is easy to show that $\mathbb{E}[e^{-rT}\mid T\leq t]=\frac{1}{rt}(1-e^{-rt})$ in this case, such that
Therefore, under discounting, the optimal forecast is no longer a uniform distribution, but a truncated exponential distribution (with rate $r>0$). Figure (ref) plots the above cdfs $G$ for different values of $r$. In the case of no discounting (i.e., $r=0$), we have that $G=F$ (black line). As $r$ increases, more probability mass is placed on earlier times, as the red, blue and green lines in Figure (ref) show.
We now examine whether actual human forecasters alter their reports in response to the possibility of early payoffs. In doing so, we build on the theoretical framework introduced in Section (ref) for discrete time and extended to continuous time in Section (ref) (cf. Remark (ref)), where forecasts $G$ regarding the probability of future events $X$ are considered. Here, we specifically analyze forecasts made on Metaculus---a reputation-based, massive online forecasting platform. This allows us to identify the causal effect of time preferences---henceforth called time preference effects---on reports, because the platform features both time-fixed events and time-varying events. Examples of time-fixed events from Metaculus are:
The verification dates of these questions are pre-specified. Examples of time-varying events from Metaculus are:
For these time-varying events it is not clear on which day outcomes can be verified.
At Metaculus, reputational tokens are awarded based on the logarithmic error of the forecasts. In principle, this is incentive-compatible, as the logarithmic error is strictly proper GR07. In practice, however, reputational tokens are awarded immediately after the outcome is determined by content moderators: for time-fixed events, this happens shortly after the scheduled date; for time-varying events, this happens soon after the event occurs, or after the scheduled end date if the event does not occur. Consequently, this introduces a time preference effect.\footnote{Additionally, Metaculus runs tournaments where there may be monetary prizes and ranks to be claimed. Most predictions are not monetarily incentivized. Any monetary reward would only materialize after a fixed time, so the monetary incentive cannot cause the behavior that we describe in the paper. Rather, the monetary reward should moderate it.}
Assessing distorted reporting behavior is a fundamentally difficult task because we can never observe true expectations. However, if forecasters do respond to incentives to misreport due to time preferences, we should observe a systematic bias in forecasts for time-varying events that is not present in forecasts for time-fixed events. Any such systematic bias will be observable through the calibration (sometimes also called empirical reliability) which is simply how predictions correspond to actually observed frequencies of events parmigiani_decision_2009. That is, “a forecaster is well-calibrated if, for example, of those events to which he assigns a probability 30 percent, the long-run proportion that actually occurs turns out to be 30 percent” dawid_well-calibrated_1982. Forecasts on time-varying events should be systematically “off the mark”, as shown in Figure (ref). Thus, keeping everything else constant, we interpret any systematic difference in calibration between predictions on time-fixed events and forecasts on time-varying events as attributable to time preference effects.
To empirically examine incentives for misreporting, we analyze all binary events listed on the Metaculus platform at the time of data collection in August 2024 ($n=2005$). We denote whether the event occurred as $X \in \{0,1\}$; i.e., they could either occur or not. 775 events were classified by hand as time-fixed events, 798 as time-varying events that can only occur early, 197 as time-varying events that can only be ruled out early. The remaining 235 events were classified as ambiguous, either because these events can both occur early and be ruled out early (the incentive to misreport is intangible) or due to ambiguity as to when the event would be considered resolved.\footnote{One example from our data includes the question “Will Liverpool win the 2021--2022 Premier League?”. As Liverpool can theoretically secure winning the Premier League before the last game, or lose any chance of winning the Premier League before the last game, it is not clear which incentive forecasters may have to strategically misreport. This depends on how probable they deem either scenario, which can be heterogeneous.} We only consider events that must have been verified by August 2024, as including events with outcomes that could have been unobservable by August 2024 introduces a “surprisingly early bias”.
Table (ref) reports basic statistics for the two biggest groups of events. We see that (i) some events have received far more predictions than others, because the mean is far higher than the median, and that (ii) time-varying events received more predictions on average. Apart from that, the sets of events are relatively similar.
We can gather suggestive visual evidence for the presence of misreporting by plotting the calibration curves for time-varying event forecasts and time-fixed event forecasts from our entire sample. In the analysis we only include time-varying events which can only occur early, i.e., turn out to be 1, as opposed to events which can be ruled out early, i.e., turn out to be 0.\footnote{We add the latter group of events in a robustness check that is included in the Online Supplement. We find that the inclusion of this sample does not affect outcomes in a meaningful way and provides additional strong evidence for the presence of time preference effects.} To estimate aggregate calibration, we take all forecasts that range from 1% to 99% assigned probability, and plot the association with the share of events that did occur, i.e., the frequency of the verified events, which results in Figure (ref).
In Figure (ref) we see that time-varying events are systematically reported to be more probable conditional on their observed frequency. Interestingly, forecasts on time-fixed events are almost perfectly calibrated. This means that forecasts roughly correspond to the frequencies of time-fixed events, whereas forecasts on time-varying events are systematically too high, just as we would expect to see if the cause of this were time preference effects.
We cannot claim the calibration difference in Figure (ref) to be a causal effect of the incentives to misreport before addressing two potential sources of confounding. Firstly, there may be significant differences between the two sets of events---time-varying and time-fixed. Although the Metaculus forecasting tournament is full of rich and diverse events, some events are by their very nature more often time-varying, such as events related to technological progress (“When will a SpaceX Starship reach orbit?”). Other events are by their very nature fixed in time, such as elections. Events on the Metaculus platform are automatically assigned a number of topics.\footnote{See: https://www.metaculus.com/notebooks/21576/streamlined-llm-enhanced-question-discovery/} Table (ref) reports how predictions are distributed conditional on topics and event type. We observe that there are great differences between the two sets of predictions. Table (ref) includes only the most frequent topics. Predictions on topics that are not included in Table (ref) are labeled “Other”. The topics are not exclusive, i.e., events labeled “Politics” can also be labeled “Elections”. Now, if forecasters consistently misjudge certain types of events---such as being overly optimistic about technological progress---then this bias may be associated with the event classification as time-varying or time-fixed, potentially confounding the effect of interest. To address differences in topics, we balance our sample so that topics are represented equally, as described in Section (ref).
Furthermore, we must acknowledge that forecasters are not forced to make forecasts on any particular event. Therefore, the second source of confounding is that time-fixed events and time-varying events may attract systematically different forecasters, potentially causing differences in calibration. We address this by using a within-subject design, eliminating this potential source of confounding. This is described in Section (ref).
To preview the main results, we find that the aggregate picture in Figure (ref) remains intact, and that time preference effects are a powerful driver of rationally dishonest reporting, as none of the added adjustments for (i) topics and (ii) forecaster self-selection meaningfully affects the results.
Finally, one might be inclined to view question difficulty or “inherent unpredictability” to be an additional confounder. After all, forecasts on time-fixed questions may be objectively easier or harder than forecasts on time-varying events. Whilst that may very well be true, differences in inherent unpredictability should not confound our analysis because we are studying the calibration, not the precision or sharpness of forecasts. Fundamentally, strictly proper scoring rules---such as the quadratic scoring rule---can be decomposed into error from systematic bias (calibration) and “how spread out the forecasts are” (precision/sharpness) brocker_reliability_2009, parmigiani_decision_2009. Changes in inherent unpredictability should affect precision, not the calibration.\footnote{ We can use archery as an analogy to explain the problem. An archer's accuracy is judged by how close their arrows land to the center of the target. Ideally, the archer should not systematically miss the target. The average landing spot of the arrows should be near the center. If the arrows tend to gather away from the center, the archer is miscalibrated, meaning their shots are consistently off target. How difficult it is to hit the target depends on how far the archer is from it. The farther away they stand, the harder it is, and the more spread out the shots will be. We assume that being further away doesn't mean the archer becomes miscalibrated; the shots are just less accurate, not consistently off-center. Similarly, we assume that forecasters calibration does not change with question difficulty, although the precision of forecasts will be affected. As we will see later, the data do support the assumption that calibration is independent of question difficulty. If question difficulty was indeed a strong confounder, impacting calibration, we would see a drastic change in calibration as a result of controlling for topics, which should mix up question difficulty. Since re-balancing the forecasting data based on topics does not seem to affect calibration to a large degree, we find it highly plausible that both topics and difficulty (which cannot be disentangled here) do not meaningfully affect calibration.}
We address the concern that differences between forecasts are driven by systematic differences in topics through weighting forecasts by topics in order to create an artificially balanced sample, which is a way of controlling for the differences in topics. This method---called inverse probability of treatment weighting (IPTW)---effectively gives more weight to forecasts that are on likely-to-be time-varying topics (Geopolitics, Economy, ...) in the control group (time-fixed events) and more weight to forecasts that are on likely-to-be time-fixed topics (Elections, Sports, ...) in the treatment group (time-varying events), thus automatically balancing the sample austin_moving_2015. We refer to Section (ref) of the Online Supplement for details. In other words, we create within-forecaster samples of time-varying event and time-fixed event predictions that are---on average---more comparable in terms of topics, if we weight the forecasts carefully. The weights are calculated by using the probability to be treated, i.e., the probability that a certain forecast is on a time-varying event:
where $tvarying$ is an indicator variable that equals $1$ for a time-varying event and $0$ otherwise. We estimate the probability $\mathbb{P}(tvarying=1 \mid \mathbf{topics})$ via the standard logistic regression
where $\mathbf{topics}$ is a vector of 80 observable covariates (e.g., Politics).\footnote{Details regarding the covariates and coefficients can be found in the Online Supplement (Section (ref)).}
IPTW is a variation of propensity score matching, which has been found to sometimes insufficiently address bias between treatment and control samples, largely because of lacking or improper control variables smith_reconciling_2001. There is no reason to believe that we are faced with such an issue in this study because there is no complex causal interdependency between forecasts and the events that they are referring to---quite unlike most observational studies where propensity score matching is used. Therefore, controlling for additional covariates/topics is unlikely to be causing systematically wrong estimates. Furthermore, we have a very rich set of covariates/topics that we control for and we do not engage in propensity score matching but simply balance our entire sample on topics. We do not find evidence that the weighting of forecasts meaningfully affects the outcome of the analysis. This result is discussed in the next section.
Another cause for the differences in calibration in Figure (ref) could be that forecasters who issue more time-varying event forecasts are systematically different from those that make more time-fixed event forecasts. In this case, differences in aggregate calibration between the two groups of questions would actually reflect differences between individual forecasters.
We can eliminate this concern by looking at the calibration within subjects, effectively controlling for a potential self-selection of forecasters to events. By using a within-subject design, we ask: What is the expected difference between forecasts on time-varying events and time-fixed events of an individual forecaster? We use the balanced sample created by IPTW (see Section (ref)).
We collect forecasts from each forecaster $i$, which yields $k=202$ panels of forecasts. The data structure is illustrated in Table (ref). We denote a probabilistic forecast by $G\in[0,1]$. We then measure the differences in calibration between time-fixed event forecasts and time-varying event forecasts within individual forecasters and combine these differences across forecasters in a random-effects meta-analysis using the R package metafor's function rma() borenstein_basic_2010,viechtbauer_conducting_2010. Thus, we arrive at an average “incentive treatment” effect and can test whether it is significantly different from zero.
For each forecaster $i$, we estimate a calibration curve $\mu_{i}(G) := \mathbb{P}(X=1 \mid i,G)$, i.e., $\mu_{i}$ is forecaster $i$'s observed event rate among predictions of value $G$. We obtain $\mu_{i}$ by binning forecaster $i$'s predictions at the 0.01-level. We estimate the calibration $\mu_{i}$ as
where $\beta_0$ denotes the intercept, $\beta_1$ the slope for the time-fixed event calibration, and $\beta_2$ and $\beta_3$ measure the change in intercept and slope for time-varying events.\footnote{As the data may be heteroskedastic, we use heteroskedasticity-robust standard errors from mackinnon_heteroskedasticity-consistent_1985.} Although calibration must not be linear per se, perfect calibration, where reported probabilities correspond to observed frequencies, is linear; see also Figure (ref).\footnote{The calibration of misreported predictions in Section (ref) is also linear. However, this depends on an arguably arbitrary modeling choice.} Thus, we use linear regression as the most valid available approach. We practically estimate two linear functions---one for time-varying event forecasts and one for time-fixed event forecasts. However, by using time-varying events as a treatment dummy we incorporate both into one model, as in (ref).
We test that individual forecasters are not systematically misreporting their expectations of time-varying events. This is equivalent to the variable $tvarying$, that signals whether the event is time-varying, having no predictive power, i.e.,
To a lesser extent, we furthermore expect forecasters to be well-calibrated, i.e., their expectations to be accurate. Since this is in line with an intercept of 0 and a slope of 1,
In order to limit the influence that a single event might have on the calibration of individual forecasters we restrict our analysis to forecasters who have made predictions on at least 40 different time-fixed and time-varying events, and at least a total of 80 predictions on time-varying and time-fixed events respectively. Forecasters ($k=202$) can make multiple predictions on the same event, which is why we require both a sufficient number of predictions and a minimum number of events that these predictions refer to. Multiple predictions on the same event are common, as they reflect new information (updates). Within-event dependence arises as the event outcome on updated predictions will be the same. Furthermore, the updated predictions could be serially correlated. We run a robustness check using only one forecast per event per forecaster (first or last). It turns out that this is not an issue, as collapsing the data in this way does not meaningfully affect our result, and refer to the supplemental materials for details. We obtain the mean estimates and standard errors of the four coefficients $\beta_0,\ldots,\beta_3$ using a standard random-effects model, and test the joint hypotheses $\mathcal{H}_0^{tvarying}$ and $\mathcal{H}_0^{calib}$ using the Bonferroni correction miller_simultaneous_1981.\footnote{We use the metafor package in R and estimates of heterogeneity across forecasters as specified in dersimonian_meta-analysis_1986. Graphical representations in the form of forestplots of the effects (Figures (ref)--(ref)) can be found in the Online Supplement (Section (ref)).}
Doing so, we reject both hypotheses $\mathcal{H}_0^{tvarying}$ and $\mathcal{H}_0^{calib}$ since $\beta_0$,$\beta_1$, and $\beta_3$ are significantly different from the hypothesized value; indeed, the Bonferroni-adjusted $p$-values are smaller than $0.001$ for both hypotheses. The detailed outcomes are reported in Table (ref), where column 3 corresponds to the regression including controls for topics (IPTW) and forecaster self-selection (Within-subject). To see how the added controls affect the result, we run a simple OLS regression on the aggregate predictions dataset, dropping both the within-subject control and topic balancing. This regression---which is essentially a fitted line in Figure (ref)---is in line with both theoretical expectations and the more carefully controlled estimates. We report estimated coefficients in column 1 of Table (ref). Additionally, we repeat the regression analysis using only the within-subject specification but without balancing for topics. These results are reported in the second column in Table (ref).
Forecasters are not perfectly calibrated because the intercept $\beta_0$ is smaller than 0 and the slope for $\beta_1$ is larger than 1. This suggests that forecasters skew their predictions toward 50%, such that the calibration curve is slightly “s-shaped”.\footnote{Such a “center bias” is commonly observed in forecasting data and is partially explainable by risk aversion hurley_experimental_2005,danz_belief_2022.} Nonetheless, the calibration of forecasters on time-fixed events is very good overall.
On the other hand, for time-varying events, the calibration of forecasters is significantly different, i.e., time preferences affect predictions to a significant degree. Forecasters report systematically higher predictions when the event is time-varying, as the difference in slope is negative ($\beta_3 <0$). A lower coefficient refers to a lower frequency ceteris paribus. Thus, a lower coefficient implies a higher prediction for any given frequency. Our analysis finds that $\beta_2 > 0$ (although the effect is not statistically significant), which means that the intercept difference between control and treatment is positive (though small).
We can combine these observations by looking at all coefficients and (model-implied) calibration in Table (ref), which reports the expected prediction $\mathbb{E}[G \mid tvarying, \mu]$ implied by (ref) for a given true event probability $\mu \in (0,1)$, for time-fixed and time-varying events, i.e., the regression inverted to solve for $G$. Predictions on time-varying events are higher than predictions on time-fixed events because the difference in slope $\beta_3$ is far larger than---and completely dominates---$\beta_2$, which is close to zero. Why is $\beta_2$ still positive? We conjecture that $\beta_2$ is positive because $\beta_0$ is already large and negative for reasons unrelated to time preferences, such as all predictions involving low-probability judgment being too high. Thus, $\beta_0$ might be masking potential differences between intercepts ($\beta_2$).
We see that the extent to which forecasters misreport their own expectation is quite large. Table (ref) shows that the expected prediction from an average forecaster for a time-fixed event which occurs with 30% probability would be 32.9%, whilst the expected prediction for an equally likely time-varying event would be 36.4%. On time-varying events that are expected to occur with 70% probability (frequency and fixed-event-prediction), forecasters reported a whopping 79% probability on average. Table (ref) reports that predictions on time-varying events that occurred with 90% frequency are at an impossible 100.4%. From Figure (ref) we can see that the observed frequency of time-varying events is never above 90%, such that it would indeed take a larger-than-100% prediction to get to a frequency of over 90% for time-varying events.
We find no pattern in how much forecasters misreport their predictions. Moderate misreporting seems to be common across forecasters and not strongly correlated with how many predictions a forecaster made. We plot the estimates for $\beta_3$ for the most active 39 forecasters in the forestplot in Figure (ref). The farther the estimates of $\beta_3$ are from the center, the stronger the difference in slope between calibration on time-fixed and time-varying events. The individual forecasters are sorted from most predictions made (top) to least predictions made. The bottom entry marks the average estimate.
Finally, we are interested in how time preferences impact the accuracy of forecasts. The effect on accuracy is negative because a strictly worse calibration will ceteris paribus lead to lower accuracy. Since any strictly proper scoring rule penalizes systematic bias (calibration) and precision (how close forecasts are to true outcomes) separately GR07, we can isolate the contribution of systematic bias to the accuracy, as measured by the mean quadratic scoring rule or Brier score in our dataset. We find that the predictions on time-fixed events have an average systematic bias that contributes $0.0023$ to the Brier score, whereas the predictions on time-varying events have an average systematic bias that contributes roughly $0.011$ to the Brier score. Setting the calibration of time-varying-event predictions equal to those of time-fixed-event predictions would have improved accuracy by more than 3.5%.\footnote{We refer to Section (ref) of the Online Supplement for more details.} This is just another way of saying that time-varying event predictions are much less well calibrated.
This study provides strong evidence that time preferences impact real-world forecasts. We bring forward intuitive, theoretical, and empirical accounts, all of which provide evidence of their own. Whilst our empirical study is observational---limiting our ability to assure equality of treatment and control---the study has the huge advantage of being conducted in the real world with actual forecasts on future events spanning multiple years. Therefore, our empirical study possesses strong external validity. Furthermore, we find time preference effects in all model specifications; see Table (ref). This leads us to believe that differences between treatment and control, which are minimized using our IPTW and the within-subject design, do not strongly confound the measured time preference effects.
We remark that additional factors need to be considered when interpreting the results. A cognitive bias that is closely associated with time-varying events cannot be distinguished from an effect that is caused by human preferences.\footnote{We thank an attendant of our talks for this remark.} However, the data suggests that the observed effect is caused by time preferences. The reason for this lies in the rule of complement: We can take any event and ask for the complement, i.e., the chance that it will not happen. Theoretically, as the predictions on events that can happen early are inflated, the complement would have to be deflated, i.e., predictions should be systematically too low. We indeed see this pattern in a limited set of events on Metaculus where the questions are formulated to ask for the complement.\footnote{We refer to Figure (ref) in the Online Supplement. We also discuss how follow-up studies could collect evidence on either hypothesis.} It seems implausible that a cognitive bias would cause distortions that are sensitive to the rule of complements.
In fact, we suspect that our empirical study far underestimates the degree to which predictions would be affected had forecasters minimized their expected discounted error. Firstly, we know from the literature on belief elicitation that humans have a strong preference for being seen as honest and actually being honest abeler_preferences_2019. Therefore, humans seem to strike a balance between honesty and reward-maximization when they can gain from misreporting their beliefs. Forecasters may be distorting their forecasts slightly but not drastically, even if they would benefit from doing so. Secondly, forecasters may be (partially) unaware of the scoring and the opportunity to gain from misreporting. Furthermore, our sample contains mainly short-term forecasts. Although a handful of forecasts is as much as 7-years-ahead, the vast majority of forecasts is less than 1-year- ahead (median: $118$ days, mean: $227$ days). Since the issue of time preferences would increase with the expected time-to-event, we expect that long-term forecasts suffer more heavily from time preference effects than short-term forecasts. Most long-term forecasts from Metaculus are not in our dataset because the events are not yet verified.
A natural question arises: How large is this distortion in practice? The magnitude of the distortion should depend on both the discount rate $r$ and the expected time-to-event. The longer the time-horizon over which events are expected to unfold, the larger the distortion should be. However, estimating the size of the distortion as a function of time-to-event (or estimating $r$) is less straightforward than it may seem. In our empirical study, simply conditioning on the observed time-to-event introduces a surprisingly-early bias, a form of selection bias: events with a very short observed time-to-event are, almost by definition, events that occurred sooner than expected. In such cases, we should not expect observed frequencies to match forecasted probabilities, even if those forecasts were perfectly calibrated.
Overall, we find that time preference effects are a serious problem that can systematically affect forecasts in all areas, and that forecasting practitioners should be aware of. Moreover, this problem extends towards all statements regarding the future which can be “right” earlier than they can be “wrong”. Failures and disruptions are more likely to occur---and become apparent---in the short-term than sustained, smooth operations or simply “business as usual”. Therefore, time preferences should lead to overestimation of risks in various areas, such as credit risk, life expectancy estimation, natural disaster forecasting, reliability engineering, technological progress forecasting, pandemic forecasts, geopolitical analysis, environmental risk, and information security.\footnote{It is commonly understood that experts may overestimate risks in order to err on the side of caution. This is often (mistakenly) called risk aversion or conservatism. However, if we wanted to err on the side of caution, being systematically misinformed about risks is not necessary. Risk-averse actors can simply choose to take the safe route based on unbiased estimates of risk. It seems far more plausible that experts overestimate risks because it is better for them, not only because of time preferences, but also because unexpected failures are certainly penalized more harshly than an undue overestimate of that risk. Thus, experts may overestimate risks in order to protect their reputation, or avoid blame.} As a result, we should expect that decision-makers and regulators apply standards in systematically suboptimal ways, over- or underemphasizing safety and preparedness. In particular, forecasts for technological milestones seem to be overly optimistic tichy_over-optimism_2004, which could---at least partially---be explained by time preference effects.
As ottaviani_strategy_2006 put it: “Forecasting is proving to be an apt laboratory for improving our understanding of strategic communication and positioning by non-partisan informed agents. The availability of data sets and the richness of institutional details can inspire and give discipline to our theorizing. The insights gained can be helpful in shedding light on a number of other social and economic problems [...]”. We remark that time preferences do not necessarily lead to a decrease in the accuracy of forecasts---as they do in the Metaculus tournament---if they happen to cancel out or mitigate other existing biases. However, in most cases we expect time preferences to be unwanted and potentially harmful.
The remaining question is how to reduce or eliminate them. We identify four different strategies to eliminate or cope with time preference effects and investigate each in turn. However, there is no solution that comes without drawbacks.
This enumeration of possible solutions suggests that forecasts of time-varying events are inherently challenging to evaluate and incentivize. Any attempt to do so risks creating misaligned incentives. Given the importance of foresight related to future events, this is an important finding of its own.
The reason for its unsolvability is obvious once we understand time preference effects (as defined in this study) as the less-well-known cousin of the general problem of long-term forecasting: it is difficult to incentivize long-term predictions given that their evaluation is far away. Since there is no easy “solution” to this problem, it follows that time-varying event predictions---even if relatively short-term---are also affected.
This study investigates the incentive to report earlier occurrence of events in order to increase (perceived) short-term benefits. If a forecasted event can occur earlier than other potential outcomes, forecasters can “bet” on this early resolution and can only be “correct” in the short term. As a result of such distortions, overall forecasting performance decreases. Utilizing forecasting data from Metaculus, we observe a substantial and significant inflation of predictions when the forecasted event can occur early. Therefore, we have reason to believe that forecasts in important domains such as technological forecasting, extreme weather forecasting, pandemic forecasting, geopolitical and financial risk may be systematically biased. We also find that there is limited practical leeway for aligning incentives with honest reporting, such that awareness related to potentially misaligned incentives is critically important. Future research in this area could involve studying time preference effects in applied areas. Furthermore, conducting a controlled experiment could deepen our understanding and provide valuable evidence in addition to this study. Such an experiment would allow us to obtain various auxiliary measurements, including the implied discount factor $r$.
\singlespacing
The authors report there are no competing interests to declare.
The code used to generate the results and figures in this paper is publicly available in a GitHub repository at \nolinkurl{https://github.com/PfadQualle/time_preferences_in_forecasting}. This repository includes all scripts necessary to reproduce the analysis, conditional on access to the underlying data. The forecasting data used in this study were obtained from Metaculus and are subject to restrictions that prevent redistribution. Researchers seeking access to these data should contact Metaculus directly. Upon reasonable request, and where permitted by the terms of our data use agreement with Metaculus, the corresponding author will be happy to provide guidance on replicating and extending the empirical analysis in this study.