EconBase
← Back to paper

Testing Forecast Rationality for Measures of Central Tendency

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

224,577 characters · 31 sections · 147 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing Forecast Rationality for Measures of Central Tendency

{ \today }

\rule{\linewidth}{.4pt} Abstract \\[.3\baselineskip] \onehalfspacing Rational respondents to economic surveys may report as a point forecast any measure of the central tendency of their (possibly latent) predictive distribution, for example the mean, median, mode, or any convex combination thereof. We propose tests of forecast rationality when the measure of central tendency used by the respondent is unknown. We overcome an identification problem that arises when the measures of central tendency are equal or in a local neighborhood of each other, as is the case for (exactly or nearly) symmetric distributions. As a building block, we also present novel tests for the rationality of mode forecasts. We apply our tests to income forecasts from the Federal Reserve Bank of New York's Survey of Consumer Expectations. We find these forecasts are rationalizable as mode forecasts, but not as mean or median forecasts. We also find heterogeneity in the measure of centrality used by respondents when stratifying the sample by past income, age, job stability, and survey experience. \\ [.3\baselineskip] \\ Keywords: forecast evaluation, partial identification, survey forecasts, mode forecasts\\ J.E.L. Codes: C53, D84, E27

Introduction

\doublespacing

Economic surveys are a rich source of information about future economic conditions, yet most economic surveys are vague about the specific statistical quantity the respondent should report. For example, the New York Federal Reserve's labor market survey asks respondents “What do you believe your annual earnings will be in four months?” A reasonable response to this question is the respondent reporting her mathematical expectation of future earnings, or her median, or her mode; all common measures of the central tendency of a distribution. When these measures coincide, as they do for symmetric unimodal distributions, this ambiguity does not affect the information content of the forecast. When these measures differ, the specific measure adopted by the respondent can influence its use in other applications, and testing rationality of forecasts becomes difficult. Subjective forecast distributions have been found to be asymmetric for many important economic variables as GDP growth adrian2019vulnerable, bekaert2019link, inflation rates garcia2007can, firm earnings Foster1986, givoly2000changing, gu2003earnings and consumer expenses Howard2022, motivating the need for reliable forecast evaluation methods for the different measures of central tendency.

While the assumption of a mean forecast\footnote{We use the phrase “mean forecasts,” or similar, as shorthand for the forecaster reporting the mean of her predictive distribution as her point forecast.} is common in the economic and statistical literature (see, e.g., coibion2015information, bordalo2020overreaction), it may be mistaken in some applications. For example, Knuppel2012 conclude that point forecasts of inflation published by central banks almost always correspond to the modes of the forecast densities, and ReifschneiderTulip2017 suggest that forecasts from the U.S. Board of Governors are best interpreted as mode forecasts. Howard2022 and zhao2022internal find that some survey forecasts are more consistent with the respondents' modes than means or medians. In an early experimental study, peterson1964mode found that respondents could accurately predict the mode and median if incentivized correctly, but had difficulty reporting accurate estimates of the mean, and a more recent study by KroegerThibaud2019 reports that participants asked to summarize their predictive distribution responded most frequently with their mode. We propose that other measures of central tendency deserve further consideration, especially given the growing evidence that point forecasts may reflect the mode.

Given the ambiguity around which specific measure of central tendency is used by survey respondents, we consider a class of such measures defined by the set of convex combinations of the mean, median, and mode.\footnote{These are the three measures described in introductory statistics textbooks (McClave2017), in previous studies (engelberg2009comparing), in central bank publications (the BoE2019's quarterly inflation reports), and in psychological work, kahneman1982judgment. Nevertheless, our approach can easily be extended to consider a broader set of measures of centrality.} {\color{black}This approach, naturally, nests each of these measures separately, and also allows for survey respondents who use a mixture of measures of central tendency, e.g., a respondent who “tilts” her mean forecast in the direction of the most likely outcome (the mode) or towards a robust measure of location (the median). Such mixtures may arise through the formal minimization of a mixture of the respective loss functions for these measures, or informally as the respondent uses “expert judgment” to adjust a forecast of a specific measure of central tendency.}

Similar to EKT2005, we propose a testing framework that nests the mean as a special case, but unlike that paper we allow for alternative forecasts within a class of measures of central tendency, rather than measures that represent other aspects of the predictive distribution (such as non-central quantiles or expectiles). Our approach faces an identification problem: for symmetric distributions the combination weight vector is unidentified, for “mildly” asymmetric distributions the weight vector is weakly identified, and even for strongly asymmetric distributions the weight vector may only be partially identified. Economic variables may fall into any of these cases, and a valid testing approach must accommodate these measures of central tendency being equal, unequal, or in a local neighborhood of each other. We use the work of StockWright2000 to obtain asymptotically valid confidence sets for the combination weights and to test forecast rationality.

Depending on the skewness of the respondent's predictive distribution, our approach may allow the researcher to distinguish mean, median, and mode forecasts based on point forecasts and realizations alone. An alternative approach was proposed in engelberg2009comparing, who combine point and density forecasts to determine whether the point forecasts are the mean, median, or mode of the predictive density. In their study of respondents to the Federal Reserve Bank of Philadelphia's Survey of Professional Forecasters, engelberg2009comparing find that most point forecasts are consistent with the bounds derived for all three functionals. zhao2022internal similarly studies inflation forecasts from the Federal Reserve Bank of New York's Survey of Consumer Expectations and finds the mode of the density forecast to be most consistent with the point forecasts, but the mean and median are also valid for a majority of forecasts. Our testing approach accommodates the fact that these functionals may be hard, or impossible, to separately identify in data.

Before implementing the above test for rationality for a general forecast of central tendency, we must first overcome a lack of rationality tests for mode forecasts. Rationality tests for mean forecasts go back to at least MincerZarnowitz1969, see ETbook2016 for a recent survey, while rationality tests for quantile forecasts (nesting median forecasts as a special case) are considered in christoff and GaglianoneJBES2011. A critical impediment to similar tests for mode forecasts is that the mode is not an “elicitable functional” Heinrich2014, meaning that it cannot be obtained as the solution to an expected loss minimization problem.\footnote{Gneiting2011 provides an overview of elicitability and identifiability of statistical functionals and shows that several important functionals such as variance, Expected Shortfall, and mode are not elicitable. FisslerZiegel2016 introduce the concept of higher-order elicitability, which facilitates the elicitation of vector-valued (stacked) functionals such as the variance and Expected Shortfall, though not the mode Dearborn2019.}$^,$\footnote{For categorical data, rationality tests for mode forecasts were established in das1999comparing and extended in madeira2018testing. Our focus is on continuously distributed target variables, and so we cannot use their results.} We obtain a test for mode forecast rationality by first proposing novel results on the asymptotic elicitability of the mode. We define a functional to be asymptotically elicitable if there exists a sequence of elicitable functionals that converges to the target functional. We consider the (elicitable) “generalized modal interval,” defined in detail in Section (ref), and show that it converges to the mode for a general class of probability distributions. We combine these results with recent work on mode regression (Kemp2012; Kemp2019) and nonparametric kernel methods to obtain mode forecast rationality tests analogous to well-known tests for mean and median forecasts. In addition to size control, we show that the proposed test has non-trivial asymptotic power against both fixed and local alternative hypotheses.

We evaluate the finite sample performance of the new mode rationality test and of the proposed method for obtaining confidence sets for measures of centrality through an extensive simulation study. We use cross-sectional and time-series data generating processes with a range of asymmetry levels. We find that our proposed mode forecast rationality test has satisfactory size properties, even in small samples, and exhibits strong power across different misspecification designs. Our simulation design allows us to consider the four identification cases that can arise in practice: strongly identified (skewed data), where the mean, median and mode differ; unidentified (symmetric unimodal data), where all centrality measures coincide; weakly identified (mildly skewed data), where the centrality measures differ but are close to each other; and partially identified (skewed location-scale data), where one centrality measure is a convex combination of the other two. We find that in the symmetric case, the resulting confidence sets contain, correctly, the entire set of convex combinations of mean, median and mode. In the asymmetric cases, our rationality test is able to identify the combination weights corresponding to the issued centrality forecast.

We apply the new tests to the income survey responses from the Survey of Consumer Expectations conducted by the Federal Reserve Bank of New York. In the full sample of respondents, we find the intriguing result that we can reject rationality with respect to the mean or median, however we cannot reject rationality when interpreting these as mode forecasts, suggesting that survey participants report the anticipated most likely outcome rather than the average or median. When allowing for cross-respondent heterogeneity, we find that forecasts from low-income younger survey respondents cannot be rationalized using any measure of central tendency, while forecasts from high-income respondents, regardless of their age, are rationalizable for many different measures of centrality. We also find evidence of learning between survey rounds kim2022learning for high-income respondents, but much less so for low-income respondents. We compare our rationality test results with those obtained using the approach of EKT2005 (EKT), which allows for rational optimism or pessimism. We find no cases where allowing for optimism or pessimism “overturns” a rejection of rationality based on a measure of centrality. As a stark example, we find that forecasts from low-income younger respondents cannot be rationalized as any centrality measure, nor as a measure in the EKT framework.

Our paper is related to the large literature on forecasting under asymmetric loss, see Granger69, christoffersen1997optimal, EKT2005, PattonTimmermann2007 and biases amongst others. The work in these papers is motivated by the fact that forecasters may wish to use a loss function other than the omnipresent squared-error loss function. The use of asymmetric loss functions generally leads to point forecasts that differ from the mean (though this is not always true, see Gneiting2011 and patton2014comparing), and generally these point forecasts are not interpretable as measures of central tendency. For example, christoffersen1997optimal show that the linex loss function implies an optimal point forecast that is a weighted sum of the mean and variance, while biases find that their sample of macroeconomic forecasters report an expectile with asymmetry parameter around 0.4. Instead of moving from the mean to a point forecast that is not a measure of location, our novel approach considers moving only within a general class of central tendency measures.

The remainder of the paper is structured as follows. In Section (ref) we propose new forecast rationality tests for the mode based on the concepts of asymptotic elicitability and identifiability. Section (ref) presents forecast rationality tests for general measures of central tendency, allowing for weak and partial identification. Section (ref) presents simulation results on the finite-sample properties of the proposed tests, and Section (ref) presents the empirical application using income survey forecasts. A supplemental appendix contains {\color{black}all proofs, further theoretical and simulation results}, and empirical analyses of the Federal Reserve Board's “Greenbook” forecasts of US GDP growth and random walk forecasts of exchange rates.

Eliciting and Evaluating Mode Forecasts

General Forecast Rationality Tests

Let $Z_t = \big( Y_{t}, X_t, \widetilde{\mathbf{h}}_t \big)$ {\color{black}with $t \in \mathbb{N}$} be a stochastic process defined on a common probability space $\big( \Omega, \mathcal{F}, \mathbb{P} \big)$. $Y_{t+1}$ denotes the (scalar) variable of interest, $\widetilde{\mathbf{h}}_t$ denotes a vector of variables known to the forecaster at the time she issues her point forecast for $Y_{t+1}$, which is denoted $X_t$.\footnote{Given our focus on survey forecasts, where the underlying model used by the respondents, if any, is unknown, we take the forecasts as given, putting this paper in the general framework of giacomini2006tests, as opposed to that of West1996.} {\color{black}We define the information set $\mathcal{F}_t = \sigma \big\{ \widetilde{\mathbf{h}}_s; s \le t \big\}$ as the $\sigma$-field containing all information known to the forecaster at time $t$. Throughout the paper, we assume that $X_t \in \mathcal{F}_t$.} We denote the distribution of $Y_{t+1}$ given $\mathcal{F}_t$ by $F_t$, and with corresponding density $f_t$. Neither the forecaster's information set, $\mathcal{F}_t$ nor her predictive distribution, $F_t$, are assumed known to the econometrician. In conducting the forecast rationality test, we assume that the econometrician uses an $\mathcal{F}_t$-measurable $(k \times 1$) vector $\mathbf{h}_t$, which can be thought of as a subset of $\widetilde{\mathbf{h}}_t$.\footnote{The intepretability of the outcome of a forecast rationality test, including ours, critically depends on the “test vector” or “instrument vector,” $\mathbf{h}_t$, being observable to the forecaster. In our empirical analysis we only use vectors that are guaranteed to satisfy this requirement.} Note that since $F_t$ is not observed, we do not know whether this distribution is strongly skewed, mildly skewed, or symmetric, and thus our inference method must be valid for all of these possibilities. Conditional expectations are denoted $\mathbb{E}_t[\cdot] =\mathbb{E}[\cdot|\mathcal{F}_t]$. We use $\cal{P}$ to denote a class of distributions.

We start by considering rationality tests (also known as calibration tests; Nolde2017) for the mean, i.e.\ we assume that the forecasts $X_t$ are one-step ahead mean forecasts for $Y_{t+1}$. We are interested in testing if these forecasts are rational , which would imply the null hypothesis:

align[align omitted — 122 chars of source]

We test this hypothesis using the “identification function” for the mean, which is simply the difference between the forecast and the realized value, i.e., the forecast error:\footnote{The identification function for a point forecast can be obtained as the first derivative of any loss function that elicits that forecast. The quadratic loss function elicits the mean, and so its identification function is simply the forecast error, up to scale and sign. In econometrics the forecast error is usually defined as $Y_{t+1}-X_t$, that is, as the negative of the identification function in equation ((ref)). Given the important role that forecast identification functions play in this paper, we adopt the definition for $\varepsilon_{t}$ given in equation ((ref)), and we refer to $\varepsilon_{t}$ as the forecast error.}

align[align omitted — 112 chars of source]

Specifically, the null hypothesis in (ref) implies that the identification function $V_{\mathrm{Mean}} \big( X_t, Y_{t+1} \big)$ is uncorrelated with any $(k \times 1)$ instrument vector $\mathbf{h}_t \in \mathcal{F}_t$, which provides a testable moment condition. Under the above null hypothesis and subject to standard regularity conditions, it is straight forward to show that $T^{-1/2} \sum_{t=1}^T V_{\mathrm{Mean}} \big(X_t, Y_{t+1} \big) \mathbf{h}_t \overset{d}{\to} \mathcal{N} \big( 0,\Omega_{\mathrm{Mean}} \big)$ as $T \rightarrow \infty$, and that

align[align omitted — 306 chars of source]

as $T \rightarrow \infty$, where $\widehat{\Omega}_{T,\mathrm{Mean}}$ is a consistent estimator of $\Omega_{\mathrm{Mean}}$. This result facilitates testing whether given forecasts $X_t$ are rational mean forecasts for the realizations $Y_{t+1}$ by using the test statistic $J_T$ in equation ((ref)) to test for uncorrelatedness of the identification function $V_{\mathrm{Mean}} \big( X_t, Y_{t+1} \big)$ and the instrument vector $\mathbf{h}_t$. As in most other tests in the literature, this is of course only a test of a necessary condition for forecast rationality, and the conclusion may be sensitive to the choice of instruments, $\mathbf{h}_t$.\footnote{Like most of the forecast evaluation literature, we assume that the vector of instruments is of fixed and finite length. A Bierens82-type test, where the length of the vector diverges with the sample size, is considered for forecast evaluation in, for example, CS02.}

The test statistic in equation ((ref)) is only informative about rationality if the forecasts are interpreted as being for the mean of $Y_{t+1}$. The decision-theoretic framework of identification functions and consistent loss functions is fundamental for generalizations to other measures of central tendency, such as the median and the mode. For a general real-valued functional $\Gamma: \mathcal{P} \longrightarrow \mathbb{R}$, a strict identification function $V_\Gamma(x,Y)$ is defined by being zero in expectation if and only if $x$ equals the functional $\Gamma(F)$. Strict identification functions are generally obtained as the derivatives of strictly consistent loss functions, which are defined as having the functional $\Gamma(F)$ as their unique minimizer (in expectation). A functional is called identifiable if a strict identification function exists, and is called elicitable if a strictly consistent loss function exists. See Gneiting2011 for a general introduction to elicitability and identifiability.

The forecast error $X_t - Y_{t+1}$ is a strict identification function for the mean, and a strict identification function for the median is given by the step function:

align[align omitted — 140 chars of source]

We obtain a test of median forecast rationality by replacing $ V_{\mathrm{Mean}}$ and $\widehat{\Omega}_{T,\mathrm{Mean}}$ by $V_{\mathrm{Med}}$ and $\widehat{\Omega}_{T,\mathrm{Med}}$ in equation ((ref)).

The Mode Functional

In contrast to the mean and the median, rationality tests for mode forecasts are more challenging to consider. The underlying reason is that there do not exist identification functions for the mode for random variables with continuous Lebesgue densities Heinrich2014, Dearborn2019. In this section we simplify notation and refer to the target variable and forecast as $Y$ and $x$. We define the mode for random variables with continuous Lebesgue densities as the global maxima of the density function.\footnote{ More generally, the mode is often defined as the limit, as $\delta \to 0$, of the modal midpoint functional $\operatorname{MMP}_\delta$, given in equation ((ref)) below Gneiting2011, Dearborn2019. This definition coincides with the global maxima of the density function for distributions with continuous Lebesgue density; and it coincides with the points of maximal probability for discrete distributions Heinrich2014. Note that our definition of the mode as the global maxima of a density function is only valid for distributions with continuous Lebesgue densities as otherwise the density function can be modified on singletons (null sets in the distribution) without altering the underlying probability measure. } We make the following distinction in the notion of unimodality.

definitionAn absolutely continuous distribution is defined as {\color{black}(a) weakly unimodal if it has a unique mode, i.e., a unique point $m \in \mathbb{R}$ such that $f(x) < f(m)$ for all $x \neq m$, and (b) strongly unimodal if there exists a unique point $m \in \mathbb{R}$ such that for its differentiable density function $f$, it holds that $f'(x) \ge 0$ if $x<m$ and $f'(x) \le 0$ if $x>m$.}

HeinrichFissler2021 show that even for the class of strongly unimodal distributions, neither strictly consistent loss functions nor strict identification functions exist for the mode. Gneiting2011 notes that it is sometimes stated informally that the mode is an optimal point forecast under the loss function $L_\delta(x,Y) = \mathds{1}_{\{ | x - Y | \le \delta \}}$ for some fixed $\delta > 0$. In fact, this loss function elicits the midpoint of the modal interval (also known as the modal midpoint, or MMP) of length $2 \delta$. The MMP of $Y \sim P$ is defined as the midpoint of the interval of length $2 \delta$ that contains the highest probability:

align[align omitted — 149 chars of source]

More formally, it holds that for any $\delta > 0$ small enough, the modal midpoint is well-defined for all distributions with unique and well-defined mode, and it holds that $ \lim_{\delta \downarrow 0}\operatorname{MMP}_\delta( P ) = \operatorname*{Mode}( P )$ Gneiting2011. In a similar manner, Eddy1980, Kemp2012 and Kemp2019 propose estimation of the mode by estimating the modal interval with an asymptotically shrinking length. We formalize these ideas in the decision-theoretical framework in the following definition.

definitionThe functional $\Gamma: \mathcal{P} \longrightarrow \mathbb{R}$ is asymptotically elicitable (identifiable) relative to $\mathcal{P}$ if there exists a sequence of elicitable (identifiable) functionals $\Gamma_\delta: \mathcal{P} \longrightarrow \mathbb{R}$, such that $\Gamma_\delta(P) \to \Gamma(P)$ as $\delta \to 0$ for all $P \in \mathcal{P}$.

As the modal midpoint converges to the mode, this establishes asymptotic elicitability for the mode functional for the class of weakly unimodal probability distributions with continuous Lebesgue densities. Unfortunately, this does not directly allow for asymptotic identifiability of the mode as any pseudo-derivative of the loss function $L_\delta$ equals zero. We establish asymptotic identifiability of the mode through the generalized modal midpoint.

definitionGiven a kernel function $K(\cdot)$ and bandwidth parameter $\delta$, the generalized modal midpoint, $\Gamma_\delta^K(P)$, of $Y \sim P$ is defined as \begin{align} \begin{aligned} \Gamma_\delta^K(P) = \operatorname*{arg\,min}_{x \in \mathbb{R}} \mathbb{E} \left[ L^K_\delta(x,Y) \right], \qquad where \qquad L^K_\delta(x,Y) = - \frac{1}{\delta} K\left( \frac{x-Y}{\delta} \right). \end{aligned} \end{align}

The familiar modal midpoint is nested in this definition by using a rectangular kernel for $K(u) = \mathds{1}_{\{ | u | \le 1 \}}$, and this definition allows for smooth generalizations. As this definition involves the argmin of a function, we first establish that this is well-defined and that it converges to the mode functional. The following theorem also considers identifiability of the generalized modal midpoint.

theoremLet $K$ be a strictly positive kernel function on the real line that is log-concave, i.e.\ $\log(K(u))$ is a concave function, and additionally let $\int K(u) \mathrm{d}u = 1$ and $\int |u| K(u) \mathrm{d}u < \infty$. Let $\mathcal{P}$ be the class of absolutely continuous and weakly unimodal distributions with bounded and Lipschitz-continuous density and let $\tilde{\mathcal{P}} \subset \mathcal{P}$ be the subclass of strongly unimodal distributions. \begin{enumerate}[label=(\alph*)] • The functional $\Gamma_\delta^K$ induced by the loss function ((ref)) is well-defined for all $\delta > 0$ and $P \in \mathcal{P}$. • It holds that $\Gamma^K_\delta(P) \to \operatorname*{Mode}(P)$ as $\delta \to 0$ for all $P \in \mathcal{P}$. • If $K$ is differentiable, for all (fixed) $\delta>0$ and $P \in \tilde{\mathcal{P}}$, it holds that the function \begin{align} V_\delta^K (x,Y) = \frac{\partial}{\partial x} L^K_\delta(x,Y) = - \frac{1}{\delta^2} K' \left( \frac{x-Y}{\delta} \right) \end{align} is a strict identification function for $\Gamma_\delta^K$. In particular, the generalized modal midpoint is identifiable and the mode is asymptotically identifiable with respect to $\tilde{\mathcal{P}}$. \end{enumerate}

This theorem shows that for the classes of weakly (strongly) unimodal distributions, the generalized modal midpoint is elicitable (and identifiable), and consequently, the mode is asymptotically elicitable (and identifiable). While $V_\delta^K (x,Y)$ being an identification function for the generalized modal midpoint is obvious from Definition (ref), Theorem (ref)(ref) establishes its strictness.

For a fixed $\delta > 0$, strict identifiability is doomed to fail when both the underlying distribution and the kernel function have bounded support as the expected identification function equals zero for values far outside both supports. Theorem (ref)(ref) shows that employing log-concave kernels with infinite support circumvents this problem. While kernels with bounded support generally exhibit a superior performance in nonparametric statistics, this proposition motivates the use of kernels with infinite support like the Gaussian density function. Furthermore, the Gaussian density function, among many others, satisfies the required log-concavity of the kernel function.\footnote{Theorem (ref)(ref) also holds if log-concavity of the underlying density, instead of the kernel function, holds. This illustrates that kernels with bounded support can be employed at the cost of restricting the class of distributions.}

{\color{black}Beyond the rationality tests proposed in the following, the concept of asymptotic elicitability is of interest in its own right. Asymptotic elicitability may facilitate forecast evaluation and comparison for novel statistical functionals. For example, schmidt2021belief propose an elicitation procedure for the maximum, a functional that generally is not elicitable bellini2015elicitable.}

Forecast Rationality Tests for the Mode

Having established the asymptotic identifiability of the mode in the previous section, we now consider rationality testing of mode forecasts, i.e.\ testing the following null hypothesis,

align[align omitted — 152 chars of source]

While classical, $\sqrt{T}$-consistent rationality tests based on strict identification functions are unavailable for the mode due to its non-identifiability, we next propose a rationality test for mode forecasts based on an asymptotically shrinking bandwidth parameter $\delta_T$. Consider the (asymptotically valid) identification function $V^K_{\delta_T}$ with bandwidth $\delta_T$, and multiplied by the instruments $\mathbf{h}_t$,

align[align omitted — 208 chars of source]

We make the following assumptions. For the remainder of the paper all limits are taken as $T \to \infty$, unless stated otherwise.

assumption$ $ \begin{enumerate}[label=(A\arabic*)] • The sequence $\big( \varepsilon_{t}, \mathbf{h}_t \big)$ for $t \in \mathbb{N}$ is $\alpha$-mixing of size $-r/(r-1)$ for some $r > 1$. • It holds that $\mathbb{E} \left[ ||\mathbf{h}_t||^{2r+\delta} \right] < \infty$ for all $t \in \mathbb{N}$ for some $\delta > 0$. • The matrix $\mathbb{E} \left[\mathbf{h}_t \mathbf{h}_t^\top \right]$ has full rank for all $t \in \mathbb{N}$. • The limit, $\Omega_{\mathrm{Mode}}$, of $\Omega_{T,\mathrm{Mode}} := \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \mathbf{h}_t \mathbf{h}_t^\top f_{t}(0) \right] \int K'(u)^2 \mathrm{d}u$ is positive definite. • For all $t \in \mathbb{N}$, the conditional distribution of $\varepsilon_t = X_t - Y_{t+1}$ given $\mathcal{F}_t$ is absolutely continuous with density $f_{t}(\cdot)$ which is three times continuously differentiable with bounded derivatives. • $K: \mathbb{R} \to \mathbb{R}$, $u \mapsto K(u)$ is a non-negative and continuously differentiable kernel function such that: (i) $\int K(u) \mathrm{d}u = 1$, (ii) $\int u K(u) \mathrm{d}u = 0$, (iii) $\sup K(u) \le c <\infty$, (iv) $\sup K'(u) \le c <\infty$, (v) $\int u^2 K(u) \mathrm{d}u <\infty$, (vi) $\int \big| K'(u) \big| \mathrm{d}u <\infty$, (vii) $\int u K' (u) \mathrm{d}u < \infty$. • $\delta_T$ is a strictly positive and deterministic sequence such that (i) $T \delta_T \to \infty $, and (ii) $T \delta_T^7 \to 0$. \end{enumerate}

The above assumptions are a combination of standard assumptions from rationality testing and nonparametric statistics. {\color{black}Importantly, for our applications, these assumptions allow for heterogeneous predictive distributions, $F_t$.} Conditions (ref) and (ref) facilitate the use of a law of large numbers and of a central limit theorem for martingale difference arrays (MDA) (that generalize MD sequences to triangular arrays; see davidson1994stochastic) and allows for both possibly non-stationary time-series and cross-sectional applications. Notice that the rate in the mixing condition (ref) is relatively weak as we only need it for a law of large numbers while we apply a central limit theorem for MDA in the proof of Theorem (ref) below. In cross-sectional applications with independent observations, this assumption can be replaced (and weakened) by the classical Lindeberg condition (see e.g.\ White2001, Section 5.2). Notice that as the kernel function $K'$ is bounded, we do not require existence of any moments of $Y_t$ or $X_t$, which makes this more flexible than rationality testing for mean forecasts. The full rank condition (ref) prevents the instruments from being perfectly colinear which in turn prevents the asymptotic covariance matrix from being singular. Condition (ref) guarantees that the asymptotic covariance matrix is well behaved for non-stationary data. Assumption (ref) assumes a relatively smooth behavior of the conditional density function which is required to apply a Taylor expansion common to the nonparametric literature. Conditions (ref) and (ref) are standard kernel and bandwidth conditions from the nonparametric literature. We discuss specific kernel and bandwidth choices in Supplemental Appendices (ref) and (ref) respectively.

theoremUnder Assumption (ref) and $\mathbb{H}_0: X_t = \operatorname*{Mode}( Y_{t+1} | \mathcal{F}_t ) ~ \forall~t$ a.s., it holds that \begin{align} \delta_T^{3/2} T^{-1/2} \sum_{t=1}^T \psi(Y_{t+1},X_t,\mathbf{h}_t, \delta_T) \overset{d}{\to} \mathcal{N} \big( 0, \Omega_{\mathrm{Mode}} \big), \end{align} where $\Omega_{\mathrm{Mode}}$ is the limit of $\Omega_{T,\mathrm{Mode}} = \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \mathbf{h}_t \mathbf{h}_t^\top f_{t}(0) \right] \int K'(u)^2 \mathrm{d}u$ as $T \to \infty$.

We obtain a test for the rationality of mode forecasts by drawing on the literature on nonparametric estimation. Unsurprisingly, therefore, the rate of convergence is slower than $\sqrt{T}$; Assumption (ref) requires $\delta_T \propto T^{-\kappa}$ with $\kappa \in (1/7, 1)$, which implies that the fastest convergence rate approaches $T^{2/7}$ and is obtained by setting $\delta_T \approx T^{-1/7}$.\footnote{Under additional assumptions, the speed of convergence of a nonparametric estimator may be increased via the use of higher-order kernel functions, see e.g.\ LiRacine2006. However, as the generalized modal midpoint introduced in Definition (ref) requires a log-concave kernel to be well-defined and unique (see Theorem (ref)), and as this assumption is automatically violated for higher-order kernels, we do not consider them here.}

Theorem (ref) guarantees the strictness of the identification function $V_\delta(x,Y)$ only if $K$ has infinite support and is strictly increasing (decreasing) left (right) of its mode. These conditions are satisfied by the familiar Gaussian kernel, which we adopt for our analysis.\footnote{We also considered the relatively efficient quartic (or biweight) kernel but did not observe a change in power relative to the Gaussian kernel. See Supplemental Appendix (ref) for further discussion of the kernel choice and simulation results.}

Following Kemp2012 and Kemp2019, we estimate the covariance matrix by its sample counterpart,

align[align omitted — 171 chars of source]

The following theorem shows consistency of the asymptotic covariance estimator without imposing our null hypothesis.

theoremGiven that $X_t \in \mathcal{F}_t$, $\forall t \in \mathbb{N}$ and Assumption (ref), it holds that $\widehat{\Omega}_{T,\mathrm{Mode}} - \Omega_{T,\textrm{Mode}} \overset{P}{\to} 0$.

We can now define the Wald test statistic:

align[align omitted — 276 chars of source]

The following statement follows directly from Theorem (ref) and Theorem (ref).

corollaryUnder Assumption (ref) and the null hypothesis $\mathbb{H}_0: X_t = \operatorname*{Mode}( Y_{t+1} | \mathcal{F}_t ) ~ \forall ~ t$ a.s., it holds that $J_T \overset{d}{\to} \chi^2_k$.

This corollary justifies an asymptotic test at level $\alpha \in (0,1)$ which rejects $\mathbb{H}_0$ when $J_T > Q_k(1-\alpha)$, where $Q_k(1-\alpha)$ denotes the $(1-\alpha)$ quantile of the $\chi^2_{k}$ distribution.\footnote{Our focus on one-step ahead forecasts allows for the application of a central limit theorem for MDAs, which only requires the existence of second moments. Multi-step ahead forecasts, on the other hand, are usually handled via CLTs for processes with more memory (e.g., mixing processes) at a cost of imposing stronger moment conditions. We leave this extension for future research.} Note that the bandwidth parameter, $\delta_T$, is introduced only to conduct the test of mode forecast rationality; the forecast itself, $X_t$, is, under the null, the true conditional mode of the target variable, not a (smoothed) modal midpoint.

We now turn to the behavior of our test statistic $J_T$ under a local alternative hypothesis,

align[align omitted — 152 chars of source]

{\color{black}for some forecast $X_t \in \mathcal{F}_t$,} some constant $c \in \mathbb{R}^k$ and some possibly stochastic sequence $a_T$ to be specified in the following Theorem (ref), which characterizes the behavior of our mode rationality test under the local alternative.

theorem{\color{black}Assume that $T \delta_T^3 \to \infty$, Assumption (ref) and the alternative $\mathbb{H}_{A, \text{loc}}$ in (ref) hold for the forecasts $X_t \in \mathcal{F}_t$. Then:} \begin{enumerate}[label=(\alph*)] • If $a_T T^{1/2} \delta_T^{3/2} \overset{P}{\to} 1$, then $J_T \overset{d}{\to} \chi^2_k \big( c^\top \Omega_{\textrm{Mode}}^{-1} c \big)$, where $\chi^2_k \big( \tilde c \big)$ denotes a $\chi^2_k$ distribution with non-centrality parameter $\tilde c \in \mathbb{R}$. • If $a_T T^{1/2} \delta_T^{3/2} \overset{P}{\to} 0$, then $J_T \overset{d}{\to} \chi^2_k$. • If $\big( a_T T^{1/2} \delta_T^{3/2} \big)^{-1} \overset{P}{\to} 0$, then $\mathbb{P} \left( J_T \ge \bar c \right) \to 1$ for any $\bar c > 0$, i.e., we have uniform power. \end{enumerate}

This theorem shows that the power of our mode rationality test is essentially driven by the condition that $\frac{1}{T} \sum_{t=1}^T f'_t(0) \mathbf{h}_t$ is non-zero, which implies that $X_t$ cannot be the mode of $F_t$. It essentially means that the instruments $\mathbf{h}_t$ are correlated with the conditional density slope at zero, $f'_t(0)$. This is analogous to the case for standard mean and median rationality tests which have power against the alternative that $\mathbb{E} \big[V(X_t, Y_{t+1}) \mathbf{h}_t] \not= 0$, where $V(X_t, Y_{t+1})$ is the identification function for the mean or median.\footnote{See e.g.\ Theorem 2 of giacomini2006tests, where their Comment 6 and Nolde2017 point out that the theory can be directly adapted to rationality testing.} When the instruments include a constant, a sufficient condition for a global alternative hypothesis, where $a_T$ is constant, is $f'_t(0) \ge c > 0$ ($f'_t(0) \le -c < 0$) for all $t \in \mathbb{N}$, i.e.\ when all forecasts are issued to the left (right) of the mode. This can be interpreted as uniform power against the class of strongly unimodal distributions. Our test will have low power to reject forecasts that are far in the tail, where the density is (almost) flat, however our test is intended for forecasts of a measure of central tendency, and such forecasts will lie broadly in the central region of the distribution where this concern does not arise.

Part (ref) of Theorem (ref) shows that if $\frac{1}{T} \sum_{t=1}^T f'_t(0) \mathbf{h}_t$ converges at rate $T^{-1/2} \delta_T^{-3/2}$, then the asymptotic distribution of $J_T$ stabilizes as a non-central $\chi^2$-distribution, implying that for a fixed alternative, our test statistic diverges at rate $T^{1/2} \delta_T^{3/2}$, which is approximately $T^{2/7}$ in practice. In contrast, it is easy to show that rationality tests for identifiable functionals as the mean and median have local power against alternatives that converge with rate $T^{-1/2}$, which is of course faster than the rate $T^{-1/2} \delta_T^{-3/2}$ of the mode test . We demonstrate in the simulations and the applications that our mode test nevertheless has satisfactory power in practice in typically encountered sample sizes.

Part (ref) of Theorem (ref) further shows that if $\frac{1}{T} \sum_{t=1}^T f'_t(0) \mathbf{h}_t$ converges to zero faster than $T^{-1/2} \delta_T^{-3/2}$, then our test behaves as under the null and has no power. Finally, part (ref) implies that our test has unit power asymptotically when $\frac{1}{T} \sum_{t=1}^T f'_t(0) \mathbf{h}_t$ converges to zero slower than $T^{-1/2} \delta_T^{-3/2}$, which nests the classical case of a fixed, global alternative.

{\color{black} Following Kemp2012 and Kemp2019, we choose $\delta_T$ proportional to $T^{-0.143}$, which is almost $T^{-1/7}$ as required by Assumption (ref). Specifically, we set $\delta_T = k_1 \cdot k_2 \cdot T^{-0.143}$ where $k_1 = 2.4\widehat{\operatorname{Med}} \big( \big| \varepsilon_{1:T} - \widehat{\operatorname{Med}} [ \varepsilon_{1:T} ] \big| \big)$, $k_2 = \exp(-3 \left| \hat \gamma \right|)$, and $\hat \gamma = 3 \, \big( \frac{1}{T} \sum_t \varepsilon_t - \widehat{\operatorname{Med}} [\varepsilon_{1:T}] \big)/\widehat{\sigma}(\varepsilon_{1:T})$, where $\widehat{\sigma}$ and $\widehat{\operatorname{Med}}$ denote the sample standard deviation and median. As in Kemp2012 and Kemp2019, we choose $k_1$ proportional to the median absolute deviation of the forecast error, a robust measure of variation, and introduce a second constant, $k_2$, to adjust for the skewness of the forecast error, measured by the absolute value of Pearson's second skewness coefficient, $\hat \gamma$. We illustrate the good finite sample properties of our bandwidth choice through simulations in Supplemental Appendix (ref).}

Evaluating Forecasts of Measures of Central Tendency

We define a class of measures of central tendency nesting the mean, median, and mode, and we propose tests of whether a forecast is rational with respect to any element of the class. Formally, we consider convex combinations of loss functions pertaining to the mean, median, and mode: $L_{\mathrm{Mean}}(x,y) = (x-y)^2$, $L_{\mathrm{Med}}(x,y) = |x-y|$, $L_{\mathrm{Mode}, \delta}(x,y) = -\delta^{1/2} K\left( (x-y)/\delta \right)$, for some kernel $K$.\footnote{Note that $L_{\mathrm{Mode}, \delta}$ differs from $L^K_{\delta}$ in equation ((ref)) by a scaling factor of $\delta^{3/2}$. This arises from the convergence rate presented in Theorem (ref).} Each vector in the unit simplex, $\Theta := \{ \theta \in \mathbb{R}^3: ||\theta||_1 = 1, \theta \ge 0 \}$, {\color{black}together with a positive weight vector $\big(w_{\mathrm{Mean}}, w_{\mathrm{Med}}, w_{\mathrm{Mode}}\big) \in \mathbb{R}_+^3$,} generates an optimal forecast

align[align omitted — 363 chars of source]

where we minimize over {\color{black}all $\mathcal{F}_t$-measurable potential forecasts $X$.} {\color{black}The scalar weights $w_{\mathrm{Mean}}$, $w_{\mathrm{Med}}$, and $w_{\mathrm{Mode}}$ can be used to equalize the influence of the different loss functions, discussed further below.} At the vertices of $\Theta$, this nests the mean, median, and generalized modal midpoint, the latter tending to the mode as $\delta \to 0$; see Theorem (ref).

Assuming that a forecaster generates her forecasts $X_t = X_t^\ast(\theta_0)$ according to (ref), we aim to determine the set of values for $\theta_0$ for which the given forecasts are optimal. Tests of forecast optimality are commonly conducted via an unconditional moment condition obtained by interacting an $\mathcal{F}_t$-measurable $(k \times 1)$ vector of instruments $\mathbf{h}_t$ with the first order condition of the optimal forecast EKT2005, Nolde2017. The latter also arise in the rationality tests introduced in equations (ref) and (ref). We continue in this tradition and test rationality across our class of central tendency measures by examining the variable $\phi_{t,T}(\theta)$, defined as:

equation[equation omitted — 338 chars of source]

The identification functions $V_{\mathrm{Mean}}$ and $V_{\mathrm{Med}}$ are given in equations (ref) and (ref), and $V_{\mathrm{Mode},\delta}( x,y) = - \delta^{-1/2} K' \left( (x-y)/\delta \right)$ is an appropriately scaled version of (ref) that guarantees asymptotic normality in Theorem (ref) and Theorem (ref) below. {\color{black}Under forecast rationality, there must exist a $\theta_0$ (possibly set-valued) such that $\phi_{t,T}(\theta_0)$ is “small” on average. We formalize this statement in Assumption (ref) below.} Notice that $\phi_{t,T}$ is a triangular array, depending on $t$ and $T$, through its dependence on the bandwidth $\delta_T$.

remark{\color{black}The forecast $X_t^*(\theta)$ from equation ((ref)) is optimal with respect to a convex combination of loss functions.} Intuitively one might be inclined to consider a convex combination of the functional values rather than of the associated loss functions, however such functionals are generally neither elicitable nor identifiable (see Proposition (ref) in the Supplemental Appendix), rendering infeasible an extension of the rationality tests introduced in Section (ref) to this case. However it can be shown (Proposition (ref)) that any forecast following equation (ref) lies between the mean, median and generalized modal midpoint, though with combination weights for loss functions that generally differ from the combination weights for the functional values.
remark{\color{black}It is possible to consider a probabilistic mixture of mean, median, and mode type forecasts, discussed further in Supplemental Appendix (ref). Such forecasts are consistent with the forecast rationality moment condition we test if only a constant is used as instrument (Proposition (ref)), but will generally be rejected for multivariate instruments. }

In {\color{black}equations (ref) and (ref)}, we allow for normalizations of each {\color{black}loss and} identification function using the scalar weights $w_{\mathrm{Mean}}$, $w_{\mathrm{Med}}$ and $w_{\mathrm{Mode}}$. This allows one to adjust the importance of each loss function in order to construct tests that are robust to linear data transformations. To do so, we use the inverse of the standard deviations of the respective identification functions in our empirical analysis, though one could instead use equal weights, or some other choice. We define $\widehat \phi_{t,T}(\theta)$ as in equation (ref), but using sample-dependent weights $\widehat{w}_{T,\mathrm{Mean}}$, $\widehat{w}_{T,\mathrm{Med}}$ and $\widehat{w}_{T,\mathrm{Mode}}$.

Consider the GMM objective function based on $\widehat \phi_{t,T}(\theta)$:

align[align omitted — 232 chars of source]

where $\widehat{\Sigma}_T^{-1}(\theta)$ denotes an $O_P(1)$ positive definite weighting matrix, which may depend on the parameter $ \theta$. Unlike the problem in EKT2005, the unknown parameter in our framework (the weight vector $\theta$) cannot be assumed to be well identified. For example, for symmetric distributions, the combination weights are completely unidentified. For distributions that exhibit only mild asymmetry a weak identification problem arises. For asymmetric distributions where one measure is a convex combination of the other two (a situation that arises naturally in location-scale processes) we have partial identification of the weight vector. The distribution of economic variables may or may not exhibit asymmetry, and so addressing this identification problem is a first-order concern.

The possibility that the true parameter $\theta_0$ is unidentified, partially identified, or weakly identified implies that the objective function $S_T(\theta)$ may be flat or almost flat in a neighborhood of $\theta_0$, ruling out consistent estimation of $\theta_0$. StockWright2000 show that, under regularity conditions, we can nevertheless construct asymptotically valid confidence bounds for $\theta_0$, by showing that the objective function $S_T$ evaluated at $\theta_0$ continues to exhibit an asymptotic $\chi^2$ distribution. This facilitates the construction of asymptotically valid confidence bounds even in a setting where the parameter vector may be strongly identified, weakly identified, or unidentified.\footnote{Alternative approaches to estimate the confidence sets under partial identification include Kleibergen2005, CHT07, BM08 and Chen2018 among others.} We further impose the following regularity conditions on our process.

assumption(A) $\mathbb{E} \left[ |\varepsilon_t|^{2+\delta} \right] < \infty$ and $\mathbb{E} \left[ ||\mathbf{h}_t||^{2+\delta} |\varepsilon_t|^{2+\delta} \right] < \infty$, (B) $\widehat{w}_{T,\mathrm{Mean}} \overset{P}{\to} w_{\mathrm{Mean}}$, $\widehat{w}_{T,\mathrm{Med}} \overset{P}{\to} w_{\mathrm{Med}}$, and $\widehat{w}_{T,\mathrm{Mode}} \overset{P}{\to} w_{\mathrm{Mode}}$ for some positive weights $w_{\mathrm{Mean}}$, $ w_{\mathrm{Med}}$ and $w_{\mathrm{Mode}}$; (C) The limit of $\Sigma_T(\theta_0)$ defined in (ref) is positive definite. (D) {\color{black}For the forecasts $X_t \in \mathcal{F}_t$, $t\in \mathbb{N}$,} there exists a {\color{black}(not necessarily unique)} $\theta_0 \in \Theta$ and triangular arrays $\phi_{t,T}^\ast(\theta_0)$ and $u_{t,T}(\theta_0)$, such that \begin{equation} \phi_{t,T}(\theta_0) = \phi_{t,T}^\ast(\theta_0) + u_{t,T}(\theta_0), \qquad where \end{equation} \begin{enumerate}[label=(\alph*), leftmargin=1.0cm] • $\big\{ T^{-1/2} \phi_{t,T}^\ast(\theta_0), \mathcal{F}_{t+1} \big\}$ is a martingale difference array, • $T^{-1} \sum_{t=1}^T || u_{t,T}(\theta_0) ||^2 \overset{P}{\to} 0$, and $\sum_{t=1}^T \mathbb{E} \left[ || T^{-1/2} u_{t,T}(\theta_0)||^{2+\delta} \right] \to 0$, • $T^{-1} \sum_{t=1}^T u_{t,T}(\theta_0) \phi_{t,T} (\theta_0)^\top \overset{P}{\to} 0$ and $T^{-1} \sum_{t=1}^T \mathbb{E} \left[ u_{t,T}(\theta_0) \phi_{t,T} (\theta_0)^\top \right] \to 0$. \end{enumerate}

{\color{black}Conditions (A), (B), and (C) are standard regularity conditions, and Assumption (D) represents the null hypothesis of forecast rationality that we test. Importantly, these conditions allow for heterogeneous predictive distributions, $F_t$. } {\color{black}The scaling with $T^{-1/2}$ in Assumption (ref) (D)(a) matches the normalization in the CLT we employ (davidson1994stochastic).} The decomposition in equation ((ref)) implies that the array $T^{-1/2} \phi_{t,T}(\theta_0)$ is an approximate MDA in the sense that $T^{-1/2} \phi_{t,T}(\theta_0)$ can be decomposed into a MDA $T^{-1/2} \phi^\ast_{t,T}(\theta_0)$ and some asymptotically vanishing array $T^{-1/2} u_{t,T}(\theta_0)$. This decomposition is required as the two standard assumptions---mixing and exact MDA conditions---are too restrictive for our application. First, the assumption that $\big\{ T^{-1/2} \phi_t(\theta_0), \mathcal{F}_{t+1} \big\}$ is a MDA does not hold for the baseline case that $X_t$ is an optimal mode forecast (see the proof of Theorem (ref) for details). Second, imposing mixing conditions is too weak for our case as CLTs for mixing processes generally require finite moments of some order $r>2$, which is not fulfilled for the mode case as these moments diverge through the bandwidth parameter $\delta_T$.

The intermediate case of Assumption (ref) (D) allows to apply a CLT based on finite second moments only, and can easily be shown to hold at the three vertices, where the forecast is the mean, median or mode. Specifically, when $X_t$ is a mean or median forecast (i.e.\ $\theta_0 = (1,0,0)$ or $\theta_0 = (0,1,0)$), set $u_{t,T}(\theta_0) = 0$ and $\big\{ T^{-1/2} \phi_{t,T}(\theta_0), \mathcal{F}_{t+1} \big\}$ is obviously a MDA. When $X_t$ is the true conditional mode of $Y_{t+1}$, (i.e.\ $\theta_0 = (0,0,1)$), set

align[align omitted — 235 chars of source]

Then, the conditions on $u_{t,T}(\theta_0)$ in Assumption (ref) (D) are fulfilled as shown in Lemma (ref) in the Supplemental Appendix.

When $X_t$ is a convex combination of a mean and median forecast, i.e., $\theta_0 = (\xi,1-\xi,0)$ for some $\xi \in [0,1]$, we set $u_{t,T}(\theta_0) = 0$ and $\big\{T^{-1/2} \phi_{t,T}(\theta_0), \mathcal{F}_{t+1} \big\}$ is again a MDA. When $X_t$ is a convex combination with non-zero weight on the mode, Assumption (ref) (D) is difficult to verify.

Theorem (ref) below presents the asymptotic distribution of the process $T^{-1/2} \sum_{t=1}^T \widehat \phi_{t,T}(\theta_0)$ at the true parameter $\theta_0$.

theorem{\color{black}Under Assumptions (ref) and (ref),} it holds that \begin{align} T^{-1/2} \sum_{t=1}^T \widehat \phi_{t,T}(\theta_0) \overset{d}{\to} \mathcal{N} \big( 0, \Sigma(\theta_0) \big), \end{align} where $\Sigma(\theta_0)$ is the limit as $T \to \infty$ of \begin{align} \begin{aligned} \Sigma_T(\theta_0) &:= \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{10}^2 w_{\mathrm{Mean}} \mathbf{h}_t \mathbf{h}_t^\top w_{\mathrm{Mean}} \varepsilon_t^2 + \theta_{20}^2 w_{\mathrm{Med}} \mathbf{h}_t \mathbf{h}_t^\top w_{\mathrm{Med}} \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big)^2 \right. \\ &\qquad \qquad \qquad + \theta_{30}^2 w_{\mathrm{Mode}} \mathbf{h}_t \mathbf{h}_t^\top w_{\mathrm{Mode}} f_t(0) \int K'(u)^2 \mathrm{d}u \\ &\qquad \qquad \qquad \left. + 2 \theta_{10} \theta_{20} w_{\mathrm{Mean}} \mathbf{h}_t \mathbf{h}_t^\top w_{\mathrm{Med}} \varepsilon_t \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big) \right]. \end{aligned} \end{align}

Under the null hypothesis, Assumption (ref) (D) implies that $T^{-1/2} \phi_{t,T}(\theta_0)$ (and hence also $T^{-1/2} \widehat \phi_{t,T}(\theta_0) $) is an approximate MDA, i.e.\ this array is approximately (as $T \to \infty$) uncorrelated. Consequently, we do not need to rely on HAC covariance estimation, and can instead estimate the asymptotic covariance matrix using the simple {\color{black}sample covariance matrix $\widehat{\Sigma}_T(\theta) = \frac{1}{T} \sum_{t=1}^T \widehat \phi_{t,T}(\theta) \widehat \phi_{t,T}(\theta)^\top$.} The next theorem shows consistency of the outer product covariance estimator.

theoremGiven Assumptions (ref) and (ref), it holds that $\widehat{\Sigma}_T(\theta_0) - \Sigma_T(\theta_0) \overset{P}{\to} 0$.
corollaryGiven Assumptions (ref) and (ref), it holds that $S_T(\theta_0) \overset{d}{\to} \chi^2_k$.

Following StockWright2000, this corollary allows one to construct asymptotically valid confidence regions for $\theta_0$ with coverage probability $(1-\alpha)\%$ by considering the set

align[align omitted — 117 chars of source]

where $Q_k(1-\alpha)$ denotes the $(1-\alpha)$ quantile of the $\chi^2_{k}$ distribution.

Given the above results, we obtain a test for forecast rationality for a general measure of central tendency by evaluating the GMM objective function using a dense grid of convex combination parameters $\theta_j \in \Theta$ for $j = 1,\dots,J$. An asymptotically valid confidence set is given by the values of $\theta_j$ for which $S_T(\theta_j) \le Q_k(1-\alpha)$. These values represent the centrality measures that “rationalize” the observed sequence of forecasts and realizations, in that rationality cannot be rejected for these measures of centrality. It is possible that the confidence set is empty, in which case we reject rationality at the $\alpha$ significance level for the entire class of general centrality measures.

The power of the rationality test depends on the instrument choice. This property is shared with many other tests in the literature EKT2005, PattonTimmermann2007,schmidt2021interpretation. Good instruments span the information set of the forecaster and are not strongly correlated with each other. Additional instruments generally improve power asymptotically (or shrink the identified set, in the partial identification case), but can deteriorate power in finite samples. An informative instrument, and one used as far back as MincerZarnowitz1969, is the forecast $X_t$ itself. In our simulations and applications we found the forecast to be a powerful instrument. Additionally, it is guaranteed to be in the information set of the forecaster, and so is a valid instrument.

The forecast evaluation problem and approach considered here is related to, but distinct from, EKT2005. These authors consider the case that a respondent's point forecast corresponds to some quantile (or expectile) of her predictive distribution. They employ a parametric loss function (“lin-lin” for quantiles, “quad-quad” for expectiles), $L(X,Y;\tau)$ with a scalar unknown parameter ($\tau$) characterizing the asymmetry of the loss. EKT2005 use GMM to estimate the $\tau$ that best describes the sequence of forecasts and realizations, and test whether forecast rationality holds at the estimated value for $\tau$. Economically, our approach differs from EKT2005 in that we consider forecasts only as measures of centrality, allowing for a wide range of such measures, while that paper considers only a single centrality measure nested within a wide range of asymmetric forecasts. Statistically, our approach differs as we are forced to address the feature that our parameter may be partially, weakly, or un-identified, which precludes point estimation.

We use convex combinations of identification functions to parametrically nest the mean, the median, and (asymptotically) the mode. Alternative approaches are possible, e.g., via the $L^p$ norm. Our convex combination approach has the advantage of separating the parameter of interest $\theta$ from the bandwidth parameter $\delta_T$, and further does not suffer from technical difficulties such as bandwidth parameters in the exponent, or identification functions with singularities at the mode.

{\color{black}Simulation Study}

Rationality tests for mode forecasts

To evaluate the finite-sample properties of our mode rationality test, we simulate data from a cross-sectional and a time series AR(1)-GARCH(1,1) data generating process (DGP), given by

align[align omitted — 176 chars of source]

where $\mathcal{SN}(0,1,\gamma)$ is a skewed standard Normal distribution, $Z_t$ denotes a vector of covariates (possibly including lagged values of $Y_{t+1}$), $\zeta$ denotes a parameter vector and $\sigma_{t+1}$ represents a conditional variance process. Using the general formulation in (ref), the cases we consider are:\footnote{In the Supplemental Appendix, we also present data for a heteroskedastic cross-sectional DGP that is as in case (1), but with $\sigma_{t+1} = 0.5 + 1.5(t + 1)/T$, and a homoskedastic AR(1) process with $Z_t = Y_{t}$, $\zeta = 0.5$ and $\sigma_{t+1} = 1$.}

eqnarray[eqnarray omitted — 368 chars of source]

where $\boldsymbol{\iota}$ is a vector of ones. These DGPs are based on a skewed Gaussian residual distribution with skewness parameter $\gamma$. This choice nests the case of a standard Gaussian distribution at $\gamma = 0$, and in this case all measures of centrality coincide. As the skewness parameter grows in magnitude, the measures of centrality differ increasingly.

For the DGP in ((ref)), optimal mode forecasts are given by $X_t = {\operatorname*{Mode}}(Y_{t+1}| \mathcal{F}_t) = Z_{t}^\top \zeta + \sigma_{t+1} \operatorname*{Mode} ( \xi_t )$, where $\operatorname*{Mode} ( \xi_t )$ depends on the skewness parameter $\gamma$. We use a range of values $\gamma \in \{ 0, 0.1, 0.25, 0.5\}$ and sample sizes $T \in \{ 100, 500, 2000, 5000 \}$. To evaluate the size of our test in finite samples, we generate optimal mode forecasts and apply the mode forecast rationality test based on the instrument choices $\mathbf{h}_t = 1$ and $\mathbf{h}_t = (1,X_t)$. All simulation results are obtained with $2,000$ simulation replications.

table[table omitted — 1,636 chars of source]

Table (ref) presents the finite-sample sizes of the test under the different DGPs, sample sizes, instrument choices and skewness parameters. In all cases we use a Gaussian kernel and set the nominal size to $5\%$.\footnote{In the Supplemental Appendix we present similar results for different kernel choices and significance levels.} We find that our mode rationality test leads to finite-sample rejection rates that are generally close to the nominal test size, across all of the different choices of DGPs, instruments, sample sizes and skewness parameters. Table (ref) reveals that an increasing degree of skewness in the underlying conditional distribution negatively influences the test's performance. This can be explained by the increasing difference between the mode and the modal midpoint (which is elicited by the corresponding identification function in finite samples) for an increasing skewness in the data. As a consequence, we choose a smaller bandwidth following the rule of thumb described at the end of Section (ref), resulting in less efficient estimates. Consequently, for highly skewed distributions, the mode rationality test requires larger sample sizes in order to converge to the nominal test size.

figure[figure omitted — 530 chars of source]

To analyze the power of the mode forecast rationality test we use the DGPs from equations (ref)-(ref) and introduce a bias to the forecasts $\tilde X_t = X_t + \kappa \varsigma$, where $\varsigma = \sqrt{\operatorname{Var}(\varepsilon_{1:T})}$ and $\kappa \in (-0.5, 0.5)$.\footnote{We also consider a misspecification where we add noise as $\tilde X_t = X_t + \mathcal{N}(0, \kappa \varsigma^2)$ in Supplemental Appendix (ref).} This type of misspecification introduces a deterministic bias, where the degree of misspecification depends on the parameter $\kappa$.

(ref) presents power plots for these “biased” forecasts for a range of sample sizes, skewness parameters and the two DGPs in equations ((ref))-((ref)), where we plot the rejection rate against the degree of misspecification $\kappa$. For all plots, we use the instrument choice $\mathbf{h}_{t} = (1,X_t)$, a Gaussian kernel, and a nominal level of $5\%$. Notice that for $\kappa = 0$ the figures show the empirical test size.

Figure (ref) reveals that the proposed mode rationality test exhibits, as expected, increasing power for an increasing degree of misspecification. Also as expected, larger sample sizes lead to tests with greater power, although even the two smaller sample sizes exhibit reasonable power. The figure also reveals that an increasing degree of skewness yields to a slight loss of power (given a fixed degree of misspecification). This is driven through the bandwidth choice, where larger values of the (empirical) skewness result in a smaller bandwidth, and consequently a lower test power (analogous to the bias-variance trade-off in the nonparametric estimation literature). We illustrate the good behavior of our finite sample bandwidth choice through further simulations in the Supplemental Appendix (ref).

Overall, our simulation results show that even though the mode rationality test converges at a slower-than-parametric rate (approximately $T^{2/7}$), it nevertheless has considerable power for the empirically relevant sample sizes of $T \ge 500$; as e.g., in our application in Section (ref).

Rationality tests for an unknown measure of central tendency

We now examine the small sample behavior of the asymptotic confidence sets for the measures of central tendency, described in (ref). As in the previous section, we consider the two DGPs described in and after equation (ref) and the same varying sample sizes $T$, skewness parameters $\gamma$, and instruments $\mathbf{h}_t$. We generate optimal one-step ahead forecasts for the mean, median and mode as the true conditional mean, median and mode of $Y_{t+1}$ given $\mathcal{F}_t$:

align[align omitted — 117 chars of source]

for $\mathsf{M} \in \{\textrm{Mean}, \textrm{Median}, \textrm{Mode}\}$, corresponding to the three vertices of $\Theta$.

For any other $\theta \in \Theta$, we generate the forecast $X_t^*(\theta)$ that arises as the optimal forecast under a convex combination of loss functions as in (ref) through a numerical optimization procedure. For the weights in (ref), we use the reciprocal of the estimated standard deviations of the respective loss functions over all $t=1,\dots,T$. In detail, for a given $\theta \in \Theta$, and for each $t =1,\dots,T$, we draw 5000 replications from the residuals $\xi_t$ in (ref) and obtain $X_t^*(\theta)$ in (ref) through a numerical optimization, where we approximate the expected value in (ref) by the empirical average over the 5000 replications. We repeat this process three times to stabilize the effect of the estimated weights through the inverted standard deviations, akin to the iterated generalized method of moments.

Table (ref) reports the coverage rates of the correct parameter $\theta_0$ in the $90\%$-confidence sets from (ref), where we simulate the forecasts following (ref) and the previous paragraph, where the true $\theta_0$ is given in the second table column. We find accurate inclusion rates for both DGPs and both, symmetric as well as skewed data. Only the true mode forecasts (compare to Section (ref)) exhibit some undercoverage for skewed data for the smallest considered sample size.

table[table omitted — 2,180 chars of source]
figure[figure omitted — 679 chars of source]

Figure (ref) shows the coverage rates of the confidence sets at an equally spaced grid of values in $\Theta$. The panels in the figure correspond to six true combination weights, which are in each case highlighted in red and are explained in the first two columns of Table (ref). Each point in the triangles corresponds to a tested centrality measure, i.e.\ to one value of $\theta \in \Theta$, and for each point we display as a percentage number and a color code how often it is contained in the $90\%$ confidence set. We use the instruments $\mathbf{h}_{t} = (1,X_t)$, and the iid DGP in equation ((ref)).

In each of the six subfigures the values highlighted in red confirm the correct coverage rates of the respective true forecasts from Table (ref). We further find clearly decreasing coverage rates around the truth illustrating the power our test exhibits. Notably, there exist central tendency forecasts, e.g., mean-mode loss combinations in our simulation, that would typically be rejected by standard tests of mean, median, or mode rationality. The confidence sets are stretching in the upper left direction, which illustrates the partial identification in the simulated process: As the median is between the mean and mode for the skewed normal distribution, it can also be identified by a convex combination of the mean and mode as can be seen in the upper right panel.\footnote{When the DGP is conditionally location-scale, one of the three centrality measures can always be expressed as a (constant) convex combination of the other two. In this DGP, this is the median, but in other applications it may be any of the three measures. When the conditional distribution exhibits variation in higher-order moments or other “shape” parameters, this restriction will generally not hold, and the variation may or may not be sufficient to separately identify the three centrality measures.}

In Supplemental Appendix (ref), we show plots similar to Figure (ref), but in each case we exchange one feature: First, when using symmetric data, all points are naturally identified in all panels. Second, the results are qualitatively unchanged when using the AR-GARCH DGP in equation ((ref)). Third, the power is lower for the smaller sample size of $T=500$. Finally, Figure (ref) illustrates that our approach identifies probabilistic mixtures of mean, median, and mode forecasts for instruments $\mathbf{h}_t = 1$, but generally rejects rationality when using $\mathbf{h}_t = (1,X_t)$.

Evaluating Survey Forecasts of Individual Income

We apply our proposed tests to survey responses to the Federal Reserve Bank of New York's Survey of Consumer Expectations.\footnote{Source: Survey of Consumer Expectations, © 2013-2019 Federal Reserve Bank of New York (FRBNY). The SCE data are available without charge at \url{http://www.newyorkfed.org/microeconomics/SCE} and may be used subject to license terms posted there. FRBNY disclaims any responsibility or legal liability for this analysis and interpretation of Survey of Consumer Expectations data.}$^{,}$\footnote{In Supplemental Appendix (ref) we consider two other empirical applications. The first is to the “Greenbook” forecasts of US GDP growth produced by economists at the Board of Governors of the Federal Reserve. In that application we find that the Greenbook forecasts are rational mean forecasts, but not mode or median forecasts. Our second application is to random walk forecasts of exchange rates, in the spirit of MeeseRogoff1983. We find random walk forecasts to be rational mean forecasts for three different exchange rates, while the mode and median are rejected for two of the three exchange rates.} We focus on the Labor Market Survey component, which is conducted each March, July, and November, and which asks participants a variety of questions, including about their current earnings and their beliefs about their earnings in four months (i.e., the date of the next survey). Using adjacent surveys over the period March 2015 to March 2018, we obtain a sample of 3,916 pairs of forecasts ($X_t$) and realizations ($Y_{t+1}$).\footnote{We drop observations that include forecasts or realizations of annualized income below \$1,000 or above \$1 million, which represent less than 1% of the initial sample. We also drop observations where the ratio of the realization to the forecast, or its inverse, is between 9 and 13, to avoid our results being affected by misplaced decimal points or by the failure to report annualized income (leading to proportional errors of around 10 to 12 respectively).}$^,$\footnote{Most respondents in our sample appear just once: the 3,916 forecast-realization pairs come from 2,628 unique individuals. Our econometric approach does not require repeated samples and so this causes no difficulty, and our results are qualitatively unchanged if we use only the first forecast-realization pair from each respondent.} In testing the rationality of these forecasts we initially assume that all participants report the same, unknown, measure of centrality as their forecast. In the next subsection, we explore potential heterogeneity in the measure of centrality used by different respondents.

table[table omitted — 1,464 chars of source]

(ref) presents the results of rationality tests for three measures of central tendency, and for a variety of instrument sets. The first instrument set includes just a constant, and the rationality test simply tests whether the forecast errors have unconditional mean, median or mode, respectively, of zero. The other instrument sets additionally include the forecast ($X_t$) itself, and other information about the respondent collected in the survey. We consider the respondent's income at the time of making the forecast, indicators for the respondent's type of employer,\footnote{The survey includes the categories government, private (for-profit), non-profit, family business, and “other.” The first two categories dominate the responses and so we only consider indicators for those.} and whether the respondent received any job offers in the past four months.

figure[figure omitted — 811 chars of source]

The first row of (ref) shows that when only a constant is used, rationality of the survey forecasts can be rejected for the mean, but cannot be rejected for the other two measures of central tendency. When we additionally include the forecast as an instrument, we can reject rationality as mean or median forecasts, but we cannot reject rational mode forecasts. We are similarly able to reject rationality as mean and median forecasts when we include additional covariates, but find no evidence against rationality when these forecasts are interpreted as mode forecasts.

(ref) shows the convex combinations of mean, median and mode forecasts that lie in the confidence set constructed using the methods for weakly-identified GMM estimation in StockWright2000. In the left panel we see that when using only a constant as the instrument, we are able to reject rationality for the mean, and for measures of centrality “close” to the mean, but we are unable to reject rationality for the median or mode and measures in a neighborhood of these. This is consistent, naturally, with the entries in the first row of (ref), which correspond directly to the three vertices in (ref). Using the results in Section (ref) and Supplemental Appendix (ref), the left panel of (ref) can be interpreted assuming either that survey respondents all use the same mixture of loss functions, or that respondents each use a single measure of centrality and are possibly heterogeneous in which measure they use. In the former case the region with black circles identifies the set of loss function weights that are consistent with rationality; in the latter case this region indicates the set of proportions of forecaster types in the population.

In the right panel of (ref), when the instrument set includes a constant and the forecast, we see that only the mode and centrality measures very close to the mode are included in the confidence set; all other forecasts can be rejected at the 5% level. Overall, we conclude that the responses to the New York Fed's income survey, taken on aggregate, are consistent with rationality when interpreted as mode forecasts, but not when interpreted as forecasts of the mean, median, or convex combinations of these measures.\footnote{Given the panel structure of our data, the target variable may be subject to common unpredictable shocks, leading to forecast errors correlated across individuals within the same wave. In Supplemental Appendix Section (ref), we repeat all of our analyses using a covariance estimator clustered at the time level and find very similar results.}

{\color{black}Holding power issues aside (discussed in the next paragraph), these test results indicate that at least some survey respondents have asymmetric predictive distributions, as the unconditional test rejects mean, but not mode and median rationality, and the conditional test rejects rationality for all centrality measures except those near the mode. Inferring (a)symmetry of the respondents' predictive distributions is not straightforward, as only point forecasts and realizations are observed. As shown by bai2005tests, symmetry of the (unconditional) distribution of the observed forecast errors is not sufficient for symmetry of the predictive distribution.}

When interpreting the results in (ref), and similar figures below, it is worth keeping in mind that the power to detect sub-optimal forecasts is not uniform across values of $\theta$: sampling variation in the mean and median vanishes at rate $T^{-1/2}$, while for the mode it vanishes only at rate approximately $T^{-2/7}$ (see Theorem (ref)). This implies that for comparably sub-optimal forecasts, power will be lower at the mode vertex than at the mean or median vertices. This unavoidable variation in power means that the information conveyed by inclusion in the confidence set differs across values of $\theta$.\footnote{StockWright2000 suggest caution when interpreting a small but nonempty confidence set, as such an outcome is consistent with both a correctly-specified model (a rational forecast, in our case) estimated precisely and also with a misspecified model (irrational forecast) facing low power. These two interpretations clearly have very different economic implications, but cannot be disentangled empirically. Given the lower power at the mode vertex, the latter explanation may be relevant here.}

The analysis of individual income survey forecasts above used $3,916$ pairs of forecasts and realizations from a total of $2,628$ unique survey respondents. This naturally raises the question of whether there is heterogeneity in the measure of centrality used by different respondents. Given that our survey respondents generally only appear once or twice in our sample, allowing for arbitrary heterogeneity is not empirically feasible. In our analysis below we instead stratify survey respondents by observable characteristics and test forecast rationality separately for each subsample. Stratifying the sample may reduce power, due to the smaller sample sizes available, or it may reveal irrationality that is undetected in a joint analysis, for example if one subsample deviates strongly from rationality while the remaining subsamples are rational, or if one subsample exhibits more asymmetric predictive distributions.

Forecast rationality and income level

figure[figure omitted — 776 chars of source]

Firstly, we consider stratifying our sample by income. This is motivated by the possibility that, in addition to a different level of future income, low-income respondents face a different shape of future income, compared with high-income respondents. This analysis may also reveal that respondents at different income levels use different centrality measures to summarize their predictive distribution. (ref) presents confidence sets for forecast rationality of measures of centrality for low-, middle-, and high-income respondents based on terciles of the distribution of reported income.\footnote{Qualitatively similar results are found if we stratify the sample into just two groups based on median income.} We see that for low-income respondents, only the mode and measures very close to the mode are contained in the confidence set. For middle- and high-income respondents, the mode, mean, and centrality measures “close” to the mean and mode are included in the confidence set. This finding is consistent with all respondents using the mode, and only the distribution for low-income respondents allowing for separate identification of the mode and the mean. Specifically, if the conditional distribution of future earnings is more skewed for low-income respondents than for high-income respondents, then the measures of centrality are better separated for the former, which allows for better identification of the functional used by survey respondents. As income is commonly correlated with numeracy and education levels, rejecting mean- and median-rationality for lower-income respondents is consistent with work in the experimental literature: KroegerThibaud2019QuestionFormats observe high-numeracy individuals to be better at mean and median forecasting.

Forecast rationality and age

figure[figure omitted — 302 chars of source]

We next stratify the sample both by income and by age, motivated by the possibility that younger respondents have less experience in the workforce and may be less able to predict their future earnings. The latest versions of the FRBNY survey contain a field for whether the respondent is above or below 40 years of age, while earlier versions asked for the respondent's specific age. To maximize coverage, we adopt the age 40 split contained in the later versions of the survey. We have 2,457 (1,332) respondents over (under) that age. We use the median income within each age category to define “low” and “high” income groups.

(ref) presents the striking result that forecasts from younger low-income workers cannot be rationalized using any measure of centrality; all points lie outside the 95% confidence set. Forecasts from older low-income respondents can be rationalized as only mode forecasts. In contrast, forecasts from both young and old high-income workers can be rationalized for a variety of measures of centrality.\footnote{Note that a large confidence set for $\theta$ does not imply that the respondents are “more rational” than if the confident set were small; a respondent can be fully rational at just a single value of $\theta$, and if the predictive distribution is sufficiently skewed then the confidence set will, in the limit, contain just that value of $\theta$. A large confidence set may reflect a predictive distribution that does not allow for the identification of a single value of $\theta$, or low finite-sample power.} This figure suggests that younger low-income workers make systematic errors when attempting to predict their income in the coming four months. Asymmetric loss functions, such as those used in EKT2005 (EKT), have the potential to explain these forecasts, and we consider this in Section (ref) below. Foreshadowing those results, we also reject rationality of forecasts from younger low-income respondents in the EKT framework.

Forecast rationality and job stability

An important component of income uncertainty is job search and new job offers, see mueller2022expectations for a recent review of survey expectations in the job search literature. Respondents naturally have different prospects of changing jobs, and these changes impact their predictive distributions of income. To proxy for the likelihood of changing jobs, we stratify respondents based on whether they reported receiving at least one job offer over the previous four months. \footnote{We find very similar results when we stratify respondents using their estimated probability of receiving a job offer in the next four months, or by their estimated probability of staying in the same job for the next four months.}

{\color{black}(ref) reveals that respondents who did not receive a job offer, and high-income respondents who did, provide forecasts that can be rationalized as mean, mode, and combinations thereof. For high-income respondents without job offer the confidence set contains also (near-)median forecasts, indicating that the predictive distributions for those respondents are more symmetric. Forecasts from low-income respondents who reported receiving a job offer can only be rationalized as mode forecasts; rationality for all other centrality measures is rejected. This is consistent with the predictive distributions for these respondents being more asymmetric, allowing for discrimination between centrality measures. Overall, these results reveal that low-income respondents with more job uncertainty differ meaningfully from the other three subgroups.}

figure[figure omitted — 323 chars of source]

Forecast rationality and survey experience

figure[figure omitted — 320 chars of source]

As economic agents gather experience, their actions and expectations change, see vanNVeldkamp2006learning and malmendier2011depression for example. In recent work, kim2022learning show that this effect is true in economic surveys as well: forecasts of inflation show less uncertainty as respondents garner experience responding to the survey. We next investigate whether having survey experience leads to more rational forecasts. The FRBNY survey includes participants for a maximum of twelve months, and within that period they are asked to predict future income and report current income three times, providing us with up to two matched pairs of forecasts and realizations for each respondent. We have a total of 1,288 respondents with two matched pairs, and we study the rationality of the first and second of these forecasts separately.

Figure (ref) reveals a stark difference in the impact of the round of elicitation between low- and high-income respondents. Forecasts from low-income respondents can be rationalized as mode forecasts, and the results change only slightly from the first to the second round: in the first round some “near mode” functionals are also rationalizable, and in the second round the mode functional is only rationalizable at the 95% significance level. For high-income respondents, in contrast, only the mode and mean functionals (and convex combinations thereof) are rationalizable in the first round, while in the second round, four months later, all functionals are rationalizable. Assuming that the true conditional distributions of future income are unaffected by survey participation, and from the second round results we can infer these are (approximately) symmetric, this suggests that high-income respondents' forecasts are more accurate in the second round, and as a result can be (correctly) rationalized via more functionals.\footnote{Recall that a large confidence set does not imply “more rational,” merely that the forecast is rationalizable by a wider variety of measures of centrality.} Our results for high-income respondents are consistent with the findings of kim2022learning for inflation forecasts, while the similarity across survey rounds for low-income respondents differs from that study.

In Supplemental Appendix (ref), we present an additional analysis, stratifying respondents in the second round by the size of their forecast error in the first round. We find that respondents with high accuracy in the first round can be rationalized using almost any measure of centrality, regardless of income, as can high-income low-accuracy respondents. In contrast, lower-income respondents with low accuracy in the first round cannot be rationalized using any measure of centrality at the 90% confidence level.

Irrational or optimistic?

The premise of the rationality test proposed in this paper is that survey respondents use a measure of central tendency when providing a point forecast but the specific centrality measure used is unknown. An alternative framework is to model the respondent as reporting either a centrality measure or some other functional of their predictive distribution. For example, rather than using the median they may use some other quantile, capturing optimism or pessimism. EKT2005 (EKT) consider such a case, and model respondents as using either “lin-lin” loss, which elicits a quantile of the predictive distribution, or “quad-quad” loss, which elicits an expectile of the predictive distribution. In both cases, the loss function contains an asymmetry parameter, and the special case of symmetry corresponds to the respondent using the median or the mean.

table[table omitted — 1,548 chars of source]

Table (ref) presents tests of rationality in the EKT framework, as well as the results for the three measures of centrality considered in this paper. These are presented for the full sample, the three income subsamples, and the $2\times2$ stratifications based on age and income. Results for the other stratifications considered above are presented in Supplementary Appendix (ref). As in our previous analysis, we use a constant and the forecast as instruments for the EKT tests.

For the full sample of respondents, in the top line of Table (ref), we see that rationality is rejected in the EKT framework. When interpreted as a quantile forecast, the respondents' estimated quantile 90% confidence interval is 0.61 to 0.64, indicating optimism, but the $p$-value for the test of rationality is less than 0.005, rejecting rationality.\footnote{The estimated shape parameter is interpretable as the one that makes the observed forecasts as close as possible to rational, and the corresponding $p$-value determines whether these forecasts are consistent or inconsistent with rationality.} When interpreted as an expectile forecast, the shape parameter straddles 0.5 (the case of symmetry, where the expectile equals the mean) but the test again rejects rationality. The only functional for which rationality is not rejected is the mode, using the new mode forecast rationality test proposed in this paper.

The results of the rationality tests stratified by income, presented in the second panel of Table (ref), show that middle- and high-income respondents can be rationalized as expectile forecasters, and in both cases the confidence interval for the estimated shape parameter includes the point of symmetry, consistent when the mean functional not being rejected for these subsamples. For low-income respondents, we see that allowing for optimism or pessimism through the EKT framework does not rationalize these forecasts: the $p$-values of the EKT rationality tests for these cases are both less than or equal to 0.01.

Finally, we stratify survey respondents by both age and income, and present the results in the lower panel of Table (ref). Recall that above we found that the younger, low-income subsample could not be rationalized as any measure of centrality. When using the EKT framework, we find that the estimated shape parameters for the lin-lin and quad-quad loss functions are both greater than 0.5, the point of symmetry, indicating optimistic forecasts. However the $p$-values in both cases are less than 0.005, strongly rejecting rationality. Thus, forecasts from younger low-income respondents cannot be rationalized as centrality measures, nor as quantiles or expectiles.

Conclusion

Reasonable people can interpret a request for their prediction of a random variable in a variety of ways. Some, including perhaps most economists, will report their expectation of the value of the variable (i.e., the mean of their predictive distribution), others might report the value such that the observed outcome is equally likely to be above or below it (i.e., the median), and others may report the value most likely to be observed (i.e, the mode). Still others might solve a loss minimization problem and report a forecast that is not a measure of central tendency. Economic surveys generally request a point forecast, despite calls for surveys to solicit distributional forecasts, see manski2004measuring for example, and the specific type of point forecast (mean, median, etc.) to be reported is generally not made explicit in the survey.

This paper proposes new methods to test the rationality of forecasts of some unknown measure of central tendency. Similar to EKT2005, we propose a testing framework that nests the mean forecast as a special case, but unlike that paper we allow for alternative forecasts within a general class of measures of central tendency, rather than measures that represent other aspects of the predictive distribution (such as non-central quantiles or expectiles). We consider a class that nests convex combinations of the mean, median and mode, and we overcome an inherent identification problem that arises using the methods of StockWright2000.

As a building block for the above test, we also present new tests for the rationality of mode forecasts. Some recent work (e.g., Knuppel2012; ReifschneiderTulip2017) suggests that point forecasts from central banks should be interpreted as mode rather than mean forecasts, and experimental studies (e.g., peterson1964mode; KroegerThibaud2019) report that participants are more likely to use the mode to summarize their predictive distribution than other measures of central tendency. While mode regression has received some attention in the recent literature (Kemp2012 and Kemp2019), tests for mode forecast rationality similar to those available for the mean and median (e.g., MincerZarnowitz1969 and GaglianoneJBES2011) are lacking. Direct analogs of existing tests are infeasible because the mode is not elicitable Heinrich2014. We introduce the concept of asymptotic elicitability and show it applies to the mode by considering a generalized modal midpoint with asymptotically vanishing length. We then present results that allow for tests similar to the famous Mincer-Zarnowitz regression for mean forecasts.

We apply our tests to individual income expectation survey data collected by the Federal Reserve Bank of New York. We reject forecast rationality when interpreting responses as mean or median forecasts, however we cannot reject rationality when interpreting them as mode forecasts. We also find evidence of heterogeneity across respondents: for example, forecasts from younger, low-income respondents are not rationalizable using any measure of centrality, while forecasts from high-income respondents, regardless of their age, are rationalizable for many, though not all, measures of centrality. Further, we find that the behavior of high-income survey participants changes as they gain experience in the survey, consistent with the learning suggested in kim2022learning.

Previous analyses of economic surveys have typically assumed that responses are mean forecasts, and deviations from that benchmark have motivated learning and attention models which can account for such irrationality bordalo2020overreaction,coibion2015information,kohlhas2021asymmetric. This paper proposes an intriguingly simple alternative to mean rationality, and one which is supported by our empirical work: point forecasts may not reflect the statistical average, but simply the most likely value, the mode.

\singlespacing {4pt}

\doublespacing \setcounter{page}{1} \setcounter{footnote}{0} \setcounter{section}{0} \setcounter{table}{0} \setcounter{figure}{0}

center[center omitted — 242 chars of source]

This appendix contains additional details, discussions, and results. It is structured as follows: {\color{black}Section (ref) contains proofs for the theorems in the main paper and (ref) contains some additional technical lemmas and their proofs.} {\color{black}Sections (ref) and (ref)} discuss the impact of the choice of kernel and bandwidth on the results for mode forecast rationality testing in Section (ref) of the main paper. Section (ref) presents a result showing that a convex combination of functionals is generally not elicitable, as stated in Remark (ref) of the main paper. {\color{black}Section (ref) relates our tested null hypothesis to a probabilistic combination of mean, median and mode forecasters.} Section (ref) presents simulation results supplementing those presented in Section (ref) of the main paper. Section (ref) presents additional results for the application in Section (ref) of the main paper. In particular, we show tests of rationality using a cluster covariance estimator, results for an additional stratification with respect to past forecast accuracy and the results of tests of rationality in the framework of EKT2005 for the full set of stratifications. Section (ref) presents two additional empirical analyses, the first to the “Greenbook” forecasts produced by the Board of Governors of the Federal Reserve, and the second to random walk forecasts of exchange rates.

Proofs

proof[Proof of Theorem (ref)] For the proof of statement (ref), let $Y \sim P \in \mathcal{P}$ with density $f$, define $\tilde K_\delta(e) = \frac{1}{\delta} K\left( \frac{e}{\delta} \right)$ and introduce the notation $\bar L^K_\delta(x,P) = \mathbb{E}_{Y\sim P} \left[ L^K_\delta(x,Y) \right]$. Then, it holds that \begin{align*} \bar L^K_\delta(x,P) = - \int \frac{1}{\delta} K\left( \frac{x-y}{\delta} \right) f(y) \, \mathrm{d}y = - \int \tilde K_\delta \left( x-y \right) f(y) \, \mathrm{d}y = - (f \ast \tilde K_\delta) (x), \end{align*} where $f \ast \tilde K_\delta$ denotes the convolution of the functions $f$ and $ \tilde K_\delta$. Ibragimov1956 shows that for any log-concave density, its convolution with any other unimodal distribution function is again unimodal.\footnote{Ibragimov1956 calls densities satisfying this property strongly unimodal. It is important to note that his notion of strong unimodality is different from ours introduced in (ref).} Thus, $\bar L^K_\delta(x,P)$ exhibits a unique minima which shows that $\Gamma_\delta^K$ is well-defined by ((ref)). We continue with statement (ref) and show that $\Gamma_\delta(P) \to \operatorname*{Mode}(P)$ for all $P \in \mathcal{P}$. Notice that \begin{align} \bar L^K_\delta(x,P) = - \int \frac{1}{\delta} K\left( \frac{x-y}{\delta} \right) f(y) \, \mathrm{d}y = - \int f(x + u \delta) K\left(u \right) \, \mathrm{d}u = - \int f(x + u \delta) \, \mathrm{d} \mathcal{K}(u), \end{align} by applying integration by substitution and by interpreting the kernel $K(\cdot)$ as the density of the probability measure $\mathcal{K}$. It holds that \begin{align} &\sup_{x \in \mathbb{R}} \left| \bar L^K_\delta(x,P) - \big( - f(x) \big) \right| = \sup_{x \in \mathbb{R}} \left| f(x) - \int f(x + u \delta) \mathrm{d} \mathcal{K}(u) \right| \\ \le \, &\sup_{x \in \mathbb{R}} \int \big| f(x) - f(x + u \delta) \big| \mathrm{d} \mathcal{K}(u) \le \sup_{x \in \mathbb{R}} \int | c \delta u | \, \mathrm{d} \mathcal{K}(u) = c \delta \int |u| {K}(u) \mathrm{d}u \to 0 \end{align} as $\delta \to 0$ as $f$ is Lipschitz continuous with constant $c\ge0$ and $\int |u| {K}(u) \mathrm{d}u < \infty$. Hence, $\bar L^K_\delta(x,P)$ converges uniformly (for all $x \in \mathbb{R}$) to $-f(x)$ as $\delta \to 0$ and it holds that $\operatorname*{arg\,min}_x \bar L^K_\delta(x,P) \to \operatorname*{arg\,min}_x (- f(x) )$ as $\delta \to 0$. Consequently, \begin{align*} \lim_{\delta \to 0} \Gamma_\delta(P) = \lim_{\delta \to 0} \left( \operatorname*{arg\,min}_{x \in \mathbb{R}} \bar L^K_\delta(x,P) \right) = \operatorname*{arg\,min}_{x \in \mathbb{R}} \left( \lim_{\delta \to 0} \bar L^K_\delta(x,P) \right) = \operatorname*{arg\,min}_{x \in \mathbb{R}} \left( - f(x) \right), \end{align*} which equals the mode for distributions with continuous Lebesgue density by definition. For the proof of statement (ref), {\color{black}let $P \in \tilde{\mathcal{P}}$ with density $f$. We} consider a fixed $\delta > 0$ and define $\bar V^K_\delta(x, P) \mathrel{\mathop:}= - \mathbb{E}_{Y\sim P} [ V_\delta(x,Y) ] = \frac{1}{\delta^2}\int K' \left( \frac{x-y}{\delta} \right) f(y) \mathrm{d}y$ and $\tilde K_\delta'(e) = \frac{1}{\delta^2} K'\left( \frac{e}{\delta} \right)$. Then, \begin{align} \bar V^K_\delta(x, P) &= \int \tilde K_\delta'(x-y) f(y) dy = \big( \tilde K_\delta' \ast f \big) (x) = \big( \tilde K_\delta \ast f' \big) (x) = \int \tilde K_\delta(x-y) f'(y) dy \\ &= \int \tilde K_\delta(x+\operatorname*{Mode}(P)-y) f'(y-\operatorname*{Mode}(P)) dy\\ &= \int \tilde K_\delta(x+\operatorname*{Mode}(P)-y) g'(y) dy \end{align} for some shifted density $g$ with mode at zero. As the kernel $\tilde K_\delta$ is log-concave it has a monotone likelihood ratio (Proposition 2.3 (b) in Saumard2014), i.e.\ for any $a \le b$ and $y \ge 0$, it holds that $\frac{\tilde K_\delta(a-y)}{\tilde K_\delta(a)} \le \frac{\tilde K_\delta(b-y)}{\tilde K_\delta(b)}$ and as $g'(y) \le 0$ for $y\ge 0$ {\color{black}(as $P \in \tilde{\mathcal{P}}$ is strongly unimodal)}, this implies that \begin{align} \frac{\tilde K_\delta(a-y)}{\tilde K_\delta(b-y)} g'(y) &\ge \frac{\tilde K_\delta(a)}{\tilde K_\delta(b)} g'(y). \end{align} Analogously, the same inequality follows for any $a \le b$ and $y \le 0$, where the monotone likelihood ratio above holds with the reverse inequality but at the same time $g'(y) \ge 0$ {\color{black}as $P \in \tilde{\mathcal{P}}$.} Thus, for any $x>z$, \begin{align} \bar V^K_\delta(x, P) &= \int \tilde K_\delta (x+\operatorname*{Mode}(P)-y) g'(y) dy \\ &= \int \frac{\tilde K_\delta (x+\operatorname*{Mode}(P)-y)}{\tilde K_\delta (z+\operatorname*{Mode}(P)-y)} \tilde K_\delta (z+\operatorname*{Mode}(P)-y) g'(y) dy \\ &\ge \frac{\tilde K_\delta (x+\operatorname*{Mode}(P))}{\tilde K_\delta (z+\operatorname*{Mode}(P))} \int \tilde K_\delta (z+\operatorname*{Mode}(P)-y) g'(y) dy \\ &= \frac{\tilde K_\delta (x+\operatorname*{Mode}(P))}{\tilde K_\delta (z+\operatorname*{Mode}(P))} \; \bar V^K_\delta(z, P). \end{align} For any root $x^\ast$ such that $\bar V^K_\delta(x^\ast, P)= 0$, it follows that $\bar V^K_\delta(x_u, P) \ge 0$ for all $x_u > x^\ast$ and $\bar V^K_\delta (x_l, P) \le 0$ for all $x_l < x^\ast$ as the kernel $\tilde K_\delta$ has infinite support. If there exist two distinct roots $x_l^\ast < x_u^\ast$ of $\bar V^K_\delta (\cdot, P)$, the above implies that $\bar V^K_\delta(x, P)=0$ for all $x \in [x_l^\ast, x_u^\ast]$, which implies that the roots of $\bar V^K_\delta (\cdot, P)$ constitute a closed interval. As $\bar L_\delta^K(x,P)$ has a unique maximum for all $P \in \tilde{\mathcal{P}} \subset \mathcal{P}$ by part (ref) and $\bar V^K_\delta(x, P) = - \frac{\partial}{\partial x} \bar L_\delta^K(x,P)$, it follows that $\bar V^K_\delta(x, P) $ must have a unique root, which implies that $V_\delta^K(x,Y)$ is a strict identification function for the generalized modal midpoint.
proof[Proof of Theorem (ref)] We define \begin{align} g_{t,T} &:= \delta_T^{3/2} T^{-1/2} \psi(Y_{t+1},X_t,\mathbf{h}_t, \delta_T) = - (T\delta_T)^{-1/2} K'\left( \frac{X_t-Y_{t+1}}{\delta_T} \right) \mathbf{h}_t, \\ g_{t,T}^e &:= \mathbb{E}_t[g_{t,T} ], \qquad and \qquad g_{t,T}^\ast :=g_{t,T} - g_{t,T}^e, \end{align} such that \begin{align} \delta_T^{3/2} T^{-1/2} \sum_{t=1}^T \psi(Y_{t+1},X_t,\mathbf{h}_t, \delta_T) = \sum_{t=1}^T g_{t,T}^e + \sum_{t=1}^T g_{t,T}^\ast. \end{align} Lemma (ref) shows that $\sum_{t=1}^T g_{t,T}^e \overset{P}{\to} 0$. Thus, it remains to shows that $\sum_{t=1}^T g_{t,T}^\ast \overset{d}{\to} \mathcal{N} \big( 0, \Omega_{\mathrm{Mode}} \big)$. For some arbitrary, but fixed $\lambda \in \mathbb{R}^k, || \lambda ||_2 = 1$, we define \begin{align} z_{t,T} := \lambda^\top g_{t,T}^\ast, \qquad \bar \omega_T^2 := \sum_{t=1}^T \operatorname{Var}(z_{t,T}), \qquad h_{t,T} := \frac{z_{t,T}}{\bar \omega_T}, \qquad and \qquad \omega^2 := \lambda^\top \Omega_\mathrm{Mode} \lambda, \end{align} and show that a univariate CLT for martingale difference arrays (MDA) holds for $\sum_{t=1}^T h_{t,T}$. It obviously holds that $\big( g_{t,T}^\ast, \mathcal{F}_{t+1} \big)$ is a MDA as $g_{s,T}^\ast \in \mathcal{F}_{t+1}$ for all $s \le t$ and $\mathbb{E}_t \big[ g_{t,T}^\ast \big] = 0$ a.s.\ by definition. Thus, $\big( z_{t,T}, \mathcal{F}_{t+1} \big)$ and $\big( h_{t,T}, \mathcal{F}_{t+1} \big)$ are also MDAs. In the following, we verify the following three conditions of Theorem 24.3 of davidson1994stochastic: (a) $\sum_{t=1}^T \operatorname{Var}( h_{t,T}) = 1$, (b) $\sum_{t=1}^T h_{t,T}^2 \overset{P}{\to} 1$, and (c) $\operatorname{max}_{1 \le t \le T} |h_{t,T}| \overset{P}{\to} 0$. Lemma (ref) shows that $ \bar \omega_T^2 = \sum_{t=1}^T \operatorname{Var}[ z_{t,T} ] \to \lambda^\top \Omega_\mathrm{Mode} \lambda = \omega^2$. Thus, as $\Omega_\text{Mode}$ is assumed to be positive definite, $\bar \omega_T^2$ is strictly positive for all $T$ sufficiently large and hence, $h_{t,T}$ is well-defined and $\sum_{t=1}^T \operatorname{Var}( h_{t,T}) = 1$, which shows condition (a). Lemma (ref) shows that $\sum_{t=1}^T z_{t,T}^2 \overset{P}{\to} \omega^2$ as $T \to \infty$ which implies condition (b), i.e. $\sum_{t=1}^T h_{t,T}^2 = \sum_{t=1}^T \frac{z_{t,T}^2}{\bar \omega_T^2} \overset{P}{\to} 1$. Eventually, Lemma (ref) shows condition (c) and we can apply Theorem 24.3 of davidson1994stochastic in order to conclude that for all $\lambda \in \mathbb{R}^k, || \lambda||_2 = 1$, it holds that $\sum_{t=1}^T h_{t,T} \overset{d}{\to} \mathcal{N}(0,1)$. As $\bar \omega_T^2 \to \omega^2$, Slutsky's theorem implies that $\sum_{t=1}^T z_{t,T} =\sum_{t=1}^T \lambda^\top g_{t,T}^\ast \overset{d}{\to} \mathcal{N} \big( 0, \omega^2 \big)$, and as this holds for all $\lambda \in \mathbb{R}^k, || \lambda||_2 = 1$, we apply the Cram\'er-Wold theorem and get that $\sum_{t=1}^T g_{t,T}^\ast \overset{d}{\to} \mathcal{N} \big( 0, \Omega_{\mathrm{Mode}} \big)$, which concludes the proof of this theorem.
proof[Proof of Theorem (ref)] Let $\lambda \in \mathbb{R}^k$, $||\lambda||_2 = 1$ be a fixed and deterministic vector. Then, \begin{align} \begin{aligned} &\lambda^\top \widehat{\Omega}_{T,\mathrm{Mode}} \lambda - \lambda^\top \Omega_{T,\mathrm{Mode}} \lambda \\ = \, &\frac{1}{T} \sum_{t=1}^T \delta_T^{-1} K' \left( \frac{X_t-Y_{t+1}}{\delta_T} \right)^2 (\lambda^\top \mathbf{h}_t)^2 - \frac{1}{T} \sum_{t=1}^T \mathbb{E}_t \left[ \delta_T^{-1} K' \left( \frac{X_t-Y_{t+1}}{\delta_T} \right)^2 (\lambda^\top \mathbf{h}_t)^2 \right] \\ + \, &\frac{1}{T} \sum_{t=1}^T \mathbb{E}_t \left[ \delta_T^{-1} K' \left( \frac{X_t-Y_{t+1}}{\delta_T} \right)^2 (\lambda^\top \mathbf{h}_t)^2 \right] - \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ (\lambda^\top \mathbf{h}_t)^2 f_{t}(0) \int K'(u)^2 \mathrm{d}u \right]. \end{aligned} \end{align} We start by showing that the last line in ((ref)) is $o_P(1)$. It holds that \begin{align} &\frac{1}{T} \sum_{t=1}^T \mathbb{E}_t \left[ \delta_T^{-1} K' \left( \frac{X_t-Y_{t+1}}{\delta_T} \right)^2 (\lambda^\top \mathbf{h}_t)^2 \right] - \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ (\lambda^\top \mathbf{h}_t)^2 f_{t}(0) \int K'(u)^2 \mathrm{d}u \right] \\ =\;&\frac{1}{T} \sum_{t=1}^T (\lambda^\top \mathbf{h}_t)^2 \int K' \left( u \right)^2 f_t(\delta_T u) \mathrm{d}u -\frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ (\lambda^\top \mathbf{h}_t)^2 \int K'(u)^2 f_{t}(0) \mathrm{d}u \right] \overset{P}{\to} 0 \end{align} as $ f_t(\delta_T u) \to f_t(0) \le c$ a.s., and by applying a law of large numbers for mixing data as $\mathbb{E} \left[ ||\mathbf{h}_t||^{2r+\delta} \right] < \infty$. We further show that the penultimate line in ((ref)) converges to zero in $L_p$ ($p$-th mean) for some $p > 1$ small enough. By applying the BahrEsseen1965 inequality for MDA (notice that the summands of the sum in the first line of (ref) naturally form an MDA without imposing our null hypothesis), we get \begin{align} \begin{aligned} &\mathbb{E} \left[ \left| \frac{1}{T} \sum_{t=1}^T \delta_T^{-1} K' \left( \frac{\varepsilon_t}{\delta_T} \right)^2 (\lambda^\top \mathbf{h}_t)^2 - \frac{1}{T} \sum_{t=1}^T \mathbb{E}_t \left[ \delta_T^{-1} K' \left( \frac{\varepsilon_t}{\delta_T} \right)^2 (\lambda^\top \mathbf{h}_t)^2 \right] \right|^p \right] \\ \le \, & 2 T^{-p} \sum_{t=1}^T \mathbb{E} \left[ \left| \delta_T^{-1} K' \left( \frac{\varepsilon_t}{\delta_T} \right)^2 (\lambda^\top \mathbf{h}_t)^2 \right|^p \right] + 2 T^{-p} \sum_{t=1}^T \mathbb{E} \left[ \left| \mathbb{E}_t \left[ \delta_T^{-1} K' \left( \frac{\varepsilon_t}{\delta_T} \right)^2 (\lambda^\top \mathbf{h}_t)^2 \right] \right|^p \right]. \end{aligned} \end{align} For the first term in (ref), we get that \begin{align*} &T^{-p} \sum_{t=1}^T \mathbb{E} \left[ \left| \delta_T^{-1} K' \left( \frac{\varepsilon_t}{\delta_T} \right)^2 (\lambda^\top \mathbf{h}_t)^2 \right|^p \right] = (T \delta_T)^{-p} \sum_{t=1}^T \mathbb{E} \left[ \left| \lambda^\top \mathbf{h}_t \right|^{2p} \int \left| K' \left( \frac{e}{\delta_T} \right) \right|^{2p} f_t(e) \mathrm{d}e \right] \\ = \, &(T \delta_T)^{1-p} \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \left| \lambda^\top \mathbf{h}_t \right|^{2p} \int \left| K' \left( u \right) \right|^{2} f_t(\delta_T u) \mathrm{d}u \right] \to 0, \end{align*} as $(T \delta_T)^{1-p} \to 0$ for any $p > 1$, $\mathbb{E} \left[ ||\mathbf{h}_t||^{2p} \right] < \infty$ for $p > 1$ small enough, the density $f_t$ is bounded from above, and $\int |K'(u)|^2 \mathrm{d}u < \infty$ by assumption. The second term in (ref) converges by a similar argument as further detailed in ((ref)) in the proof of Lemma (ref). As $L_p$ convergence for any $p > 1$ implies convergence in probability, the result of the theorem follows.
proof[Proof of Theorem (ref)] We let \begin{align*} g_{t,T} &:= \delta_T^{3/2} T^{-1/2} \psi(Y_{t+1},X_t,\mathbf{h}_t, \delta_T) = - (T\delta_T)^{-1/2} K'\left( \frac{X_t-Y_{t+1}}{\delta_T} \right) \mathbf{h}_t, \\ g_{t,T}^e &:= \mathbb{E}_t[g_{t,T} ], \qquad and \qquad g_{t,T}^\ast :=g_{t,T} - g_{t,T}^e, \end{align*} such that $\sum_{t=1}^T g_{t,T} = \sum_{t=1}^T g_{t,T}^e + \sum_{t=1}^T g_{t,T}^\ast$ as in the proof of Theorem (ref). Then, it holds that $\sum_{t=1}^T g_{t,T}^\ast \overset{d}{\to} \mathcal{N}(0,\Omega_{\mathrm{Mode}})$ as in the proof of Theorem (ref). Notice for this that the employed Lemmas (ref)--(ref) do not require the null hypothesis in (ref). We continue to analyze the term $\sum_{t=1}^T g_{t,T}^e$ as in the proof of Lemma (ref). However, the Taylor expansion in (ref) now entails another term as we now impose $\mathbb{H}_{A, \text{loc}}$ in (ref) instead of using $f_t'(0) = 0$ that resulted from $\mathbb{H}_{0}$ in (ref). Hence, using that $\int K( u ) \mathrm{d}u = 1$, we get \begin{align} \sum_{t=1}^T g_{t,T}^e = T^{-1/2} \delta_T^{3/2} \sum_{t=1}^T f'_{t} (0) \mathbf{h}_t + o_P(1) &= c \cdot \left( a_T T^{1/2} \delta_T^{3/2} + o_P(1) \right) + o_P(1) \end{align} For statement (ref), $a_T T^{1/2} \delta_T^{3/2} \overset{P}{\to} 1$, and thus, $\sum_{t=1}^T g_{t,T}^e = c + o_P(1)$. As Theorem (ref) shows that $\widehat{\Omega}_{T,\mathrm{Mode}} - \Omega_{T, \textrm{Mode}} \overset{P}{\to} 0$ (without imposing our null hypothesis in (ref)), applying Slutzky's theorem implies $\widehat{\Omega}_{T,\mathrm{Mode}}^{-1/2} \sum_{t=1}^T g_{t,T} \overset{d}{\to} \mathcal{N} \big( \Omega_{\mathrm{Mode}}^{-1/2} c, I_k \big)$. Hence, $J_T$ converges to a non-central $\chi^2$-distribution, \begin{align*} J_T = \left( \widehat{\Omega}_{T,\mathrm{Mode}}^{-1/2} \sum_{t=1}^T g_{t,T} \right)^\top \left( \widehat{\Omega}_{T,\mathrm{Mode}}^{-1/2} \sum_{t=1}^T g_{t,T} \right) \overset{d}{\to} \chi^2_k \big( c^\top \Omega_{Mode}^{-1} c \big). \end{align*} For statement (ref), $a_T T^{1/2} \delta_T^{3/2} = o_P(1)$ and hence, $\sum_{t=1}^T g_{t,T}^e = o_P(1)$ putting us exactly in the situation of Theorem (ref). For statement (ref), it holds that $\left( a_T T^{1/2} \delta_T^{3/2} \right)^{-1} \overset{P}{\to} 0$, which implies that for any $ \bar c \in \mathbb{R}$, $\mathbb{P} \left( \left| \sum_{t=1}^T g_{t,T}^e \right| \ge \bar c \right) \to 1$, and consequently also $\mathbb{P} \left( \left| \sum_{t=1}^T g_{t,T} \right| \ge \bar c \right) \to 1$. As furthermore $J_T = \left( \sum_{t=1}^T g_{t,T} \right)^\top \widehat{\Omega}_{T,\mathrm{Mode}}^{-1} \left( \sum_{t=1}^T g_{t,T} \right)$ and $\widehat{\Omega}_{T,\mathrm{Mode}} - \Omega_{T,\mathrm{Mode}} \overset{P}{\to} 0$ by Theorem (ref), where $\Omega_{T,\mathrm{Mode}}$ is uniformly positive definite for $T$ large enough, the conditions of Theorem 8.13 of White1994 are satisfied and we can conclude that for any $\bar c \in \mathbb{R}$, $\mathbb{P} \left( \left| J_T \right| \ge \bar c \right) \to 1$, which concludes the proof of this theorem.
proof[Proof of Theorem (ref)] For all fixed $\lambda \in \mathbb{R}^k$ such that $||\lambda||_2 = 1$, we define \begin{align} \begin{aligned} \sigma_T^{2} := \lambda^\top \Sigma_T(\theta_0) \lambda &= \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{10}^2 \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right)^2 \varepsilon_t^2 + \theta_{20}^2 \left( \mathbf{h}_t^\top w_{\mathrm{Med}} \lambda \right)^2 \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big)^2 \right. \\ &\qquad \qquad \quad+ \theta_{30}^2 \left( \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right)^2 f_t(0) \int K'(u)^2 \mathrm{d}u \\ &\qquad \qquad \quad \left. + 2 \theta_{10} \theta_{20}\left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right) \left( \mathbf{h}_t^\top w_{\mathrm{Med}} \lambda \right) \varepsilon_t \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big) \right]. \end{aligned} \end{align} Lemma (ref) shows that $\sum_{t=1}^T \operatorname{Var} \big(T^{-1/2} \phi_{t,T}^\ast (\theta_0) \lambda \big) - \sigma_T^2 \to 0$. As $\sigma^2 := \lambda^\top \Sigma(\theta_0) \lambda$, which is strictly positive by assumption, is the limit of $\sigma_T^2$, the latter must be strictly positive for $T$ large enough and consequently, $\sigma_T^{-1} T^{-1/2} \sum_{t=1}^T \phi_{t,T}^\ast (\theta_0) \lambda$ is well-defined. In the following, we first show that $\sigma_T^{-1} T^{-1/2} \sum_{t=1}^T \phi_{t,T}^\ast (\theta_0) \lambda \overset{d}{\to} \mathcal{N} \big( 0, 1 \big)$ by applying Theorem 24.3 in davidson1994stochastic and by verifying that the respective conditions hold (with $X_{t,T} = \sigma_T^{-1} T^{-1/2} \phi_{t,T}^\ast (\theta_0) \lambda$). Lemma (ref) shows that $T^{-1} \sum_{t=1}^T \big( \phi_{t,T}^\ast (\theta_0) \lambda \big)^2 - \sigma_T^{2} \overset{P}{\to} 0$, which implies condition (a) of Theorem 24.3 of davidson1994stochastic, i.e.\ $\sigma_T^{-2} T^{-1} \sum_{t=1}^T \big( \phi_{t,T}^\ast (\theta_0) \lambda \big)^2 \overset{P}{\to} 1$. Lemma (ref) shows condition (b), i.e.\ $\max_{t=1,\dots,T} \big| \sigma_T^{-1} T^{-1/2} \phi_{t,T}^\ast (\theta_0) \lambda \big| \overset{P}{\to} 0$. Thus, we can apply Theorem 24.3 of davidson1994stochastic and conclude that $\sigma_T^{-1} T^{-1/2} \sum_{t=1}^T \phi_{t,T}^\ast (\theta_0) \lambda \overset{d}{\to} \mathcal{N} \big( 0, 1 \big)$. As $\sigma_T^2$ has the limit $\sigma^2$, Slutsky's theorem implies that $T^{-1/2} \sum_{t=1}^T \phi_{t,T}^\ast (\theta_0) \lambda \overset{d}{\to} \mathcal{N}(0, \sigma^2)$ and as this holds for all $\lambda \in \mathbb{R}^k$ such that $||\lambda||_2 = 1$, we can conclude that $T^{-1/2} \sum_{t=1}^T \phi_{t,T}^\ast (\theta_0) \overset{d}{\to} \mathcal{N} \big( 0, \Sigma(\theta_0) \big)$. Furthermore, $ || T^{-1/2} \sum_{t=1}^T (\phi_{t,T}^\ast (\theta_0) - \phi_{t,T} (\theta_0)) || = || T^{-1/2} \sum_{t=1}^T u_{t,T}(\theta_0) || \overset{P}{\to} 0$ by Assumption (ref) (D) and \begin{align*} T^{-1/2} \sum_{t=1}^T \left( \widehat \phi_{t,T} (\theta_0) - \phi_{t,T} (\theta_0) \right) = T^{-1/2} \sum_{t=1}^T \theta_0 \cdot \begin{pmatrix} \mathbf{h}_t^\top \left( \widehat{w}_{T,\mathrm{Mean}} - w_{\mathrm{Mean}} \right) \varepsilon_t \\ \mathbf{h}_t^\top \left( \widehat{w}_{T,\mathrm{Med}} - w_{\mathrm{Med}} \right) \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big) \\ \mathbf{h}_t^\top \left( \widehat{w}_{T,\mathrm{Mode}} - w_{\mathrm{Mode}} \right) \delta_T^{-1/2} K' \left( \frac{- \varepsilon_t}{\delta_T} \right) \end{pmatrix} \overset{P}{\to} 0, \end{align*} as it holds that $\widehat{w}_{T,\mathrm{Mean}} \overset{P}{\to} w_{\mathrm{Mean}}$, $\widehat{w}_{T,\mathrm{Med}} \overset{P}{\to} w_{\mathrm{Med}}$ and $\widehat{w}_{T,\mathrm{Mode}} \overset{P}{\to} w_{\mathrm{Mode}}$ by assumption. Hence we can conclude that $T^{-1/2} \sum_{t=1}^T \widehat \phi_{t,T} (\theta_0) \overset{d}{\to} \mathcal{N} \big( 0, \Sigma(\theta_0) \big)$, which concludes the proof of this theorem.
proof[Proof of Theorem (ref)] For notational simplicity, we show consistency of the covariance estimator by considering the bilinear forms $\lambda^\top \left( \frac{1}{T} \sum_{t=1}^T \widehat \phi_{t,T} (\theta_0) \widehat \phi_{t,T} (\theta_0)^\top \right) \lambda$ and $\sigma_T^{2} = \lambda^\top \Sigma_T(\theta_0) \lambda$, given in ((ref)), for some arbitrary but fixed $\lambda \in \mathbb{R}^k$ such that $||\lambda||_2 = 1$. Then, we get that \begin{align} &\lambda^\top \left( \frac{1}{T} \sum_{t=1}^T \widehat \phi_{t,T} (\theta_0) \widehat \phi_{t,T} (\theta_0)^\top \right) \lambda = \frac{1}{T} \sum_{t=1}^T \left\{ \theta_{10}^2 \left( \mathbf{h}_t^\top \widehat{w}_{T,\mathrm{Mean}} \lambda \right)^2 \varepsilon_t^2 \right. \\ &\qquad+ \theta_{20}^2 \left( \mathbf{h}_t^\top \widehat{w}_{T,\mathrm{Med}} \lambda \right)^2 \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big)^2 + \theta_{30}^2 \left( \mathbf{h}_t^\top \widehat{w}_{T,\mathrm{Mode}} \lambda \right)^2 \delta_T^{-1} K' \left( \frac{-\varepsilon_t}{\delta_T} \right)^2 \\ &\qquad+ 2 \theta_{10} \theta_{20}\left( \mathbf{h}_t^\top \widehat{w}_{T,\mathrm{Mean}} \lambda \right) \left( \mathbf{h}_t^\top \widehat{w}_{T,\mathrm{Med}} \lambda \right) \varepsilon_t \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big) \\ &\qquad+ 2 \theta_{10} \theta_{30} \left( \mathbf{h}_t^\top \widehat{w}_{T,\mathrm{Mean}} \lambda \right) \left( \mathbf{h}_t^\top \widehat{w}_{T,\mathrm{Mode}} \lambda \right) \varepsilon_t \ \delta_T^{-1/2} K' \left( \frac{-\varepsilon_t}{\delta_T} \right) \\ &\qquad \left.+ 2 \theta_{20} \theta_{30} \left( \mathbf{h}_t^\top \widehat{w}_{T,\mathrm{Med}} \lambda \right) \left( \mathbf{h}_t^\top \widehat{w}_{T,\mathrm{Mode}} \lambda \right) \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big) \ \delta_T^{-1/2} K' \left( \frac{-\varepsilon_t}{\delta_T} \right) \right\}. \end{align} We show convergence in probability for the individual matrix components for the first term, \begin{align} &\frac{1}{T} \sum_{t=1}^T \theta_{10}^2 \left( \mathbf{h}_t^\top \widehat{w}_{T,\mathrm{Mean}} \lambda \right)^2 \varepsilon_t^2 - \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{10}^2 \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right)^2 \varepsilon_t^2 \right] \overset{P}{\to} 0 \end{align} Convergence of the remaining terms follows analogously by considering the terms component-wisely and by applying similar arguments as in Lemma (ref).

Technical Lemmas and Their Proofs

lemmaGiven Assumption (ref) and the null hypothesis in ((ref)), it holds that $\sum_{t=1}^T g_{t,T}^e \overset{P}{\to} 0$.
proofApplying integration by parts yields that \begin{align} g_{t,T}^e &= - \mathbb{E}_t \left[ (T \delta_T)^{-1/2} K' \left( \frac{ \varepsilon_t }{\delta_T} \right) \mathbf{h}_t \right] = - (T \delta_T)^{-1/2} \mathbf{h}_t \int K' \left( \frac{ e }{\delta_T} \right) f_{t} (e) \, \mathrm{d}e \\ &= T^{-1/2} \delta_T^{1/2} \mathbf{h}_t \int K \left( \frac{ e }{\delta_T} \right) f'_{t} (e) \, \mathrm{d}e - T^{-1/2} \delta_T^{1/2} \mathbf{h}_t \left[ K \left( \frac{ e }{\delta_T} \right) f_{t} (e) \right]_{e=-\infty}^{e=\infty}. \end{align} As $\lim_{e \to \pm \infty} K(e) = 0$ and $f_t$ is bounded from above, the latter term is zero a.s.\ for all $t \le T$. By transformation of variables, it further holds that \begin{align} g_{t,T}^e = T^{-1/2} \delta_T^{1/2} \mathbf{h}_t \int K \left( \frac{ e }{\delta_T} \right) f'_{t} (e) \, \mathrm{d}e = T^{-1/2} \delta_T^{3/2} \mathbf{h}_t \int K \left( u \right) f'_{t} (\delta_T u) \, \mathrm{d}u. \end{align} A Taylor expansion of $f'_t (\delta_T u)$ around zero is given by \begin{align} f'_{t} (\delta_T u) = f'_{t} (0) +(\delta_T u) f”_{t} (0) +\frac{(\delta_T u)^2}{2} f”'_{t} (\zeta \delta_T u), \end{align} for some $\zeta \in [0,1]$ and $f'_{t} (0) = 0$ holds under the null hypothesis specified in ((ref)). Consequently, \begin{align} \sum_{t=1}^T g_{t,T}^e = \, & T^{-1/2} \delta_T^{5/2} \sum_{t=1}^T f”_{t} (0) \mathbf{h}_t \int u K \left( u \right) \, \mathrm{d}u \\ &+ \frac{1}{2} T^{-1/2} \delta_T^{7/2} \sum_{t=1}^T \mathbf{h}_t \int u^2 K \left( u \right) f”'_{t} (\zeta \delta_T u) \, \mathrm{d}u. \end{align} As $\int u K \left( u \right) \, \mathrm{d}u = 0 $ by assumption (ref), the first term is zero for all $T \in \mathbb{N}$. As $\sup_x f'''_{t}(x) \le c$ by Assumption (ref) and $\int u^2 K \left( u \right) \mathrm{d}u \le c < \infty$ by Assumption (ref), we obtain \begin{align} \frac{1}{2} T^{-1/2} \delta_T^{7/2} \sum_{t=1}^T \mathbf{h}_t \int u^2 K \left( u \right) f”'_{t} (\zeta \delta_T u) \, \mathrm{d}u \le \frac{1}{2} c^2 \big( T \delta_T^{7} \big)^{1/2} \frac{1}{T} \sum_{t=1}^T \mathbf{h}_t \overset{P}{\to} 0, \end{align} as $T \delta_T^7 \to 0$ for $T \to \infty$ by Assumption (ref) and $ \frac{1}{T} \sum_{t=1}^T \mathbf{h}_t \overset{P}{\to} \mathbb{E}[\mathbf{h}_t]$ by a law of large numbers for $\alpha$-mixing sequences White2001. The result of the lemma follows.
lemmaGiven Assumption (ref), it holds that $\sum_{t=1}^T \operatorname{Var} \left( z_{t,T} \right) \to \omega^2 = \lambda^\top \Omega_{\mathrm{Mode}} \lambda$.
proofWe first observe that $ \operatorname{Var} \left( z_{t,T} \right) = \mathbb{E} \left[ \big( \lambda^\top \big( g_{t,T} - g_{t,T}^e \big) \big)^2 \right]$ as $\mathbb{E} \left[ \lambda^\top \big( g_{t,T} - g_{t,T}^e \big) \right] = 0$. Hence, \begin{align} \operatorname{Var} \left( z_{t,T} \right) = \mathbb{E} \left[ \left( \lambda^\top g_{t,T} \right)^2 \right] - \mathbb{E} \left[ \left( \lambda^\top g_{t,T}^e \right)^2 \right], \end{align} as $\mathbb{E} \left[ \big( \lambda^\top g_{t,T}^e \big) \cdot \big( \lambda^\top g_{t,T} \big) \right] = \mathbb{E} \left[ \big( \lambda^\top g_{t,T}^e \big) \cdot \mathbb{E}_t \big[ \lambda^\top g_{t,T} \big] \right] = \mathbb{E} \left[ \big( \lambda^\top g_{t,T}^e \big)^2 \right]$. For the first term in ((ref)), we obtain \begin{align} \mathbb{E} \left[ \left( \lambda^\top g_{t,T} \right)^2 \right] &= \mathbb{E} \left[ (T \delta_T)^{-1} (\lambda^\top \mathbf{h}_t)^2 \mathbb{E}_t \left[ K'\left( \frac{X_t-Y_{t+1}}{\delta_T} \right)^2 \right] \right] \\ &= \mathbb{E} \left[ (T \delta_T)^{-1} (\lambda^\top \mathbf{h}_t)^2 \int K'\left( \frac{e}{\delta_T} \right)^2 f_t(e) \, \mathrm{d}e \right] \\ &= \frac{1}{T} \, \mathbb{E} \left[ (\lambda^\top \mathbf{h}_t)^2 \int K' \left( u \right)^2 f_{t} (\delta_T u) \, \mathrm{d}u \right]. \end{align} As $\delta_T \to 0$ when $T \to \infty$ and as $\Omega_{T,\mathrm{Mode}} \to \Omega_{\mathrm{Mode}}$ from Assumption (ref), we get \begin{align} \begin{aligned} &\sum_{t=1}^T \mathbb{E} \left[ \left( \lambda^\top g_{t,T} \right)^2 \right] - \lambda^\top \Omega_{\mathrm{Mode}} \lambda \\ =\,&\sum_{t=1}^T \mathbb{E} \left[ \left( \lambda^\top g_{t,T} \right)^2 \right] - \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ (\lambda^\top \mathbf{h}_t )^2 f_{t} (0) \right] \int K'\left( u \right)^2 \mathrm{d}u \\ &\qquad + \lambda^\top \Omega_{T,\mathrm{Mode}} \lambda - \lambda^\top \Omega_{\mathrm{Mode}} \lambda \\ = \;& \frac{1}{T} \sum_{t=1}^T \left(\mathbb{E} \left[ (\lambda^\top \mathbf{h}_t)^2 \int K' \left( u \right)^2 f_{t} (\delta_T u) \, \mathrm{d}u \right] - \lambda^\top \mathbb{E} \left[ f_{t} (0) \mathbf{h}_t \mathbf{h}_t^\top \right] \lambda \int K'\left( u \right)^2 \mathrm{d}u \right) + o(1) \to 0. \end{aligned} \end{align} For the second term in ((ref)), inserting the equality in ((ref)) yields \begin{align*} \left( \lambda^\top g_{t,T}^e \right)^2 = \left( \delta_T^{3/2} T^{-1/2} (\lambda^\top \mathbf{h}_t) \int K'\left( u \right) f'_{t} (\delta_T u) \mathrm{d}u \right)^2 \le \delta_T^3 T^{-1} ||\lambda||^2 ||\mathbf{h}_t||^2 \left| \int K'\left( u \right) f'_{t} (u \delta_T) \mathrm{d}u \right|^2. \end{align*} As $\sup_x | f'_{t} (x) | \le c$ for all $t \in \mathbb{N}$ and $\left| \int K'\left( u \right) \mathrm{d}u \right| \le c$ by assumption, it holds that \begin{align} \sum_{t=1}^T \mathbb{E} \left[ \left( \lambda^\top g_{t,T}^e\right)^2 \right] &\le \delta_T^3 c^2 ||\lambda||^2 \left( \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ ||\mathbf{h}_t||^2 \right] \right) \to 0, \end{align} as $\delta_T^3 \to 0 $ as $T \to \infty$. The result of the lemma follows by combining ((ref)) and ((ref)).
lemmaGiven Assumption (ref), it holds that $\sum_{t=1}^T z_{t,T}^2 \overset{P}{\to} \omega^2 = \lambda^\top \Omega_{\mathrm{Mode}} \lambda$.
proofWe define \begin{align} h_{1,T} := \sum_{t=1}^T \left( z_{t,T}^2 - \mathbb{E}_t \left[ z_{t,T}^2 \right] \right) \qquad and \qquad h_{2,T} := \sum_{t=1}^T \mathbb{E}_t \left[ z_{t,T}^2 \right] - \omega^2, \end{align} such that $\sum_{t=1}^T z_{t,T}^2 - \bar \omega^2 = h_{1,T} + h_{2,T}$. We first show that $h_{1,T} \stackrel{L_p}{\longrightarrow} 0$ for some $ 1 < p < 2$ sufficiently small enough and thus $h_{1,T} \overset{P}{\to} 0$. For this, first notice that $z_{t,T}^2 - \mathbb{E}_t \left[ z_{t,T}^2 \right]$ is a $\mathcal{F}_{t+1}$-MDA by definition. Thus, we can apply the BahrEsseen1965-inequality for some $p \in (1,2)$ for MDA (in the first line) in order to conclude that \begin{align} \mathbb{E} \left[ \left| h_{1,T} \right|^p \right] &= \mathbb{E} \left[ \left| \sum_{t=1}^T z_{t,T}^2 - \mathbb{E}_t \left[ z_{t,T}^2 \right] \right|^p \right] \le 2 \sum_{t=1}^T \mathbb{E} \left[ \left| z_{t,T}^2 - \mathbb{E}_t \left[ z_{t,T}^2 \right] \right|^p \right] \\ &\le 2 \sum_{t=1}^T 2^{p-1} \left( \mathbb{E} \left[ \left| z_{t,T}^2 \right|^p \right] + \mathbb{E} \left[ \left| \mathbb{E}_t \left[ z_{t,T}^2 \right] \right|^p \right] \right) \le 2^{p+1} \sum_{t=1}^T \mathbb{E} \left[ \left| z_{t,T} \right|^{2p} \right], \end{align} where we use in the second inequality that $|a+b|^p \le 2^{p-1} \big( |a|^p + |b|^p \big)$ for any $a,b \in \mathbb{R}$. Using the same inequality again yields \begin{align} \mathbb{E} \left[ \left| z_{t,T} \right|^{2p} \right] = \mathbb{E} \left[ \left| \lambda^\top g_{t,T} - \lambda^\top g_{t,T}^e \right|^{2p} \right] \le 2^{2p-1} \left( \mathbb{E} \left[ \left| \lambda^\top g_{t,T} \right|^{2p} \right] + \mathbb{E} \left[ \left| \lambda^\top g_{t,T}^e \right|^{2p} \right] \right). \end{align} Thus, \begin{align} &\sum_{t=1}^T \mathbb{E} \left[ \left| \lambda^\top g_{t,T} \right|^{2p} \right] = (T \delta_T)^{1-p} \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \left| \lambda^\top \mathbf{h}_t \right|^{2p} \int \left| K'\left( u \right) \right|^{2p} f_{t}(\delta_T u) \mathrm{d}u \right] \to 0, \end{align} as $(T \delta_T)^{1-p} \to 0$ for all $p \in (1,2)$, $\mathbb{E} \left[ \left| \lambda^\top \mathbf{h}_t \right|^{2p} \right] < \infty$ for $p > 1$ sufficiently small such that $2p \le 2 + \delta$ (for the $\delta > 0$ from Assumption (ref)), and $\int \left| K'\left( u \right) \right|^{2p} f_{t}(\delta_T u) \mathrm{d}u \le c c^{2p-1} \int \left| K'\left( u \right) \right| \mathrm{d}u < \infty$ as $\int \left| K'\left( u \right) \right| \mathrm{d}u < \infty$, $\sup_{u} \left| K'\left( u \right) \right| \le c$ and $\sup_{x} f_{t}(x) \le c$ a.s.\ by assumption. Similarly, we obtain \begin{align} &\sum_{t=1}^T \mathbb{E} \left[ \left| \lambda^\top g_{t,T}^e \right|^{2p} \right] = \delta_T^{2p-1} (T \delta_T)^{1-p} \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \left| \lambda^\top \mathbf{h}_t \right|^{2p} \left| \int K'\left( u \right) f_{t}(\delta_T u) \mathrm{d}u \right|^{2p} \right] \to 0, \end{align} which shows that $h_{1,T} \stackrel{L_p}{\longrightarrow} 0$ for some $p > 1$ sufficiently small which implies that $h_{1,T} \overset{P}{\to} 0$. We continue by showing that $h_{2,T} \overset{P}{\to} 0$. Using the same argument as in ((ref)), we split \begin{align} h_{2,T} = \sum_{t=1}^T \mathbb{E}_t \left[ z_{t,T}^2 \right] - \omega^2 = \sum_{t=1}^T \mathbb{E}_t \left[ (\lambda^\top g_{t,T})^2 \right] - \sum_{t=1}^T (\lambda^\top g_{t,T}^e)^2 - \omega^2. \end{align} Applying a transformation of variables yields \begin{align} \sum_{t=1}^T (\lambda^\top g_{t,T}^e)^2 &= \delta_T \frac{1}{T} \sum_{t=1}^T (\lambda^\top \mathbf{h}_t)^2 \left( \int K'(u) f_{t}(\delta_T u) \mathrm{d}u \right)^2 \\ &\le \delta_T \left( \frac{1}{T} \sum_{t=1}^T (\lambda^\top \mathbf{h}_t)^2 \right) \left( c \int |K'(u)| \mathrm{d}u \right)^2 \overset{P}{\to} 0, \end{align} as $\delta_T \to 0$, $ \frac{1}{T} \sum_{t=1}^T (\lambda^\top \mathbf{h}_t)^2 = \mathbb{E} \left[ (\lambda^\top \mathbf{h}_t)^2 \right] + o_P(1)$ and $\left(\int |K'(u)| \mathrm{d}u \right)^2 \le \int |K'(u)|^2 \mathrm{d}u < \infty$ by assumption. Furthermore, \begin{align} &\sum_{t=1}^T \mathbb{E}_t \left[ \big( \lambda^\top g_{t,T} \big)^2 \right] - \omega^2 \\ = \, &\left( \frac{1}{T} \sum_{t=1}^T (\lambda^\top \mathbf{h}_t)^2 \right) \int K'(u)^2 f_t(\delta_T u) \mathrm{d}u - \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ f_{t}(0) (\lambda^\top \mathbf{h}_t)^2 \right] \int K'(u)^2 \mathrm{d}u + \bar \omega_T^2 - \omega^2 \overset{P}{\to} 0, \end{align} as for all $u \in \mathbb{R}$, $\frac{1}{T} \sum_{t=1}^T (\lambda^\top \mathbf{h}_t)^2 f_t(\delta_T u) - \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ f_{t}(0) (\lambda^\top \mathbf{h}_t)^2 \right] \overset{P}{\to} 0$. Thus, we find $h_{2,T} \overset{P}{\to} 0$ and consequently $\sum_{t=1}^T z_{t,T}^2 - \omega^2 \overset{P}{\to} 0$, which concludes this proof.
lemmaGiven Assumption (ref), it holds that $\max_{1\le t \le T} | h_{t,T} | \overset{P}{\to} 0$.
proofLet $\zeta > 0$ and $\delta > 0$ (sufficiently small such that $\mathbb{E} \left[ ||\mathbf{h}_t||^{2+\delta} \right] < \infty$). Then, \begin{align} \begin{aligned} \mathbb{P} \left( \max_{1\le t \le T} | h_{t,T} | > \zeta \right) &= \mathbb{P} \left( \max_{1\le t \le T} | h_{t,T} |^{2+\delta} > \zeta^{2+\delta} \right) \le \sum_{t=1}^T \mathbb{P} \left( |h_{t,T}|^{2+\delta} > \zeta^{2+\delta} \right) \\ &\le \zeta^{-2-\delta} \sum_{t=1}^T \mathbb{E} \left[ |h_{t,T}|^{2+\delta} \right] = \zeta^{-2-\delta} \bar \omega_T^{-2-\delta} \sum_{t=1}^T \mathbb{E} \left[ |z_{t,T}|^{2+\delta} \right], \end{aligned} \end{align} by Markov's inequality. Employing the same steps as in the proof of Lemma (ref) following Equation ((ref)) and replacing the exponent “$2p$” by “$2+ \delta$” yields that $\sum_{t=1}^T \mathbb{E} \left[ |z_{t,T}|^{2+\delta} \right] \to 0$. As $\bar \omega_T \to \omega^2 > 0$, this directly implies that $\mathbb{P} \left( \max_{1\le t \le T} | h_{t,T} | > \zeta \right) \to 0$.
lemmaGiven Assumption (ref) and Assumption (ref), for all $\lambda \in \mathbb{R}^k$ such that $||\lambda||_2 = 1$, it holds that $\sum_{t=1}^T \operatorname{Var} \big( T^{-1/2} \phi_{t,T}^\ast (\theta_0)^\top \lambda \big) - \sigma_T^{2} \to 0$.
proofAs $T^{-1/2} \phi_{t,T}^\ast(\theta_0)$ is a $\mathcal{F}_{t+1}$-MDA, it holds that $\mathbb{E} \big[ T^{-1/2} \phi_{t,T}^\ast (\theta_0)^\top \lambda \big] = 0$ and thus, $\operatorname{Var} \big( T^{-1/2} \phi_{t,T}^\ast (\theta_0)^\top \lambda \big) = \mathbb{E} \left[ \big( T^{-1/2} \phi_{t,T}^\ast (\theta_0)^\top \lambda \big)^2 \right]$. We further have \begin{align} & \quad \sum_{t=1}^T \mathbb{E} \left[ \big( T^{-1/2} \phi_{t,T}^\ast (\theta_0)^\top \lambda \big)^2 \right] \\ &= \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{10}^2 \left( \lambda^\top w_{\mathrm{Mean}} \mathbf{h}_t \right)^2 \varepsilon_t^2 \right] + \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{20}^2 \left( \lambda^\top w_{\mathrm{Med}} \mathbf{h}_t \right)^2 \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big)^2 \right] \nonumber \\ &+ \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{30}^2 \left( \lambda^\top w_{\mathrm{Mode}} \mathbf{h}_t \right)^2 \delta_T^{-1} K' \left( \frac{- \varepsilon_t}{\delta_T} \right)^2 \right] \nonumber \\ &+ \frac{2}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{10} \theta_{20} \left( \lambda^\top w_{\mathrm{Mean}} \mathbf{h}_t \right) \left( \lambda^\top w_{\mathrm{Med}} \mathbf{h}_t \right) \varepsilon_t \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big) \right] \nonumber \\ &+ \frac{2}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{10} \theta_{30} \left( \lambda^\top w_{\mathrm{Mean}} \mathbf{h}_t \right) \left( \lambda^\top w_{\mathrm{Mode}} \mathbf{h}_t \right) \varepsilon_t \delta_T^{-1/2} K' \left( \frac{- \varepsilon_t}{\delta_T} \right) \right] \nonumber \\ &+ \frac{2}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{20} \theta_{30} \left( \lambda^\top w_{\mathrm{Med}}\mathbf{h}_t \right) \left( \lambda^\top w_{\mathrm{Mode}} \mathbf{h}_t \right) \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big) \delta_T^{-1/2} K' \left( \frac{- \varepsilon_t}{\delta_T} \right) \right] \nonumber \\ &+\frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ (\lambda^\top u_{t,T}(\theta_0))^2 \right] - \frac{2}{T} \sum_{t=1}^T \mathbb{E} \left[ \big( \lambda^\top u_{t,T}(\theta_0)\big) \big( \lambda^\top \phi_{t,T}(\theta_0) \big) \right]. \nonumber \end{align} For the last two terms, we have $T^{-1} \sum_{t=1}^T \mathbb{E} \left[ \big( \lambda^\top u_{t,T}(\theta_0) \big)^2 \right] \to 0$ and $T^{-1} \sum_{t=1}^T \mathbb{E} \left[ \big( \lambda^\top u_{t,T}(\theta_0) \big) \big( \lambda^\top \phi_{t,T}(\theta_0) \big) \right] \to 0$ by Assumption (ref) (D). For the fifth term, \begin{align} &\frac{2}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{10} \theta_{30} \left( \lambda^\top w_{\mathrm{Mean}} \mathbf{h}_t \right) \left( \lambda^\top w_{\mathrm{Mode}} \mathbf{h}_t \right) \delta_T^{-1/2} \mathbb{E}_t \left[ \varepsilon_t K' \left( \frac{- \varepsilon_t}{\delta_T} \right) \right] \right] \\ = \, &- \frac{2}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{10} \theta_{30} \left( \lambda^\top w_{\mathrm{Mean}} \mathbf{h}_t \right) \left( \lambda^\top w_{\mathrm{Mode}} \mathbf{h}_t \right) \delta_T^{3/2} \int u K' \left( u \right)f_t(\delta_T u) \mathrm{d}u \right] \to 0, \end{align} as $\delta_T^{3/2} \to 0$, $\int u K' \left( u \right) \mathrm{d}u < \infty$ and the respective moments are finite. The sixth term converges to zero by a similar argument by bounding $\big| \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big| \le 1$. For the third term, it holds that \begin{align*} &\frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{30}^2 \left( \lambda^\top w_{\mathrm{Mode}} \mathbf{h}_t \right)^2 \delta_T^{-1} K' \left( \frac{- \varepsilon_t}{\delta_T} \right)^2 \right] - \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{30}^2 \left( \lambda^\top w_{\mathrm{Mode}} \mathbf{h}_t \right)^2 f_t(0) \int K' \left( u \right)^2 \mathrm{d}u \right] \\ = \, &\frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{30}^2 \left( \lambda^\top w_{\mathrm{Mode}} \mathbf{h}_t \right)^2 \int K' \left( u \right)^2 \big( f_t(\delta_T u) - f_t(0) \big) \mathrm{d}u \right] \to 0 \end{align*} The remaining first, second and fourth terms in (ref) appear equivalently in $\sigma_T^2$, which concludes the proof of this lemma.
lemmaGiven Assumption (ref) and Assumption (ref), for all $\lambda \in \mathbb{R}^k$ such that $||\lambda||_2 = 1$, it holds that $T^{-1} \sum_{t=1}^T \big( \phi_{t,T}^\ast (\theta_0)^\top \lambda \big)^2 - \sigma_T^{2} \overset{P}{\to} 0$.
proofWe apply the same factorization as in ((ref)) (however without the expectation operator). By applying a law of large numbers for mixing sequences (Corollary 3.48 in White2001), we obtain that \begin{align*} &\frac{1}{T} \sum_{t=1}^T \left( \theta_{10} \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right) \varepsilon_t \right)^2 - \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \left( \theta_{10} \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right) \varepsilon_t \right)^2 \right] \overset{P}{\to} 0, \\ &\frac{1}{T} \sum_{t=1}^T \left( \theta_{20} \left( \mathbf{h}_t^\top w_{\mathrm{Med}} \lambda \right) \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big) \right)^2 - \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \left( \theta_{20} \left( \mathbf{h}_t^\top w_{\mathrm{Med}} \lambda \right) \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big) \right)^2 \right] \overset{P}{\to} 0, \\ &\frac{2}{T} \sum_{t=1}^T \left( \theta_{10} \theta_{20} \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right) \left( \mathbf{h}_t^\top w_{\mathrm{Med}} \lambda \right) \varepsilon_t \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big) \right)^2 \\ &\qquad - \frac{2}{T} \sum_{t=1}^T \mathbb{E} \left[ \left( \theta_{10} \theta_{20} \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right) \left( \mathbf{h}_t^\top w_{\mathrm{Med}} \lambda \right) \varepsilon_t \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big) \right)^2 \right] \overset{P}{\to} 0. \end{align*} Furthermore, a slight modification of Lemma (ref) (multiplying with $\theta_{30}^2 \big( \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \big)^2$ instead of $ \big( \mathbf{h}_t^\top \lambda \big)^2$) yields that \begin{align} \frac{1}{T} \sum_{t=1}^T \theta_{30}^2 \left( \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right)^2 \delta_T^{-1} K' \left( \frac{-\varepsilon_t}{\delta_T} \right)^2 - \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \theta_{30}^2 \left( \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right)^2 f_t(0) \int K' \left( u \right)^2 \mathrm{d} u \right] \overset{P}{\to} 0. \end{align} We now show that the remaining four terms vanish asymptotically (in probability). For the mixed mean/mode term, we apply a similar addition of a zero (adding and subtracting $\mathbb{E}_t[\dots]$) as in the proof of Lemma (ref). For this, we first note that \begin{align} &\frac{2}{T} \sum_{t=1}^T\theta_{10} \theta_{30} \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right) \left( \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right) \delta_T^{-1/2} \mathbb{E}_t \left[ \varepsilon_t K' \left( \frac{-\varepsilon_t}{\delta_T} \right) \right] \\ = -\,&\frac{2}{T} \sum_{t=1}^T\theta_{10} \theta_{30} \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right) \left( \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right) \delta_T^{3/2} \int u K' \left( u \right) f_t(\delta_T u) \mathrm{d}u \overset{P}{\to} 0, \end{align} as $\delta_T^{3/2} \to 0$. In the following, we further show that \begin{align*} \frac{2}{T} \sum_{t=1}^T\theta_{10} \theta_{30} \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right) \left( \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right) \delta_T^{-1/2} \left\{ \varepsilon_t K' \left( \frac{-\varepsilon_t}{\delta_T} \right) - \mathbb{E}_t \left[ \varepsilon_t K' \left( \frac{-\varepsilon_t}{\delta_T} \right) \right] \right\} \stackrel{L_p}{\longrightarrow} 0, \end{align*} for any $p \in (1,2)$ small enough. As in the proof of Lemma (ref), we apply the BahrEsseen1965 inequality and Minkowski's inequality in order to conclude that \begin{align*} &\mathbb{E} \left[ \left| \frac{2}{T} \sum_{t=1}^T\theta_{10} \theta_{30} \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right) \left( \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right) \delta_T^{-1/2} \left\{ \varepsilon_t K' \left( \frac{\varepsilon_t}{\delta_T} \right) - \mathbb{E}_t \left[ \varepsilon_t K' \left( \frac{\varepsilon_t}{\delta_T} \right) \right] \right\} \right|^p \right] \\ \le \, & \frac{2^{p+2}}{T^p} \sum_{t=1}^T \left\{ \mathbb{E} \left[ \left| \theta_{10} \theta_{30} \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right) \left( \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right) \delta_T^{-1/2} \varepsilon_t K' \left( \frac{\varepsilon_t}{\delta_T} \right) \right|^p \right] \right. \\ &\qquad \qquad \; + \left. \mathbb{E} \left[ \left| \theta_{10} \theta_{30} \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right) \left( \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right) \delta_T^{-1/2} \mathbb{E}_t \left[ \varepsilon_t K' \left( \frac{\varepsilon_t}{\delta_T} \right) \right] \right|^p \right] \right\}, \end{align*} where the first term is bounded from above by \begin{align*} & \frac{2^{p+2}}{T^p} \sum_{t=1}^T \mathbb{E} \left[ \left| \theta_{10} \theta_{30} \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right) \left( \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right) \right|^p \delta_T^{-p/2} \int \left| e K' \left( \frac{e}{\delta_T} \right) \right|^p f_t(e) \mathrm{d}e \right] \\ = \, & 2^{p+2} \delta_T^{1+p/2} T^{1-p} \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \left| \theta_{10} \theta_{30} \left( \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right) \left( \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right) \right|^p \int \left| u K' \left( u \right) \right|^p f_t(\delta_T u) \mathrm{d}u \right] \to 0, \end{align*} as $\delta_T^{1+p/2}\to 0 $, $T^{1-p} \to 0$ for any $p > 1$ and as the respective moments are bounded by assumption. Similar arguments also yield that the second term converges to zero (compare to ((ref))). Applying the same line of reasoning for the mixed median/mode terms shows that \begin{align} \frac{2}{T} \sum_{t=1}^T\theta_{20} \theta_{30} \left( \mathbf{h}_t^\top w_{\mathrm{Med}} \lambda \right) \left( \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right) \delta_T^{-1/2} \mathbb{E}_t \left[ \big( \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big) K' \left( \frac{-\varepsilon_t}{\delta_T} \right) \right] \overset{P}{\to} 0. \end{align} For the fourth and last term, $T^{-1} \sum_{t=1}^T \big( u_{t,T}(\theta_0)^\top \lambda \big)^2 \overset{P}{\to} 0$ and $T^{-1} \sum_{t=1}^T \big( u_{t,T}(\theta_0)^\top \lambda \big) \big( \phi_{t,T}(\theta_0)^\top \lambda \big) \overset{P}{\to} 0$ by Assumption (ref) (D), which concludes this proof.
lemmaGiven Assumption (ref) and Assumption (ref), for all $\lambda \in \mathbb{R}^k$ such that $||\lambda||_2 = 1$, it holds that $\max_{1 \le t \le T} \big| \sigma_T^{-1} T^{-1/2} \phi_{t,T}^\ast (\theta_0)^\top \lambda \big| \overset{P}{\to} 0$.
proofLet $\zeta > 0$ and $\delta > 0$ (sufficiently small such that $\mathbb{E} \left[ ||\mathbf{h}_t||^{2+\delta} \right] < \infty$ holds). Then, as in ((ref)) in the proof of Lemma (ref), we get that \begin{align} \mathbb{P} \left( \max_{1\le t \le T} \big| \sigma_T^{-1} T^{-1/2} \phi_{t,T}^\ast (\theta_0)^\top \lambda \big| > \zeta \right) \le \zeta^{-2-\delta} \sigma_T^{-2-\delta} \sum_{t=1}^T \mathbb{E} \left[ \big| T^{-1/2} \phi_{t,T}^\ast (\theta_0)^\top \lambda \big|^{2+\delta} \right], \end{align} by Markov's inequality. Furthermore, we get that \begin{align} 4^{-2-\delta} \sum_{t=1}^T \mathbb{E} \left[ \big| T^{-1/2} \phi_{t,T}^\ast (\theta_0)^\top \lambda \big| ^{2+\delta} \right] & \le \, \sum_{t=1}^T \mathbb{E} \left[ \left| T^{-1/2} u_{t,T}(\theta_0)^\top \lambda \right|^{2+\delta} \right] \\ & + \theta_{10}^{2+\delta} T^{-\frac{\delta}{2}} \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \left| \mathbf{h}_t^\top w_{\mathrm{Mean}} \lambda \right|^{2+\delta} |\varepsilon_t|^{2+\delta} \right] \\ & + \theta_{20}^{2+\delta} T^{-\frac{\delta}{2}} \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \left| \mathbf{h}_t^\top w_{\mathrm{Med}} \lambda \right|^{2+\delta} \big| \mathds{1}_{\{\varepsilon_t > 0\}} - \mathds{1}_{\{\varepsilon_t < 0 \}} \big|^{2+\delta} \right] \\ & + \theta_{30}^{2+\delta} T^{-\frac{\delta}{2}} \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \left| \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right|^{2+\delta} \delta_T^{-\frac{2+\delta}{2}} \left| K' \left( \frac{-\varepsilon_t}{\delta_T} \right) \right|^{2+\delta} \right]. \end{align} The first term converges to zero by Assumption (ref). The second and third term converge to zero as $T^{-\frac{\delta}{2}} \to 0$ and the respective moments are bounded by assumption. For the last term, we obtain convergence equivalently to the proof of Lemma (ref), \begin{align} &\theta_{30}^{2+\delta} T^{-\frac{\delta}{2}} \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \left| \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right|^{2+\delta} \delta_T^{-\frac{2+\delta}{2}} \left| K' \left( \frac{-\varepsilon_t}{\delta_T} \right) \right|^{2+\delta} \right] \\ \le \, &\theta_{30}^{2+\delta} (T \delta_T)^{-\frac{\delta}{2}} \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \left| \mathbf{h}_t^\top w_{\mathrm{Mode}} \lambda \right|^{2+\delta} \int \left| K' \left( u \right) \right|^{2+\delta}f_t(\delta_T u) \mathrm{d}u \right], \end{align} that converges to zero as $(T \delta_T)^{-\frac{\delta}{2}} \to 0$ and the respective moments are bounded by assumption.
lemmaIf $X_t$ is the mode of $F_t$ for all $t \in \mathbb{N}$, the choice of $u_{t,T}(\theta_0)$ in (ref) satisfies Assumption (ref) (D).
proofIf $X_t$ is the mode of $F_t$, we set $\theta_0 = (0,0,1)$ and thus, $\phi_{t,T}(\theta_0) = - \omega_\text{Mode} \delta_T^{-1/2} K'(\varepsilon_t / \delta_T) \mathbf{h}_t$ and we get $\phi_{t,T}(\theta_0) = \phi_{t,T}^\ast(\theta_0) + u_{t,T}(\theta_0)$ by setting \begin{align*} T^{-1/2} \phi_{t,T}^\ast(\theta_0) = \omega_Mode \, g_{t,T}^\ast \qquad and \qquad T^{-1/2} u_{t,T}(\theta_0) = \omega_Mode \, g_{t,T}^e, \end{align*} as in the proof of Theorem (ref). Thus, $T^{-1/2} \phi_{t,T}^\ast(\theta_0) = \omega_\text{Mode} \, g_{t,T}^\ast$ is a MDA satisfying condition (a). For the remaining conditions (b) and (c) of Assumption (ref) (D), as in the proof of Lemma (ref), we get \begin{align*} g_{t,T}^e = \frac{1}{2} T^{-1/2} \delta_T^{7/2} \mathbf{h}_t \int u^2 K \left( u \right) f”'_{t} (\zeta \delta_T u) \, \mathrm{d}u. \end{align*} As $\delta_T \to 0$ and $\int u^2 K \left( u \right) f'''_{t} (\zeta \delta_T u)$ is bounded as argued in the proof of Lemma (ref), we get \begin{align*} T^{-1} \sum_{t=1}^T || u_{t,T}(\theta_0) ||^2 = \sum_{t=1}^T || \omega_Mode \, g_{t,T}^e ||^2 = \delta_T^{7} \frac{\omega_Mode^2}{4} \, \frac{1}{T} \sum_{t=1}^T \left| \left| \mathbf{h}_t \int u^2 K \left( u \right) f”'_{t} (\zeta \delta_T u) \, \mathrm{d}u \right|\right|^2 \overset{P}{\to} 0. \end{align*} Furthermore, by using similar arguments as above, we get that \begin{align*} T^{-1} \sum_{t=1}^T u_{t,T}(\theta_0) \phi_{t,T}(\theta_0)^\top &= - \omega_Mode^2 \sum_{t=1}^T g_{t,T}^e \cdot (T \delta_T)^{-1/2} K'(\varepsilon_t / \delta_T) \mathbf{h}_t^\top \\ &= - \omega_\text{Mode}^2 \sum_{t=1}^T \left( \frac{1}{2} T^{-1/2} \delta_T^{7/2} \mathbf{h}_t \int u^2 K \left( u \right) f”'_{t} (\zeta \delta_T u) \, \mathrm{d}u \right) (T \delta_T)^{-1/2} K'(\varepsilon_t / \delta_T) \mathbf{h}_t^\top \\ &= \delta_T^3 \frac{\omega_\text{Mode}^2}{2} \frac{1}{T} \sum_{t=1}^T K'(\varepsilon_t / \delta_T) \int u^2 K \left( u \right) f”'_{t} (\zeta \delta_T u) \, \mathrm{d}u \; \mathbf{h}_t\mathbf{h}_t^\top \overset{P}{\to} 0. \end{align*} The condition $T^{-1} \sum_{t=1}^T \mathbb{E} \left[ u_{t,T}(\theta_0) \phi_{t,T} (\theta_0)^\top \right] \to 0$ follows by exactly the same arguments, but working under expectation. Eventually, we get that \begin{align*} \sum_{t=1}^T \mathbb{E} \left[ || T^{-1/2} u_{t,T}(\theta_0)||^{2+\delta} \right] = T^{-\delta/2} \delta_T^{(14 + 7 \delta)/2} \; \frac{\omega_\text{Mode}^{2+\delta}}{2^{2+\delta}} \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \left| \left| \mathbf{h}_t \int u^2 K \left( u \right) f”'_{t} (\zeta \delta_T u) \, \mathrm{d}u \right|\right|^{2+\delta} \right] \to 0, \end{align*} as $T^{-\delta/2} \delta_T^{(14 + 7 \delta)/2} \to 0$ and the remaining term is bounded, which concludes the proof.

{\color{black}Bandwidth Choice}

As discussed at the end of Section (ref), for the Gaussian kernel we choose the bandwidth as

align[align omitted — 499 chars of source]

which extends the rule-of-thumb proposed by Kemp2012 and Kemp2019 by additionally taking into account the empirical skewness of the data. While the convergence rate of $T^{-0.143}$, which is almost $T^{-1/7}$, follows from the theoretical considerations in Section (ref), we assess the performance of the data-driven constants $k_1$ and $k_2$ in the following through simulations.

In particular, we consider the simulation setup of Section (ref) and use the modified bandwidth choices given by

align[align omitted — 97 chars of source]

with the additional scaling factors $c \in \{0.5, 0.75, 1, 1.5, 2\}$.

table[table omitted — 3,160 chars of source]

Table (ref) reports the size of our mode rationality test for the adjustments with $c \in \{0.5, 0.75, 1, 1.5, 2\}$, for three skewness parameters of $\gamma \in \{0, 0.25, 0.5\}$ and the four considered sample sizes for the four considered DGPs in (ref). Besides the homoskedastic cross-sectional and the AR-GARCH DGPs described below (ref), we also use a heteroskedastic cross-sectional DGP that is as the homoskedastic one, but with $\sigma_{t+1} = 0.5 + 1.5(t + 1)/T$, and a simple AR(1) process with $Z_t = Y_{t}$, $\zeta = 0.5$ and $\sigma_{t+1} = 1$.

For symmetric data with $\gamma = 0$, all bandwidth choices result in a similar behavior under the null with approximately correct rejection rates. This easily follows from the fact that any modal midpoint (with any value of $\delta > 0$) coincides with the mode for symmetric data such that any bandwidth choice works theoretically. For the skewed data, however, an increasing bandwidth through a higher value of $c$ results in inflated test sizes, where our choice of $c=1$ with values below $8\%$ in reasonable sample sizes behaves satisfactory.

figure[figure omitted — 697 chars of source]

Figure (ref) continues to illustrate the test's power by plotting the rejection rates against the degree of misspecification $\kappa$ for the homoskedastic and AR-GARCH DGPs, the four skewness values and a fixed sample size of $T=500$. Notice that the rejection rates at $\kappa = 0$ display a subset of the values shown in Table (ref). We see a uniformly increasing power for increasing values of $c$ that can be explained by the fact that larger bandwidth choices effectively utilize more data. However, as already seen in Table (ref), larger values of $c$ inflate the rejection rates under the null hypothesis with $\kappa = 0$, hence leading to a severely oversized test.

Overall, as common in non-parametric statistics, the finite sample bandwidth choice is a delicate balancing of test size and power (or in the classical estimation world, bias and variance), where our choice of $c=1$ achieves a relatively balanced behavior of the test.

Kernel Choice

The asymptotic results presented in Section (ref) rely on the chosen kernel $K$ satisfying Assumption (ref). Besides the normalization $\int K(u) \mathrm{d}u = 1$ and boundedness assumptions, we impose the first-order kernel condition $\int u K(u) \mathrm{d}u = 0$ (and $\int u^2 K(u) \mathrm{d}u >0$ follows from the non-negativity of $K$). As discussed in LiRacine2006, higher-order kernels allow one to apply a Taylor expansion of higher order and can thereby obtain a faster rate of convergence, which could in theory be made arbitrarily close to $\sqrt{T}$, at the cost of stronger smoothness assumptions on the underlying density function. However, in our application of kernel functions to the generalized modal midpoint in Definition (ref), we need to ensure that the limit of this quantity is well-defined and unique, and that the identification is strict. For this, we assume in Theorem (ref) that the kernel function is log-concave which is automatically violated for higher-order kernels. Consequently, we do not consider higher-order kernels in this work.

It is also well-known in the literature on nonparametric statistics that kernels with bounded support can be more efficient LiRacine2006. However, note that the kernel choice enters the asymptotic variance of nonparametric density estimation through the quantity $\int K(u)^2 \mathrm{d}u$, while the covariance $\Omega_{\mathrm{Mode}}$ in Theorem (ref) depends upon $\int K'(u)^2 \mathrm{d}u$, revealing that the efficiency of our mode rationality tests depends upon a different quantity.

figure[figure omitted — 722 chars of source]

{\color{black}Figure (ref) compares the Gaussian kernel (that we use in our main analyses) with a biweight kernel within the simulation setup of Section (ref). For the biweight kernel, we adjust the bandwidth choice given in Section (ref) by multiplying with a factor of 2.5 in order to account for the different kernel shape.} Figure (ref) illustrates that the test power does not increase by employing a biweight kernel, which has bounded support, and is usually found to be relatively efficient in nonparametric estimation. Strict identifiability of the generalized modal midpoint---and hence, asymptotic identifiability of the mode---only holds for kernel functions with unbounded support, which is satisfied by the Gaussian kernel, but not so for the biweight kernel.

Convex combination of functional values

Here, we illustrate that a convex combination of functionals is generally neither elicitable nor identifiable. This result shows that---as stated in Remark (ref)---testing forecast rationality directly for convex functional combinations is impossible. For this, we adapt the simplified notation of Section (ref).

propositionLet $\mathcal{P}$ be a convex class of distributions and let $\Gamma_\beta(P)= \beta \Gamma^l(P) + (1-\beta) \Gamma^n(P)$, $\beta \in [0,1]$ be the convex combination of a linear functional $\Gamma^l: \mathcal{P} \to \mathbb{R}$ and a non-linear functional $\Gamma^n: \mathcal{P} \to \mathbb{R}$, which are both continuous (in the distribution $P$) and translation equivariant, i.e., if $P\in \mathcal{P}$, then for $c \in \mathbb{R}$ the shifted $P+c \in \mathcal{P}$ and $\Gamma^l(P+c)=\Gamma^l(P)+c$ and $\Gamma^n(P+c)=\Gamma^n(P)+c$. Then, the functional $\Gamma_\beta$ is neither elicitable nor identifiable.
proof[Proof of Proposition (ref)] Theorem 6 of Gneiting2011 and Proposition 3.11 of FisslerHoga2021 show that convex level sets, i.e., \begin{align} for P_1,P_2 \in \mathcal{P} with \Gamma_\beta(P_1)=\Gamma_\beta(P_2) \quad \Longrightarrow \quad \Gamma_\beta(\alpha P_1 + (1-\alpha) P_2) = \Gamma_\beta(P_1), \end{align} for $\alpha \in (0,1)$ are a necessary condition for elicitability and identifiability of a functional. We show that the functional $\Gamma_\beta$ does not have convex level sets. As $\Gamma^n$ is not linear, there exists $\alpha \in (0,1)$, and $P_1$, $P_2 \in \mathcal{P}$ such that $$\Gamma^n(\alpha P_1 + (1-\alpha) P_2) \neq \alpha \Gamma^n(P_1) + (1-\alpha) \Gamma^n(P_2).$$ Define $P_2' = P_2 - \Gamma_\beta(P_2) + \Gamma_\beta(P_1)$ such that $\Gamma_\beta(P_2')=\Gamma_\beta(P_1)$ and $$\Gamma^n(\alpha P_1 + (1-\alpha) P_2') \neq \alpha \Gamma^n(P_1) + (1-\alpha) \Gamma^n(P_2')$$ as $\Gamma^n(P+c)=\Gamma^n(P)+c$. It follows that \begin{align*} \Gamma_\beta(\alpha P_1 + (1-\alpha) P_2') &= \beta \Gamma^l(\alpha P_1 + (1-\alpha) P_2') + (1-\beta) \Gamma^n(\alpha P_1 + (1-\alpha) P_2') \\ &\neq \beta \alpha \Gamma^l(P_1) + \beta (1-\alpha) \Gamma^l(P_2') + (1-\beta) \alpha \Gamma^n(P_1) + (1-\beta) (1-\alpha) \Gamma^n(P_2')\\ &= \alpha \Gamma_\beta(P_1) + (1-\alpha) \Gamma_\beta(P_2')\&= \Gamma_\beta(P_1) \end{align*} and hence $\Gamma_\beta$ does not have convex level sets.

As functionals are linear if and only if they are expectations abernethy2012characterization, the mean is linear and the median is non-linear for classes $\mathcal{P}$ sufficiently rich enough such that it contains distributions for which the median does not equal the mean (i.e., asymmetric distributions). Hence, Proposition (ref) shows that a convex combination of the mean and median is generally neither elicitable nor identifiable. Eliciting a convex combination that may further include the mode (which is itself only asymptotically elicitable) is only possible in the unusual case where a convex (sub-)combination of the nonlinear median and mode functionals becomes linear, an outcome that does not generally hold.

Fortunately, testing rationality for functionals elicited through a convex combination of loss functions is feasible and its interpretability is supported by the following result that convexity of the combination weights is preserved when moving from a functional elicited by a convex combination of loss (identification) functions to a convex combination of functional values.

The loss functions $L_\mathrm{Mean}$, $L_\mathrm{Med}$ and $L_\mathrm{Mode, \delta} = \delta^{3/2} L_\delta^K$ are defined just before equation (ref) and the identification functions $V_\mathrm{Mean}$, $V_\mathrm{Med}$ and $V_\mathrm{Mode, \delta} = \delta^{3/2} V_\delta^K$ at equations (ref), (ref) and (ref). We use the scaling by $\delta^{3/2}$ for the asymptotic mode loss and identification functions to be consistent with Section (ref).

propositionLet $\mathcal{P}$ be some class of distributions such that the mean, $\mu$, the median, $m$, and the generalized modal midpoint with parameter $\delta$, $\Gamma^K_\delta$, exist and are elicited by their loss functions $L_\mathrm{Mean}$, $L_\mathrm{Med}$, and $L_\mathrm{Mode, \delta}$. Let $x$ be the functional defined by \begin{align} x(P) = \underset{\tilde x \in \mathbb{R}}{\operatorname*{arg\,min}} \; \mathbb{E}_{Y \sim P} \left[ \theta_0^\top \Big( L_\mathrm{Mean}(\tilde x,Y), L_\mathrm{Med}(\tilde x,Y), L_\mathrm{Mode, \delta}(\tilde x,Y) \Big)^\top \right] \end{align} for some $\theta_0 \in \Theta$. Then, for every $P \in \mathcal{P}$ there exists some $\beta_0 \in \Theta$, such that $x(P) = \beta_0^\top \big(\mu(P), m(P) , \Gamma^K_\delta(P) \big)^\top$.
proof[Proof of Proposition (ref)] Let $P \in \mathcal{P}$, where we assume without loss of generality that the three functionals are not all equal. For notational convenience we drop $P$ when denoting functional values, e.g., we write $\mu$ instead of $\mu(P)$. For the elicited forecast $x$, it holds that \begin{align} \overline{V}(x) := \theta_0^\top \Big( \overline{V}_\mathrm{Mean}(x,P), \overline{V}_\mathrm{Med}(x,P), \overline{V}_\mathrm{Mode, \delta}(x,P) \Big)^\top = 0. \end{align} We define $L := \min(\mu, m, \Gamma^K_\delta )$ and $U := \max(\mu, m, \Gamma^K_\delta )$ as the lower and upper functional values where it holds that $L<U$. Further let $\bar V_L( x,P)$ and $\bar V_U( x,P)$ denote the corresponding expected identification functions for the distribution $P$. Suppose that $x < L$. Then, it must hold that $\overline{V}(x) > 0$ as all three expected identification functions have the same sign as they are oriented in the sense of steinwart. Similarly, if $x > U$, it must hold that $\overline{V}(x) < 0$. Hence, we can conclude that $x \in [L,U]$, which implies that there exists $\zeta \in [0,1]$ such that $ x = \zeta L + (1-\zeta) U$. Thus, $ x$ can be constructed as a convex combination of the functional values, i.e., there exists a $\beta_0 \in \Theta$ such that $x = \beta_0^\top \big(\mu, m , \Gamma^K_\delta \big)^\top$.

{\color{black}Heterogeneity of Forecasters}

In this section, we discuss heterogeneity in the measure of centrality adopted by survey respondents. For this discussion we adopt the simplified decision-theoretic notation of Section (ref). In particular, consider a forecaster who issues a (centrality) forecast $x$ for her outcome variable $Y \sim P$, whose distribution $P \in \mathcal{P}$ is in some suitable class of distributions $\mathcal{P}$.

Formally, we represent the mixture of forecaster types through a probabilistic mixture of mean, median, and the generalized modal midpoint (with bandwidth $\delta$) forecasts with combination weights (probabilities) $\gamma = (\gamma_\mathrm{Mean},\gamma_\mathrm{Med},\gamma_\mathrm{Mode, \delta}) \in \Theta$. Given a random variable $Z \sim C_\gamma$ that is categorically distributed with values $\{(1,0,0), \, (0,1,0), \, (0,0,1)\}$ and corresponding probabilities $\gamma$, we define the probabilistic mixture (with the symbol $ \otimes_\gamma$) as

equation[equation omitted — 194 chars of source]

Notice that opposed to the deterministic forecasts $x$ considered in Section (ref), the probabilistic mixture $x_\gamma^\dagger$ is random through its definition relying on the random variable $Z$.

propositionLet $\mathcal{P}$ be some class of distributions such that the mean, $\mu = \mu(P)$, the median, $m = m(P)$, and the generalized modal midpoint with parameter $\delta$, $\Gamma^K_\delta = \Gamma^K_\delta(P)$, exist and are unique for all $P \in \mathcal{P}$. Let $x_\gamma^\dagger$ be the functional defined by Equation ((ref)) for some $\gamma \in \Theta$. Then, for every $P \in \mathcal{P}$ there exists some $\theta_0 \in \Theta$, such that $$ {\mathbb{E}_{Y\sim P, Z \sim C_\gamma}} \left[ \theta_0^\top \Big({V}_\mathrm{Mean}(x_\gamma^\dagger,Y),{V}_\mathrm{Med}(x_\gamma^\dagger,Y),{V}_{\mathrm{Mode, \delta}}(x_\gamma^\dagger,Y)\Big)^\top \right] = 0 $$
proofUsing the notation $V_\theta(x,Y)=\theta^\top \Big({V}_\mathrm{Mean}(x,Y),{V}_\mathrm{Med}(x,Y),{V}_{\mathrm{Mode, \delta}}(x,Y)\Big)^\top$ and \\ $\overline{V}(x,P)= \mathbb{E}_{Y\sim P, Z \sim C_\gamma}[V(x,Y)]$, the linearity of the expectation operator implies \begin{align} \overline{V}_\theta(x,P) = \theta^\top \Big( \overline{V}_\mathrm{Mean}(x,P), \overline{V}_\mathrm{Med}(x,P), \overline{V}_{\mathrm{Mode, \delta}}(x,P) \Big)^\top. \end{align} Using $x= \otimes_\gamma (\mu(P),m(P),\Gamma^K_\delta(P))$ and the law of total expectations, we get for the first component that \begin{align} \overline{V}_\mathrm{Mean}(x,P)=\gamma^\top\Big( \overline{V}_{\mathrm{Mean}}(\mu,P), \overline{V}_{\mathrm{Mean}}(m,P), \overline{V}_{\mathrm{Mean}}(\Gamma^K_\delta,P)\Big)^\top \end{align} where it naturally holds that $\overline{V}_{\mathrm{Mean}}(\mu,P)=0$. Analogous formulas to (ref) hold for $\overline{V}_\mathrm{Med}(x,P)$ and $\overline{V}_{\mathrm{Mode, \delta}}(x,P)$. Plugging these into (ref) gives \begin{align} \overline{V}_\theta(x,P)=\theta_\mathrm{Mean} \Big(\gamma_\mathrm{Med} \overline{V}_{\mathrm{Mean}}(m,P)+ \gamma_{\mathrm{Mode, \delta}} \overline{V}_{\mathrm{Mean}}(\Gamma^K_\delta,P)\Big)+ \nonumber \\ \theta_\mathrm{Med} \Big(\gamma_\mathrm{Mean} \overline{V}_{\mathrm{Med}}(\mu,P)+ \gamma_{\mathrm{Mode, \delta}} \overline{V}_{\mathrm{Med}}(\Gamma^K_\delta,P)\Big)+\\ \theta_{\mathrm{Mode, \delta}} \Big(\gamma_\mathrm{Mean} \overline{V}_{\mathrm{Mode, \delta}}(\mu,P)+ \gamma_\mathrm{Med} \overline{V}_{\mathrm{Mode, \delta}}(m,P)\Big). \nonumber \end{align} Let $L = L(P) = \min(\mu(P),m(P),\Gamma^K_\delta(P))$ denote the “lowest” functional, $U$ the “highest” functional, and $M$ the third functional in the “middle”. As the three identification functions are oriented in the sense of steinwart, it holds that $\overline{V}_{L}(U,P)\ge 0 $ and $\overline{V}_{L}(M,P) \ge 0$ such that \[ \gamma_U \overline{V}_{L}(U,P)+\gamma_M \overline{V}_{L}(M,P) \ge 0. \] Analogously, it follows that \[ \gamma_L \overline{V}_{U}(L,P)+\gamma_M \overline{V}_{U}(M,P) \le 0. \] Thus, there is at least one non-negative and one non-positive term in Equation ((ref)) and for any $P$ and any $\gamma$ there exits a $\theta$ such that $\overline{V}_\theta(x,P)=0$. (If two functionals are the same, a similar argument applies.)

Proposition (ref) illustrates that in a setting of multiple forecasters, which have identical forecast distributions $P$, but are heterogeneous in either forecasting the mean, median or mode (modal midpoint), the asymptotic moment condition that there exists a $\theta_0 \in \Theta$ such that $\frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \phi_{t,T} (\theta_0) \right] \to 0$, which is closely related to Assumption (ref) (D), holds for the choice $\mathbf{h}_t = 1$. Moreover, the value of $\theta_0 \in \Theta$ does not change for any location shift $P+c$ of the distribution $P$, as all three identification functions fulfill the translation invariance property $V(x+c,y+c)=V(x,y)$. Thus, Proposition (ref) implies the existence of a unique $\theta_0 \in \Theta$ for which the moment condition is satisfied for $\mathbf{h}_t = 1$, even if different forecasters (that report heterogeneous functionals) have shifted distributions. For $\mathbf{h}_t = 1$, Proposition (ref) can also be extended to the more general assumption that the heterogeneous forecast distributions only have identical order of mean, median, and modal midpoint.\footnote{ In that case, we define $\mathcal{P}^O \subseteq \mathcal{P}$ as the subclass of distributions that have identical order of mean, median, and modal midpoint and consider a probability distribution $Q$ on $\mathcal{P}^O$ (that informally represents the different distributions of the different forecasters). The proof of Proposition (ref) then works equivalently by considering the expected identification functions $\overline{V}(x,Q)$ to also capture the expectation over $Q$, which represents (the distribution over) the different subjective distributions.} However, multiple instruments (as e.g., our baseline case of $\mathbf{h}_t = (1,X_t)$) generally imply different values of $\theta_0$ for each component of the instrument vector, such that there exists no $\theta_0$ that ensures the (vector-valued) validity of Assumption (ref) (D).

Also note that in the case of a probabilistic mixture, the parameter $\theta_0$ contains information on the ratio of forecasters $\gamma$, but their values are not necessarily equal. Using Equation ((ref)) and (for simplicity) considering the situation without mode (modal midpoint) forecasters, we observe that $$\frac{\theta_{\mathrm{Mean}}}{\theta_\mathrm{Med}}=\frac{\gamma_\mathrm{Mean}}{\gamma_\mathrm{Med}} \frac{\overline{V}_{\mathrm{Med}}(\mu,P)}{\overline{V}_{\mathrm{Mean}}(m,P)},$$ such that $\theta_0$ reflects the proportion of forecasters if the identification functions are suitably standardized.

figure[figure omitted — 631 chars of source]

We next illustrate via simulations that our confidence sets from (ref) develop power against probabilistic combinations of functional values as in (ref) for sufficiently informative instruments. For this, we utilize the iid DGP from (ref) and simulate according to probabilistic combinations in (ref) using mean-median combinations with $\gamma = (0.5, 0.5,0)$, mean-mode combinations with $\gamma = (0.5, 0, 0.5)$, and median-mode combinations with $\gamma = (0, 0.5, 0.5)$. We use the two instrument choices $\mathbf{h}_t = 1$ and $\mathbf{h}_t= (1,X_t)$.

Figure (ref) depicts the confidence set coverage rates for these six settings. As illustrated by Proposition (ref), when using a constant instrument $\mathbf{h}_t = 1$, we can always find a $\theta_0 \in \Theta$ (even partially identified lines in $\Theta$) with approximately nominal coverage rates of $90\%$. In contrast, when using the more informative instruments $\mathbf{h}_t= (1,X_t)$, we see rejection rates clearly below the nominal level of $90\%$ showing that our confidence sets develop power against the entire unit simplex $\Theta$.

In summary, this section illustrates that our confidence sets for convex combinations of centrality measures contain probabilistic mixtures of forecaster types only for unconditional tests using the constant as the only instrument. In the baseline case of our empirical application, $\mathbf{h}_t= (1,X_t)$, our procedure develops power against probabilistic mixtures.

{\color{black}Additional Simulation Results}

table[table omitted — 2,668 chars of source]
table[table omitted — 2,832 chars of source]
figure[figure omitted — 608 chars of source]
figure[figure omitted — 582 chars of source]
figure[figure omitted — 630 chars of source]
figure[figure omitted — 207 chars of source]
figure[figure omitted — 203 chars of source]
figure[figure omitted — 194 chars of source]

Additional results for the Survey of Consumer Expectations

Clustered Covariance Estimator

Figures (ref) to (ref) below are equivalent to Figures (ref) to (ref) after substituting the covariance estimator of Theorem (ref) with a clustered covariance estimator $\widehat{\Sigma}^{CL}_T$. Let $\phi_{i,t,T}(\theta)$ denote the moment function of individual $i$ at time $t$. $\mathcal{T}$ denotes the number of waves and $n_t$ the number of observations within wave such that $T=\sum_{t=1}^\mathcal{T} n_t$.

align[align omitted — 213 chars of source]

Overall, the results are robust to clustering at the time level. While the mean rejection is less pronounced in Figure (ref), the confidence sets are sharper for the subpopulations. In Figure (ref) to (ref) mean rationality is consistently rejected for lower income individuals at the 5% level.

figure[figure omitted — 333 chars of source]
figure[figure omitted — 354 chars of source]
figure[figure omitted — 453 chars of source]
figure[figure omitted — 463 chars of source]
figure[figure omitted — 472 chars of source]

\FloatBarrier

Forecast rationality and past forecast accuracy

In this section we investigate the relationship between past forecast accuracy and the rationality of subsequent forecasts. dacunto2019cognitive find absolute forecast errors to be negatively correlated with cognitive ability measures for inflation expectations, which may suggest that respondents with larger past forecast errors are more likely to issue irrational forecasts. In contrast, vanNVeldkamp2006learning propose that economic agents learn more about the economy when more signals are available, and a large forecast error may signal to the respondent that they need to pay more attention to their forecast. We address this question using the set of 1,288 respondents to the FRBNY survey for which we have two matched pairs of predictions and realizations. We classify respondents as having “small” or “large” past forecast errors using the cross-respondent median of absolute percentage forecast errors in the first round as a cutoff. In addition to the (relative) size of their past forecast errors, we continue to stratify by income.

figure[figure omitted — 325 chars of source]

Figure (ref) shows that respondents who had a small forecast error in the first round provide second-round forecasts that can be rationalized using many different measures of centrality, and there are few differences across low- and high-income respondents: for low-income respondents, only the mean leads to a rejection, while for high-income respondents the median and measures between the median and the mode are rejected. For respondents with a large forecast error in the first round, the rationality test results are starkly different for low and high income respondents: high-income respondents can be rationalized using any measure of centrality, while low-income respondents can be rationalized only as mean, or near-mean, forecasters, but only at the 95% confidence level; at the 90% confidence level rationality is rejected for all centrality measures. An alternative interpretation of Figure (ref) is that high-income respondents produce forecasts that are rationalizable for almost all measures of centrality, regardless of past forecast accuracy, while the rationality of low-income respondents' forecasts varies greatly with past forecast errors: those with large past forecast errors are not rationalizable as any measure of centrality at the 90% confidence level, while those with small forecast errors are rationalizable as almost any measure of centrality.

Figure (ref) reports results corresponding to Figure (ref), but based on the clustered standard errors described in Section (ref).

figure[figure omitted — 465 chars of source]

Further results for the EKT tests on subsamples

Table (ref) shows the test results for the subsamples considered in Section (ref) that were omitted in Table (ref).

table[table omitted — 2,031 chars of source]

Additional Empirical Applications

In this section, we present two additional economic applications, to “Greenbook” forecasts of US GDP growth, and to random walk forecasts of exchange rates. In some of this applications we find evidence against mode rationality but not against mean rationality. This confirms that our proposed mode forecast rationality test has non-trivial power in relevant applications.

Greenbook forecasts of U.S.\ GDP growth

First, we consider one-quarter-ahead forecasts of U.S.\ GDP growth produced by the staff of the Board of Governors of the Federal Reserve (the so-called “Greenbook” forecasts), from 1967Q2 until 2015Q2, a total of $192$ observations.\footnote{Greenbook forecasts are only available to the public after a five-year lag.} These forecasts are prepared in preparation for each meeting of the Federal Open Market Committee, and substantial resources are devoted to them, see e.g.\ RomerRomer2000. Greenbook forecasts are available several times each quarter; for this analysis we take the single forecast closest to the middle date in each quarter. Broadly similar results are found when using the first, or last, forecast within each quarter.

figure[figure omitted — 713 chars of source]

(ref) presents the confidence set for the measures of centrality that can be rationalized for these forecasts. As GDP growth is measured with error and official values are often revised, we present results for three different “vintages” of the realized value: the first, second and most recent release. For the first and second vintages, we see that only measures of centrality “close to” the mean can be rationalized as optimal, while the mode, median and similar measures can all be rejected. This is particularly noteworthy given the known lower power at the mode vertex. Using the most recent vintage for GDP growth, both the mean and median, and centrality measures between and near those, are included in the confidence set. That the Greenbook GDP forecasts are rational when interpreted as mean forecasts, but not when taken as mode or median forecasts, is consistent with the Fed staff using econometric models for these forecasts, as such models almost invariably focus on the mean.\footnote{ReifschneiderTulip2017 discuss the ambiguity in the specific centrality measure reported in the Greenbook forecasts, but write that they are “typically viewed as modal forecasts” by the Federal Reserve staff. Our results suggest that they are better interpreted as mean forecasts.}

Random walk forecasts of exchange rates

For our final empirical application we revisit the famous result of MeeseRogoff1983, that exchange rate movements are approximately unpredictable when evaluated by the squared-error loss function, implying that the lagged exchange rate is an optimal mean forecast. See Rossi2013 for a more recent survey of the literature on forecasting exchange rates. We use daily data from the European Central Bank's “Statistical Data Warehouse” on the USD/EUR, JPY/EUR and AUD/EUR exchange rates, over the period January 2000 to July 2020, a total of $5,265$ trading days. Note that our sample period has no overlap with that of MeeseRogoff1983, and so their conclusions about the mean-optimality of the random walk forecast need not hold in our data.

figure[figure omitted — 660 chars of source]

(ref) presents the results of our tests for rationality, all of which use a constant and the forecast as the instrument set. The middle and right panels reveal that for the JPY/EUR and AUD/EUR exchange rates the lagged exchange rate is not rejected as a mean forecast, while it is rejected when taken as a mode or median forecast. Thus the rationality of the random walk forecast critically depends, for these exchange rates, on whether it is interpreted as a mean, median or mode forecast. For the USD/EUR exchange rate we cannot reject rationality with respect to any of convex combination of these measures of central tendency, implying that the random walk forecast is consistent with rationality under any of these measures.\footnote{Results for the GBP/EUR and CAD/EUR exchange rates are identical to those for the USD/EUR.} The mean vertex being included in the confidence set for all three exchange rates, indicating no evidence against rationality of the random walk model when interpreted as a mean forecast, is consistent with the conclusion of MeeseRogoff1983.

Summary of results across empirical applications

Table (ref) summarizes the rationality tests for the mean, median, and mode functionals, as well as the EKT2005 rationality test that allows for optimism and pessimism. For the Federal Reserve Bank of New York (FRBNY) survey data, mode rationality is not rejected, but all other functionals are rejected. For the Greenbook GDP forecasts, the mode is rejected a the 5% level, but the mean and expectiles close to the mean are consistent with the data. For the exchange rates random walk forecasts, the mean and central quantiles and expectiles are consistent with the data, while mode rationality is rejected for two of the three exchange rates at the 5% level.

table[table omitted — 1,646 chars of source]