The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
25,468 characters
Characterizing M-estimators
\title{Characterizing M-estimators}
\author{Timo Dimitriadis\thanks{Heidelberg University, Alfred Weber Institute of Economics, Bergheimer Str.\ 58, 69115 Heidelberg, Germany and
Heidelberg Institute for Theoretical Studies, 69118 Heidelberg, Germany, e-mail: [email removed]}
\and Tobias Fissler\thanks{Vienna University of Economics and Business (WU), Department of Finance, Accounting and Statistics, Welthandelsplatz 1, 1020 Vienna, Austria,
e-mail: [email removed]} \and Johanna Ziegel\thanks{University of Bern, Department of Mathematics and Statistics, Institute of Mathematical Statistics and Actuarial Science, Alpeneggstrasse 22, 3012 Bern, Switzerland,
e-mail: [email removed]}
}
\maketitle
\begin{abstract}
We characterize the full classes of M-estimators for semiparametric models of general functionals by formally connecting the theory of consistent loss functions from forecast evaluation with the theory of M-estimation.
This novel characterization result opens up the possibility for theoretical research on efficient and equivariant M-estimation and, more generally, it allows to leverage existing results on loss functions known from the literature of forecast evaluation in estimation theory.
\end{abstract}
\textit{Keywords:}
M-estimation, loss function, strict consistency, characterization
\section{Introduction}
\label{sec:intro}
\onehalfspacing
The task of regression is to model the effect of covariates $X$ on a response variable $Y$, or more precisely, the effect of $X$ on a functional $\Gamma$ of the conditional distribution of $Y$ given $X$, $F_{Y|X}$.
The typical example for $\Gamma$ is the mean, resulting in mean regression.
In many application, one is interested in other functionals such as quantiles (Value at Risk, VaR), expectiles, variance, or Expected Shortfall (ES) \citep{Koenker1978, Bollerslev1986, Patton2019}.
A correctly specified parametric model $m(X,\theta)$ satisfies
$\Gamma(F_{Y|X}) = m(X,\theta_0)$ for some unique parameter $\theta_0\in\Theta$.
The statistician's task is to estimate the parameter $\theta_0$ based on data $(Y_t,X_t)$, $t=1, \ldots, N$.
For the standard situation of linear mean regression, $\Gamma(F_{Y|X}) = \mathbb{E}[Y|X]$, and $m(X, \theta) = X^\intercal \theta$,
one often employs the ordinary-least-squares (OLS) estimator of the form
$\widehat \theta_{T} = \operatorname*{arg\,min}_{\theta \in\Theta} \frac{1}{T} \sum_{t=1}^T \big(Y_t - X_t^\intercal \theta\big)^2$, or a related closed-form solution thereof.
The OLS estimator is a special instance of an M-estimator \citep{Huber1967, NeweyMcFadden1994},
\begin{equation} \label{eq:M-est}
\widehat \theta_{T} = \operatorname*{arg\,min}_{\theta \in\Theta} \frac{1}{T} \sum_{t=1}^T \rho \big(Y_t , m(X_t,\theta)\big),
\end{equation}
based on a loss function $\rho$, which is the key ingredient of an M-estimator.
Note that our results carry over to time-varying loss functions $\rho_t$.
The core condition on $\rho$ for consistency of $\widehat \theta_{T}$ is that
$\mathbb{E}\big[\rho \big(Y_t , m(X_t,\theta_0)\big)\big]<\mathbb{E}\big[\rho \big(Y_t , m(X_t,\theta)\big)\big]$ for all $\theta\neq\theta_0$ and for all $t\in\mathbb N$, which we call \emph{strict model-consistency} of $\rho$ for $m$.
Apart from special cases, the classes of such loss functions $\rho$ are unfortunately not well understood in M-estimation yet.
Our main result, Theorem \ref{thm:hierarchy consistency}, establishes that a loss function $\rho$ is (strictly) model-consistent for $m$ if and only if it is (strictly) \emph{consistent} for the target functional $\Gamma$, meaning that
$
\int \rho \big(y,\Gamma(F)\big)\mathrm{d} F(y) \le \int \rho(y,\xi)\mathrm{d} F(y)
$
for all $\xi \in \mathbb{R}^k$ and for all distributions $F$ in a sufficiently rich class, where strictness means that equality implies $\xi=\Gamma(F)$.
Since there are well-understood characterization results for strictly consistent losses from the literature of forecast evaluation \citep{Gneiting2011, FisslerZiegel2016}, Theorem \ref{thm:hierarchy consistency} lifts these results to a novel characterization of the classes of consistent M-estimators for general (vector-valued) functionals.
Our result provides the first characterization of the full classes of consistent M-estimators for semiparametric models of general, possibly vector-valued functionals by formally connecting the two strands of literature on M-estimation and forecast evaluation.
Understanding the full classes of M-estimators facilitates a deeper understanding of the (im)possibilities, e.g., in the following areas.
It allows to derive M-estimation efficiency bounds, similar to efficient generalized method of moments estimation of \citet{Chamberlain1987}.
Drawing on the literature of homogeneous and equivariant loss functions \citep{NoldeZiegel2017, FisslerZiegel2019}, our connection allows to classify such M-estimators having e.g.~the advantage that they are invariant to linear rescaling in finite samples.
Furthermore, estimators with the most beneficial integrability conditions in the sense that $\mathbb{E}[ \big| \rho \big(Y_t , m(X_t,\theta)\big) \big|]$ is finite can be derived.
Our result further allows to leverage recent discoveries in the forecast evaluation literature such as mixture representations \citep{EhmETAL2016} or score decompositions \citep{DGJ_2021} for the purpose of M-estimation.
Theorem \ref{thm:hierarchy consistency} also provides the first formal argument why consistent M-estimation is basically impossible if there do not exist (strictly) consistent loss functions for the target functional $\Gamma$. This argument has already been used informally for the Expected Shortfall \citep{Patton2019, DimiBayer2019}, the Range Value at Risk \citep{Barendse2020}, and the mode \citep{Kemp2012}.
\section{Notation and Definitions}
\label{subsec:Notation and Setting}
Let $(\Omega, \mathcal A, \mathbb P)$ be a non-atomic, complete probability space where all random variables are defined.
We introduce a class $\mathcal{Y}$ of $\mathbb{R}^d$-valued possible response variables, a corresponding class $\mathcal{X}$ or $\mathbb{R}^p$-valued regressors, and $\mathcal{Z}$ the class of possible response--regressor pairs $(Y,X)$.
The class of marginal distributions $F_Y$ of $Y\in\mathcal{Y}$ is denoted by $\mathcal{F}_\mathcal{Y}$, with a corresponding notation $\mathcal{F}_\mathcal{X}$.
$\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ is the class of regular versions of conditional distributions $F_{Y|X}$ for any $(Y,X)\in \mathcal{Z}$; see Appendix \ref{sec:App1} for technical details.
We will identify cumulative distribution functions with their corresponding measures where convenient.
Let $\Gamma \colon\mathcal{F}_{\mathcal{Y}|\mathcal{X}}\to\Xi\subseteq\mathbb{R}^k$ be some $k$-dimensional, single-valued functional of the conditional distribution of $Y$ given $X$.
Let $\Theta\subseteq \mathbb{R}^q$ be some parameter space with non-empty interior, $\operatorname{int}(\Theta)$, and $m\colon \mathbb{R}^p\times \Theta\to \Xi$ a parametric model for the functional $\Gamma$.
We shall work under the following assumption of a correctly specified model with a uniquely identified true model parameter.
\begin{assumption}
\label{ass:unique model}
For all $Z=(Y, X)\in\mathcal{Z}$ there is a unique parameter $\theta_0 = \theta_0(F_Z) \in \operatorname{int}(\Theta)$ such that almost surely
\begin{equation} \label{eq:unique model}
m(X,\theta_0) = \Gamma(F_{Y|X}).
\end{equation}
\end{assumption}
Assumption \ref{ass:unique model} is semiparametric in the sense that the finite-dimensional parameter $\theta_0$ in \eqref{eq:unique model} does not fully describe the distribution of $(Y,X)$, but only $\Gamma(F_{Y|X})$, the component of the conditional distribution we are essentially interested in.
In general, an estimator should be valid on a class of random variables $\mathcal{Z}$ which is as large as possible, allowing to apply the estimator in many different situations and under uncertainty of the distributions of the underlying data.
Hence, Assumption \ref{ass:unique model} is a minimal condition on the random variables that allows for semiparametric modeling of a functional of the conditional distribution.
We continue by recalling the use of loss functions in the closely related area of forecast evaluation, where the notion of a \emph{strictly consistent} loss function is a crucial concept in the literature on forecast evaluation, since it incentivizes truthful reports \citep{MurphyDaan1985}.
Making use of a similar decision-theoretic terminology as in \cite{Gneiting2011} and \cite{FisslerZiegel2016},
let $\mathcal{F}$ be some generic class of probability distributions on $\mathbb{R}^d$, which is our {observation domain}, and $\Xi\subseteq \mathbb{R}^k$, our {action domain}.
\begin{definition}[Consistency and elicitability]
\label{defn:consistency}
A loss function $\rho\colon \mathbb{R}^d\times \Xi\to\mathbb{R}$ is called \emph{$\mathcal{F}$-consistent} for a functional $\Gamma\colon \mathcal{F}\to\Xi$ if $\rho(\cdot, \xi)$ is $F$-integrable for all $F\in\mathcal{F}$ and for all $\xi\in\Xi$, and if
\begin{equation} \label{eq:strict consistency}
\int \rho\big(y,\Gamma(F)\big) \,\mathrm{d} F(y) \le \int \rho (y, \xi)\,\mathrm{d} F(y) \qquad \text{for all } F\in\mathcal{F}, \ \text{for all } \xi\in\Xi\,.
\end{equation}
If equality in \eqref{eq:strict consistency} implies $\xi = \Gamma(F)$, then the loss function is called \emph{strictly $\mathcal{F}$-consistent} for $\Gamma$.
A functional $\Gamma\colon \mathcal{F}\to\Xi$ is \emph{elicitable} if there is a strictly $\mathcal{F}$-consistent loss function for it.
\end{definition}
On the class of distributions with a finite second moment,
the squared loss $\rho(y,\xi) = (y-\xi)^2$ is strictly consistent for the mean functional.
More generally, subject to regularity and integrability conditions on $\rho$ and richness conditions on $\mathcal{F}$, $\rho$ is (strictly) $\mathcal{F}$-consistent for the mean if and only if it is a so-called \emph{Bregman loss}
\begin{align}
\label{eqn:BregmanLoss}
\rho(y,\xi) = \phi(y) - \phi(\xi) +\phi'(\xi)(\xi-y) + \kappa(y),
\end{align}
where $\phi$ is a (strictly) convex function on $\mathbb{R}$ with subgradient $\phi'$ and $\kappa$ is any function of $y$ \citep{Savage1971, Gneiting2011}.
Likewise, the well known pinball or asymmetric absolute loss $\rho(y,\xi) = (\mathds{1}\{y\le \xi\} - \alpha)(\xi - y)$ is strictly consistent for the lower $\alpha$-quantile on the class of distributions with finite mean and where the lower $\alpha$-quantile coincides with the upper $\alpha$-quantile.
Moreover, subject to regularity and integrability conditions on $\rho$ and richness conditions on $\mathcal{F}$, a loss $\rho$ is (strictly) consistent for the lower $\alpha$-quantile, $\alpha\in(0,1)$, if and only if it is a \emph{generalized piecewise linear loss function}
\begin{align}
\label{eqn:GPLLoss}
\rho(y,\xi) = (\mathds{1}\{y\le \xi\} - \alpha)(g(\xi) - g(y)) + \kappa(y),
\end{align}
where $g$ is (strictly) increasing \citep{Gneiting2011b}.
Similar characterization results exist for other functionals such as expectiles or the pairs consisting of the mean and variance or the quantile and Expected Shortfall \citep{Gneiting2011, FisslerZiegel2016}.
The characterization results in \eqref{eqn:BregmanLoss} and \eqref{eqn:GPLLoss} rely on the fact that the related classes $\mathcal{F}$ are convex and rich enough.
E.g., if we restrict attention to symmetric distributions only, the mean equals the median, and loss functions of the form given in \eqref{eqn:BregmanLoss} or \eqref{eqn:GPLLoss} would elicit the mean and median.
The following definition develops similar notions of consistency for the setting of M-estimation with an underlying class $\mathcal{Z}$ implying that $\Gamma$ is defined on the class of conditional distributions $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$.
\begin{definition}[Model-consistency]
\label{defn:model-consistency}
Suppose Assumption \ref{ass:unique model} holds for the parametric model $m\colon\mathbb{R}^p\times\Theta\to\Xi$ and the functional $\Gamma\colon\mathcal{F}_{\mathcal{Y}|\mathcal{X}}\to\Xi$.
Let $\rho\colon \mathbb{R}\times \Xi\to\mathbb{R}$ be a loss function such that $\mathbb{E}\big|\rho\big(Y,m(X,\theta)\big)\big|<\infty$ for all $(Y,X) \in\mathcal{Z}$ and for all $\theta\in\Theta$.
\begin{enumerate}[label=(\roman*)]
\item
The loss $\rho$ is \emph{unconditionally $\mathcal{F}_\mathcal{Z}$-model-consistent} for the model $m$ if
\begin{equation} \label{eq:unconditional model consistency}
\mathbb{E}\big[\rho\big(Y,m(X,\theta_0)\big)\big]\le \mathbb{E}\big[\rho\big(Y,m(X,\theta)\big)\big]
\qquad \text{for all } (Y,X)\in\mathcal{Z}, \ \text{for all } \theta\in\Theta\,.
\end{equation}
Moreover, $\rho$ is \emph{strictly unconditionally $\mathcal{F}_\mathcal{Z}$-model-consistent} for the model $m$ if equality in \eqref{eq:unconditional model consistency} implies that $\theta = \theta_0$.
\item
The loss $\rho$ is \emph{conditionally $\mathcal{F}_\mathcal{Z}$-model-consistent} for the model $m$ if
\begin{equation} \label{eq:conditional model consistency}
\mathbb{E}\big[\rho\big(Y,m(X,\theta_0)\big)\big|X\big]\le \mathbb{E}\big[\rho\big(Y,m(X,\theta)\big)\big|X\big]\ \text{a.s.}
\quad \text{for all } (Y,X)\in\mathcal{Z}, \ \text{for all } \theta\in\Theta\,.
\end{equation}
Moreover, $\rho$ is \emph{strictly conditionally $\mathcal{F}_\mathcal{Z}$-model-consistent} for the model $m$ if almost sure equality in \eqref{eq:conditional model consistency} implies that $\theta=\theta_0$.
\end{enumerate}
\end{definition}
The concept of unconditional model-consistency is the central condition for consistent M-estimation: \citet[Property 3.3 and 3.4]{Gourieroux1987} show equivalence of these two notions in terms of first order conditions and under some regularity assumptions.
Also see condition (i) in \citet[Theorem 2.1]{NeweyMcFadden1994}, \citet[Assumption (A-4)]{Huber1967}, where their remaining assumptions are merely regularity conditions.
In contrast, the conditional notion is appealing since it bridges the gap between consistency for $\Gamma$ according to Definition \ref{defn:consistency} and unconditional model-consistency for $m$ in the proof of Theorem \ref{thm:hierarchy consistency} below. It can still be practically useful when we resort to nonparametric kernel regressions or in the presence of repeated observations of $X$, e.g.\ if $X$ consists of categorical variables only.
While the classical notion of consistency in \eqref{eq:strict consistency} and the unconditional model version in \eqref{eq:unconditional model consistency} are closely related, they are crucially different in that in the latter, the expectation is also taken with respect to the covariates.
Establishing and finding reasonable conditions for their equivalence is indeed not trivial as shown in Theorem \ref{thm:hierarchy consistency} below.
\section{Main Result and Discussion}
\label{sec:main result}
To present our main Theorem \ref{thm:hierarchy consistency}, we introduce and discuss the following two assumptions.
\begin{assumption}\label{ass:separability}
For all $X\in\mathcal{X}$, the map $m(X,\cdot)\colon\Theta\to\Xi$ is surjective almost surely.
For all $(Y,X)\in\mathcal{Z}$ the conditional expectation $\mathbb{E}\big[\rho\big(Y,m(X,\theta)\big)\big|X\big]$ is continuous in $\theta$ almost surely.
\end{assumption}
The surjectivity in Assumption \ref{ass:separability} can usually be fulfilled by a sensible choice of $\Xi$, and the smoothness condition on the expected loss is standard in the literature, see e.g., \citet[Section 2.3]{NeweyMcFadden1994}.
\begin{assumption}\label{ass:reweighting}
For any $Z = (Y,X)\in\mathcal{Z}$ and any event $A\in \sigma(X)$ with positive probability $\mathbb P(A)>0$, there is some $\widetilde Z\in\mathcal{Z}$ such that $\mathbb P\big( \widetilde Z \in B\big) = \mathbb P\big(Z\in B\,|A\big)$ for all Borel sets $B\subseteq \mathbb{R}^{d+p}$.
\end{assumption}
Assumption \ref{ass:reweighting} is a richness condition on the class of possible data generating processes (DGP) in $\mathcal{Z}$ as for any process $Z = (Y,X) \in \mathcal{Z}$, and any set $A$ with positive probability, it stipulates that $\mathcal{Z}$ is rich enough to contain a process $\widetilde Z = (\widetilde Y, \widetilde X)$ as specified in Assumption \ref{ass:reweighting}.
Crucially, it yields that $\mathbb P(\widetilde Y\in C\, | \widetilde X) = \mathbb P(Y\in C\, | X,A)$ for all Borel sets $C\subseteq \mathbb{R}$. This, together with Assumption \ref{ass:unique model}, implies that the correctly specified parameter and hence the semiparametric model, is the same under the distributions $F_{\widetilde Z}$ and $F_Z$; in formulae, $\theta_0(F_{\widetilde Z}) = \theta_0(F_Z)$.
Recall that in estimation, $\mathcal{Z}$ captures the flexibility about the underlying and in practice unknown DGP, such that a large $\mathcal{Z}$ is desirable in order to obtain an estimation method which is applicable to a wide range of distributions of $Z \in \mathcal{Z}$.
Thus, Assumption \ref{ass:reweighting} intuitively means that given a certain plausible and correctly specified DGP, and given a measurable set $B\subset \mathbb{R}^p$ of possible values for the covariates $X$ which is attained with positive probability, i.e.\ $\mathbb P(X\in B)>0$, restricting the DGP to these values of covariates must be feasible.
E.g., if income $Y$ is studied in dependence of years after graduation, $X_1$, and further covariates $X_2, \ldots, X_p$, one might as well study income of persons at most 5 years after their graduation, $X_1\le 5$.
Then, in a correctly specified model, the true but unknown parameter $\theta_0$ remains the same, no matter whether considering the whole population or only persons within 5 years after their graduation.
Further recall that as discussed after \eqref{eqn:GPLLoss}, the characterization results for strictly consistent loss functions already rely on richness conditions on the classes of distributions.
\begin{theorem}
\label{thm:hierarchy consistency}
Under Assumption \ref{ass:unique model} the following holds for a loss $\rho\colon \mathbb{R}\times \Xi\to\mathbb{R}$.
\begin{enumerate}[label=\normalfont \rmfamily(\roman*)]
\item
If $\rho$ is (strictly) $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent for $\Gamma$ then it is (strictly) conditionally $\mathcal{F}_\mathcal{Z}$-model-consistent for the model $m$.
\item
Under Assumption \ref{ass:separability} if $\rho$ is conditionally $\mathcal{F}_\mathcal{Z}$-model-consistent for $m$,
there is a modification $\widetilde \mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ of $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ such that $\rho$ is $\widetilde \mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent for $\Gamma$.
\item
If $\rho$ is (strictly) conditionally $\mathcal{F}_\mathcal{Z}$-model consistent for $m$ then it is (strictly) unconditionally $\mathcal{F}_\mathcal{Z}$-model-consistent for $m$.
\item
Under Assumption \ref{ass:reweighting} if $\rho$ is (strictly) unconditionally $\mathcal{F}_\mathcal{Z}$-model-consistent for $m$ then it is also (strictly) conditionally $\mathcal{F}_\mathcal{Z}$-model-consistent for $m$.
\end{enumerate}
\end{theorem}
Theorem \ref{thm:hierarchy consistency}, whose proof can be found in Appendix \ref{sec:proof}, provides two main implications; see Figure \ref{fig:implications} for a visualisation:
First, a combination of (i) and (iii) justifies the use of strictly $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent losses for $\Gamma$ in the context of M-estimation. This is well known in the literature, e.g.\ \citet[Section 9]{GneitingRaftery2007} describe this under the term optimum score estimation.
The proofs of (i) and (iii) are straight-forward, and for special cases they can be found, e.g.\ in the proof of \citet[Theorem 1]{Patton2019}.
\begin{figure}
\centering
\scalebox{0.9}{
\begin{tikzpicture}
\node[state, minimum size=4cm] (q1) {\small \begin{tabular}{c} $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistency \\ for $\Gamma$ \end{tabular}};
\node[state, right of=q1, minimum size=4cm] (q2) {\small \begin{tabular}{c} conditional \\ $\mathcal{F}_{\mathcal{Z}}$-consistency \\ for $m$ \end{tabular}};
\node[state, right of=q2, minimum size=4cm] (q3) {\small \begin{tabular}{c} unconditional \\ $\mathcal{F}_{\mathcal{Z}}$-consistency \\ for $m$ \end{tabular}};
\draw
(q1) edge[->, bend left, above, line width=2pt] node{(i)} (q2)
(q2) edge[bend left, below, line width=2pt] node{(ii)} (q1)
(q2) edge[bend left, above, line width=2pt] node{(iii)} (q3)
(q3) edge[bend left, below, line width=2pt] node{(iv)} (q2);
\end{tikzpicture}
}
\caption{A visualisation of the implications of Theorem \ref{thm:hierarchy consistency}.}
\label{fig:implications}
\end{figure}
Second, and more important for our purposes is the reverse implication, combining (ii) and (iv).
It asserts that, under appropriate assumptions, an unconditionally $\mathcal{F}_{\mathcal{Z}}$-model-consistent loss for $m$ is necessarily $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent for $\Gamma$.
Thus, exploiting known characterization results for $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent losses for many relevant functionals $\Gamma$, it constitutes an effective and original bound on the class of consistent M-estimators.
Notice that \emph{strictness} of the $\widetilde \mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent losses cannot be established in part (ii) of Theorem \ref{thm:hierarchy consistency}.
While a stronger version of this result including the strictness would be desirable, its lack hardly diminishes the applicability of the results since characterization results for non-strict $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent losses are available in the literature and these are almost as strong as the ones for strictly consistent losses \citep{Gneiting2011}.
E.g., we obtain all consistent losses for the mean functional when possibly non-strictly convex functions $\phi$ are used in \eqref{eqn:BregmanLoss}, and for quantiles when possibly non-strictly increasing functions $g$ are used in \eqref{eqn:GPLLoss}.
Further note that for practical or intuitive purposes, the technical distinction between $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ and a modification $\widetilde \mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ thereof is inessential, see Appendix \ref{sec:App1}.
E.g., Theorem \ref{thm:hierarchy consistency} implies that the class of M-estimators for conditional mean models are characterized by loss functions of the form \eqref{eqn:BregmanLoss} while for conditional quantiles, a loss of the form \eqref{eqn:GPLLoss} must be used.
This enables to find beneficial choices of $\phi$, $g$ and $\kappa$ in terms of efficiency and equivariance \citep{DFZ2020} or integrability of the loss.
Similarly, one can leverage the flexibility of the class of consistent losses for VaR and ES provided in \cite{FisslerZiegel2016}.
As of yet, implications in the direction of the points (ii) and (iv) of Theorem \ref{thm:hierarchy consistency} have only been provided for special cases or under much stronger conditions:
First, \citet{Gourieroux1987} consider M-estimators for general semiparametric models by restricting attention to parameters identified by a set of conditional moment restrictions, as given by their equation (4.5).
As a consequence, their main result, Property 4.7, characterizes the functional form of the losses' first-order conditions and is hence closest to our Proposition S3 in the Supplementary Material.
In contrast, our Theorem \ref{thm:hierarchy consistency} allows us to conveniently connect M-estimation to known classes of strictly consistent loss functions from the literature on forecast evaluation.
\citet{Gourieroux1987} further operate under a stronger and less interpretable richness condition on the class of distributions; compare their Assumption A.7 i) to our Assumption \ref{ass:reweighting}.
Second, \citet[Theorem 2]{Komunjer2005} shows a necessary condition for M-estimation if $\Gamma$ is some quantile.
However, a richness condition, corresponding to our Assumption \ref{ass:reweighting}, is only assumed implicitly in the proof when quantifying over all model classes before their equation (19).
In contrast, our Theorem \ref{thm:hierarchy consistency} rigorously shows this relation for semiparametric models for any elicitable functional.
\section*{Acknowledgement}
T.~Dimitriadis gratefully acknowledges support of the German Research Foundation (DFG) through grant number 502572912, of the Heidelberg Academy of Sciences and Humanities and the Klaus Tschira Foundation. J.~Ziegel gratefully acknowledges support of the Swiss National Science Foundation.
We are very grateful to Jana Hlavinov\'a for a careful proofreading and valuable feedback on an earlier version of this paper.
\section*{Supplementary Material}
\label{SM}
The Supplementary Material derives results corresponding to Theorem \ref{thm:hierarchy consistency} for zero (Z) estimators and identification functions.