EconBase
← Back to paper

Characterizing M-estimators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

25,649 characters · 5 sections · 28 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Characterizing M-estimators

abstractWe characterize the full classes of M-estimators for semiparametric models of general functionals by formally connecting the theory of consistent loss functions from forecast evaluation with the theory of M-estimation. This novel characterization result opens up the possibility for theoretical research on efficient and equivariant M-estimation and, more generally, it allows to leverage existing results on loss functions known from the literature of forecast evaluation in estimation theory.

Keywords: M-estimation, loss function, strict consistency, characterization

Introduction

\onehalfspacing

The task of regression is to model the effect of covariates $X$ on a response variable $Y$, or more precisely, the effect of $X$ on a functional $\Gamma$ of the conditional distribution of $Y$ given $X$, $F_{Y|X}$. The typical example for $\Gamma$ is the mean, resulting in mean regression. In many application, one is interested in other functionals such as quantiles (Value at Risk, VaR), expectiles, variance, or Expected Shortfall (ES) Koenker1978, Bollerslev1986, Patton2019. A correctly specified parametric model $m(X,\theta)$ satisfies $\Gamma(F_{Y|X}) = m(X,\theta_0)$ for some unique parameter $\theta_0\in\Theta$.

The statistician's task is to estimate the parameter $\theta_0$ based on data $(Y_t,X_t)$, $t=1, \ldots, N$. For the standard situation of linear mean regression, $\Gamma(F_{Y|X}) = \mathbb{E}[Y|X]$, and $m(X, \theta) = X^\intercal \theta$, one often employs the ordinary-least-squares (OLS) estimator of the form $\widehat \theta_{T} = \operatorname*{arg\,min}_{\theta \in\Theta} \frac{1}{T} \sum_{t=1}^T \big(Y_t - X_t^\intercal \theta\big)^2$, or a related closed-form solution thereof. The OLS estimator is a special instance of an M-estimator Huber1967, NeweyMcFadden1994,

equation[equation omitted — 159 chars of source]

based on a loss function $\rho$, which is the key ingredient of an M-estimator. Note that our results carry over to time-varying loss functions $\rho_t$.

The core condition on $\rho$ for consistency of $\widehat \theta_{T}$ is that $\mathbb{E}\big[\rho \big(Y_t , m(X_t,\theta_0)\big)\big]<\mathbb{E}\big[\rho \big(Y_t , m(X_t,\theta)\big)\big]$ for all $\theta\neq\theta_0$ and for all $t\in\mathbb N$, which we call strict model-consistency of $\rho$ for $m$. Apart from special cases, the classes of such loss functions $\rho$ are unfortunately not well understood in M-estimation yet. Our main result, Theorem (ref), establishes that a loss function $\rho$ is (strictly) model-consistent for $m$ if and only if it is (strictly) consistent for the target functional $\Gamma$, meaning that $ \int \rho \big(y,\Gamma(F)\big)\mathrm{d} F(y) \le \int \rho(y,\xi)\mathrm{d} F(y) $ for all $\xi \in \mathbb{R}^k$ and for all distributions $F$ in a sufficiently rich class, where strictness means that equality implies $\xi=\Gamma(F)$. Since there are well-understood characterization results for strictly consistent losses from the literature of forecast evaluation Gneiting2011, FisslerZiegel2016, Theorem (ref) lifts these results to a novel characterization of the classes of consistent M-estimators for general (vector-valued) functionals.

Our result provides the first characterization of the full classes of consistent M-estimators for semiparametric models of general, possibly vector-valued functionals by formally connecting the two strands of literature on M-estimation and forecast evaluation. Understanding the full classes of M-estimators facilitates a deeper understanding of the (im)possibilities, e.g., in the following areas. It allows to derive M-estimation efficiency bounds, similar to efficient generalized method of moments estimation of Chamberlain1987. Drawing on the literature of homogeneous and equivariant loss functions NoldeZiegel2017, FisslerZiegel2019, our connection allows to classify such M-estimators having e.g. the advantage that they are invariant to linear rescaling in finite samples. Furthermore, estimators with the most beneficial integrability conditions in the sense that $\mathbb{E}[ \big| \rho \big(Y_t , m(X_t,\theta)\big) \big|]$ is finite can be derived. Our result further allows to leverage recent discoveries in the forecast evaluation literature such as mixture representations EhmETAL2016 or score decompositions DGJ_2021 for the purpose of M-estimation. Theorem (ref) also provides the first formal argument why consistent M-estimation is basically impossible if there do not exist (strictly) consistent loss functions for the target functional $\Gamma$. This argument has already been used informally for the Expected Shortfall Patton2019, DimiBayer2019, the Range Value at Risk Barendse2020, and the mode Kemp2012.

Notation and Definitions

Let $(\Omega, \mathcal A, \mathbb P)$ be a non-atomic, complete probability space where all random variables are defined. We introduce a class $\mathcal{Y}$ of $\mathbb{R}^d$-valued possible response variables, a corresponding class $\mathcal{X}$ or $\mathbb{R}^p$-valued regressors, and $\mathcal{Z}$ the class of possible response--regressor pairs $(Y,X)$. The class of marginal distributions $F_Y$ of $Y\in\mathcal{Y}$ is denoted by $\mathcal{F}_\mathcal{Y}$, with a corresponding notation $\mathcal{F}_\mathcal{X}$. $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ is the class of regular versions of conditional distributions $F_{Y|X}$ for any $(Y,X)\in \mathcal{Z}$; see Appendix (ref) for technical details. We will identify cumulative distribution functions with their corresponding measures where convenient. Let $\Gamma \colon\mathcal{F}_{\mathcal{Y}|\mathcal{X}}\to\Xi\subseteq\mathbb{R}^k$ be some $k$-dimensional, single-valued functional of the conditional distribution of $Y$ given $X$. Let $\Theta\subseteq \mathbb{R}^q$ be some parameter space with non-empty interior, $\operatorname{int}(\Theta)$, and $m\colon \mathbb{R}^p\times \Theta\to \Xi$ a parametric model for the functional $\Gamma$. We shall work under the following assumption of a correctly specified model with a uniquely identified true model parameter.

assumptionFor all $Z=(Y, X)\in\mathcal{Z}$ there is a unique parameter $\theta_0 = \theta_0(F_Z) \in \operatorname{int}(\Theta)$ such that almost surely \begin{equation} m(X,\theta_0) = \Gamma(F_{Y|X}). \end{equation}

Assumption (ref) is semiparametric in the sense that the finite-dimensional parameter $\theta_0$ in (ref) does not fully describe the distribution of $(Y,X)$, but only $\Gamma(F_{Y|X})$, the component of the conditional distribution we are essentially interested in. In general, an estimator should be valid on a class of random variables $\mathcal{Z}$ which is as large as possible, allowing to apply the estimator in many different situations and under uncertainty of the distributions of the underlying data. Hence, Assumption (ref) is a minimal condition on the random variables that allows for semiparametric modeling of a functional of the conditional distribution.

We continue by recalling the use of loss functions in the closely related area of forecast evaluation, where the notion of a strictly consistent loss function is a crucial concept in the literature on forecast evaluation, since it incentivizes truthful reports MurphyDaan1985. Making use of a similar decision-theoretic terminology as in Gneiting2011 and FisslerZiegel2016, let $\mathcal{F}$ be some generic class of probability distributions on $\mathbb{R}^d$, which is our {observation domain}, and $\Xi\subseteq \mathbb{R}^k$, our {action domain}.

definition[Consistency and elicitability] A loss function $\rho\colon \mathbb{R}^d\times \Xi\to\mathbb{R}$ is called $\mathcal{F}$-consistent for a functional $\Gamma\colon \mathcal{F}\to\Xi$ if $\rho(\cdot, \xi)$ is $F$-integrable for all $F\in\mathcal{F}$ and for all $\xi\in\Xi$, and if \begin{equation} \int \rho\big(y,\Gamma(F)\big) \,\mathrm{d} F(y) \le \int \rho (y, \xi)\,\mathrm{d} F(y) \qquad for all F\in\mathcal{F}, \ for all \xi\in\Xi\,. \end{equation} If equality in (ref) implies $\xi = \Gamma(F)$, then the loss function is called strictly $\mathcal{F}$-consistent for $\Gamma$. A functional $\Gamma\colon \mathcal{F}\to\Xi$ is elicitable if there is a strictly $\mathcal{F}$-consistent loss function for it.

On the class of distributions with a finite second moment, the squared loss $\rho(y,\xi) = (y-\xi)^2$ is strictly consistent for the mean functional. More generally, subject to regularity and integrability conditions on $\rho$ and richness conditions on $\mathcal{F}$, $\rho$ is (strictly) $\mathcal{F}$-consistent for the mean if and only if it is a so-called Bregman loss

align[align omitted — 107 chars of source]

where $\phi$ is a (strictly) convex function on $\mathbb{R}$ with subgradient $\phi'$ and $\kappa$ is any function of $y$ Savage1971, Gneiting2011. Likewise, the well known pinball or asymmetric absolute loss $\rho(y,\xi) = (\mathds{1}\{y\le \xi\} - \alpha)(\xi - y)$ is strictly consistent for the lower $\alpha$-quantile on the class of distributions with finite mean and where the lower $\alpha$-quantile coincides with the upper $\alpha$-quantile. Moreover, subject to regularity and integrability conditions on $\rho$ and richness conditions on $\mathcal{F}$, a loss $\rho$ is (strictly) consistent for the lower $\alpha$-quantile, $\alpha\in(0,1)$, if and only if it is a generalized piecewise linear loss function

align[align omitted — 113 chars of source]

where $g$ is (strictly) increasing Gneiting2011b.

Similar characterization results exist for other functionals such as expectiles or the pairs consisting of the mean and variance or the quantile and Expected Shortfall Gneiting2011, FisslerZiegel2016. The characterization results in (ref) and (ref) rely on the fact that the related classes $\mathcal{F}$ are convex and rich enough. E.g., if we restrict attention to symmetric distributions only, the mean equals the median, and loss functions of the form given in (ref) or (ref) would elicit the mean and median.

The following definition develops similar notions of consistency for the setting of M-estimation with an underlying class $\mathcal{Z}$ implying that $\Gamma$ is defined on the class of conditional distributions $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$.

definition[Model-consistency] Suppose Assumption (ref) holds for the parametric model $m\colon\mathbb{R}^p\times\Theta\to\Xi$ and the functional $\Gamma\colon\mathcal{F}_{\mathcal{Y}|\mathcal{X}}\to\Xi$. Let $\rho\colon \mathbb{R}\times \Xi\to\mathbb{R}$ be a loss function such that $\mathbb{E}\big|\rho\big(Y,m(X,\theta)\big)\big|<\infty$ for all $(Y,X) \in\mathcal{Z}$ and for all $\theta\in\Theta$. \begin{enumerate}[label=(\roman*)] • The loss $\rho$ is unconditionally $\mathcal{F}_\mathcal{Z}$-model-consistent for the model $m$ if \begin{equation} \mathbb{E}\big[\rho\big(Y,m(X,\theta_0)\big)\big]\le \mathbb{E}\big[\rho\big(Y,m(X,\theta)\big)\big] \qquad for all (Y,X)\in\mathcal{Z}, \ for all \theta\in\Theta\,. \end{equation} Moreover, $\rho$ is strictly unconditionally $\mathcal{F}_\mathcal{Z}$-model-consistent for the model $m$ if equality in (ref) implies that $\theta = \theta_0$. • The loss $\rho$ is conditionally $\mathcal{F}_\mathcal{Z}$-model-consistent for the model $m$ if \begin{equation} \mathbb{E}\big[\rho\big(Y,m(X,\theta_0)\big)\big|X\big]\le \mathbb{E}\big[\rho\big(Y,m(X,\theta)\big)\big|X\big]\ a.s. \quad \text{for all } (Y,X)\in\mathcal{Z}, \ \text{for all } \theta\in\Theta\,. \end{equation} Moreover, $\rho$ is \emph{strictly conditionally $\mathcal{F}_\mathcal{Z}$-model-consistent} for the model $m$ if almost sure equality in (ref) implies that $\theta=\theta_0$. \end{enumerate}

The concept of unconditional model-consistency is the central condition for consistent M-estimation: Gourieroux1987 show equivalence of these two notions in terms of first order conditions and under some regularity assumptions. Also see condition (i) in NeweyMcFadden1994, Huber1967, where their remaining assumptions are merely regularity conditions. In contrast, the conditional notion is appealing since it bridges the gap between consistency for $\Gamma$ according to Definition (ref) and unconditional model-consistency for $m$ in the proof of Theorem (ref) below. It can still be practically useful when we resort to nonparametric kernel regressions or in the presence of repeated observations of $X$, e.g.\ if $X$ consists of categorical variables only.

While the classical notion of consistency in (ref) and the unconditional model version in (ref) are closely related, they are crucially different in that in the latter, the expectation is also taken with respect to the covariates. Establishing and finding reasonable conditions for their equivalence is indeed not trivial as shown in Theorem (ref) below.

Main Result and Discussion

To present our main Theorem (ref), we introduce and discuss the following two assumptions.

assumptionFor all $X\in\mathcal{X}$, the map $m(X,\cdot)\colon\Theta\to\Xi$ is surjective almost surely. For all $(Y,X)\in\mathcal{Z}$ the conditional expectation $\mathbb{E}\big[\rho\big(Y,m(X,\theta)\big)\big|X\big]$ is continuous in $\theta$ almost surely.

The surjectivity in Assumption (ref) can usually be fulfilled by a sensible choice of $\Xi$, and the smoothness condition on the expected loss is standard in the literature, see e.g., NeweyMcFadden1994.

assumptionFor any $Z = (Y,X)\in\mathcal{Z}$ and any event $A\in \sigma(X)$ with positive probability $\mathbb P(A)>0$, there is some $\widetilde Z\in\mathcal{Z}$ such that $\mathbb P\big( \widetilde Z \in B\big) = \mathbb P\big(Z\in B\,|A\big)$ for all Borel sets $B\subseteq \mathbb{R}^{d+p}$.

Assumption (ref) is a richness condition on the class of possible data generating processes (DGP) in $\mathcal{Z}$ as for any process $Z = (Y,X) \in \mathcal{Z}$, and any set $A$ with positive probability, it stipulates that $\mathcal{Z}$ is rich enough to contain a process $\widetilde Z = (\widetilde Y, \widetilde X)$ as specified in Assumption (ref). Crucially, it yields that $\mathbb P(\widetilde Y\in C\, | \widetilde X) = \mathbb P(Y\in C\, | X,A)$ for all Borel sets $C\subseteq \mathbb{R}$. This, together with Assumption (ref), implies that the correctly specified parameter and hence the semiparametric model, is the same under the distributions $F_{\widetilde Z}$ and $F_Z$; in formulae, $\theta_0(F_{\widetilde Z}) = \theta_0(F_Z)$. Recall that in estimation, $\mathcal{Z}$ captures the flexibility about the underlying and in practice unknown DGP, such that a large $\mathcal{Z}$ is desirable in order to obtain an estimation method which is applicable to a wide range of distributions of $Z \in \mathcal{Z}$. Thus, Assumption (ref) intuitively means that given a certain plausible and correctly specified DGP, and given a measurable set $B\subset \mathbb{R}^p$ of possible values for the covariates $X$ which is attained with positive probability, i.e.\ $\mathbb P(X\in B)>0$, restricting the DGP to these values of covariates must be feasible. E.g., if income $Y$ is studied in dependence of years after graduation, $X_1$, and further covariates $X_2, \ldots, X_p$, one might as well study income of persons at most 5 years after their graduation, $X_1\le 5$. Then, in a correctly specified model, the true but unknown parameter $\theta_0$ remains the same, no matter whether considering the whole population or only persons within 5 years after their graduation. Further recall that as discussed after (ref), the characterization results for strictly consistent loss functions already rely on richness conditions on the classes of distributions.

theoremUnder Assumption (ref) the following holds for a loss $\rho\colon \mathbb{R}\times \Xi\to\mathbb{R}$. \begin{enumerate}[label=\normalfont (\roman*)] • If $\rho$ is (strictly) $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent for $\Gamma$ then it is (strictly) conditionally $\mathcal{F}_\mathcal{Z}$-model-consistent for the model $m$. • Under Assumption (ref) if $\rho$ is conditionally $\mathcal{F}_\mathcal{Z}$-model-consistent for $m$, there is a modification $\widetilde \mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ of $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ such that $\rho$ is $\widetilde \mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent for $\Gamma$. • If $\rho$ is (strictly) conditionally $\mathcal{F}_\mathcal{Z}$-model consistent for $m$ then it is (strictly) unconditionally $\mathcal{F}_\mathcal{Z}$-model-consistent for $m$. • Under Assumption (ref) if $\rho$ is (strictly) unconditionally $\mathcal{F}_\mathcal{Z}$-model-consistent for $m$ then it is also (strictly) conditionally $\mathcal{F}_\mathcal{Z}$-model-consistent for $m$. \end{enumerate}

Theorem (ref), whose proof can be found in Appendix (ref), provides two main implications; see Figure (ref) for a visualisation: First, a combination of (i) and (iii) justifies the use of strictly $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent losses for $\Gamma$ in the context of M-estimation. This is well known in the literature, e.g.\ GneitingRaftery2007 describe this under the term optimum score estimation. The proofs of (i) and (iii) are straight-forward, and for special cases they can be found, e.g.\ in the proof of Patton2019.

figure[figure omitted — 965 chars of source]

Second, and more important for our purposes is the reverse implication, combining (ii) and (iv). It asserts that, under appropriate assumptions, an unconditionally $\mathcal{F}_{\mathcal{Z}}$-model-consistent loss for $m$ is necessarily $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent for $\Gamma$. Thus, exploiting known characterization results for $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent losses for many relevant functionals $\Gamma$, it constitutes an effective and original bound on the class of consistent M-estimators. Notice that strictness of the $\widetilde \mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent losses cannot be established in part (ii) of Theorem (ref). While a stronger version of this result including the strictness would be desirable, its lack hardly diminishes the applicability of the results since characterization results for non-strict $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent losses are available in the literature and these are almost as strong as the ones for strictly consistent losses Gneiting2011. E.g., we obtain all consistent losses for the mean functional when possibly non-strictly convex functions $\phi$ are used in (ref), and for quantiles when possibly non-strictly increasing functions $g$ are used in (ref). Further note that for practical or intuitive purposes, the technical distinction between $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ and a modification $\widetilde \mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ thereof is inessential, see Appendix (ref).

E.g., Theorem (ref) implies that the class of M-estimators for conditional mean models are characterized by loss functions of the form (ref) while for conditional quantiles, a loss of the form (ref) must be used. This enables to find beneficial choices of $\phi$, $g$ and $\kappa$ in terms of efficiency and equivariance DFZ2020 or integrability of the loss. Similarly, one can leverage the flexibility of the class of consistent losses for VaR and ES provided in FisslerZiegel2016.

As of yet, implications in the direction of the points (ii) and (iv) of Theorem (ref) have only been provided for special cases or under much stronger conditions: First, Gourieroux1987 consider M-estimators for general semiparametric models by restricting attention to parameters identified by a set of conditional moment restrictions, as given by their equation (4.5). As a consequence, their main result, Property 4.7, characterizes the functional form of the losses' first-order conditions and is hence closest to our Proposition S3 in the Supplementary Material. In contrast, our Theorem (ref) allows us to conveniently connect M-estimation to known classes of strictly consistent loss functions from the literature on forecast evaluation. Gourieroux1987 further operate under a stronger and less interpretable richness condition on the class of distributions; compare their Assumption A.7 i) to our Assumption (ref). Second, Komunjer2005 shows a necessary condition for M-estimation if $\Gamma$ is some quantile. However, a richness condition, corresponding to our Assumption (ref), is only assumed implicitly in the proof when quantifying over all model classes before their equation (19). In contrast, our Theorem (ref) rigorously shows this relation for semiparametric models for any elicitable functional.

Acknowledgement

T. Dimitriadis gratefully acknowledges support of the German Research Foundation (DFG) through grant number 502572912, of the Heidelberg Academy of Sciences and Humanities and the Klaus Tschira Foundation. J. Ziegel gratefully acknowledges support of the Swiss National Science Foundation. We are very grateful to Jana Hlavinov\'a for a careful proofreading and valuable feedback on an earlier version of this paper.

Supplementary Material

The Supplementary Material derives results corresponding to Theorem (ref) for zero (Z) estimators and identification functions.