EconBase
← Back to paper

Kullback-Leibler-based characterizations of score-driven updates

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

100,998 characters · 0 sections · 100 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Expected Kullback-Leibler-based characterizations of score-driven updates

center[center omitted — 545 chars of source]

\footnotetext[1]{ E-mail addresses: [email removed] (Ramon de Punder), [email removed] (Timo Dimitriadis), [email removed] (Rutger-Jan Lange).}

abstractScore-driven (SD) models are a standard tool in statistics and econometrics, with applications in hundreds of published articles in the past decade. We provide an information-theoretic characterization of SD updates based on reductions in the expected Kullback–Leibler (EKL) divergence relative to the true---but unknown---data-generating density. EKL reductions occur if and only if the expected update direction aligns with the expected score; i.e., their inner product should be positive. This equivalence condition uniquely identifies SD updates (including scaled or clipped variants) as being EKL reducing, even in non-concave, multivariate, and misspecified settings. We further derive explicit bounds on admissible learning rates in terms of score moments, linking SD methods to adaptive optimization techniques. By contrast, alternative performance measures in the literature impose stronger conditions (e.g., concave logarithmic densities) and do not {\it characterize} SD updates: other updating rules may improve these measures, while SD updates need not. Our results provide a rigorous justification for SD models and establish EKL as their natural information-theoretic foundation.

{\it Keywords:} generalized autoregressive score (GAS); dynamic conditional score (DCS); Kullback Leibler; scoring rule, divergence

{8.75in} \setlength\abovedisplayskip{6pt} \setlength\belowdisplayskip{6pt}

\normalem

bibunit[chicago] \section{Introduction} The use of score-driven (SD) models has proliferated over the last decade. They were originally introduced by creal2013generalized and harvey2013dynamic and known by different names and acronyms; recent literature (e.g., artemova2022score1,artemova2022score2; harvey2022score) has converged on the {terminology} of SD models. These models specify a distribution with time-varying parameters---governing, for example, intensity, location, scale, or shape---whose dynamics are driven by the score, the derivative of the log-likelihood with respect to the time-varying parameter vector. SD updates can be interpreted as stochastic gradient ascent steps that adjust the parameter after each observation. Unlike in optimization, the parameter does not converge but remains perpetually responsive. Applications are numerous; see \href{www.gasmodel.com}{www.gasmodel.com} for a list of over 400 publications. Much of this literature assumes that the SD filter coincides with the true data-generating process. Yet it remains unclear whether SD updates possess theoretical properties that uniquely characterize them in more general, possibly misspecified settings, and, if so, what these are. In this paper, we resolve this question by showing that, in expectation, sufficiently small parameter adjustments improve the distributional fit if and only if the update direction is driven by the score, precisely the principle on which SD models are based. To explain this result, let $y_t \in \mathcal{Y}$ with $t \in \mathbb{N}$ denote the realization of $Y_t$ with true density $p_t$ given the information at time $t-1$. The researcher postulates (typically misspecified) parametric densities $f_{t|t-1}(\cdot) \equiv f(\cdot | \vartheta_{t|t-1})$ before updating and $f_{t|t}(\cdot) \equiv f(\cdot | \vartheta_{t|t})$ after updating. Here, $\vartheta_{t|t-1}$ and $\vartheta_{t|t} \equiv \vartheta_{t|t}(y_t)$ denote the predicted and updated parameters based on information up to times $t-1$ and $t$, respectively; i.e., the update takes account of the latest observation $y_t$. The SD update is \begin{equation} \vartheta_{t|t} \equiv \vartheta_{t|t}(y_t) = \vartheta_{t|t-1} + A \mathcal{S}_{t-1} \, {s}(y_t,\vartheta_{t|t-1}), \end{equation} where ${s}(y_t,\vartheta_{t\vert t-1}) := (\partial/\partial\vartheta) \log f(y_t \vert \vartheta)\vert_{\vartheta_{t\vert t-1}}$ is the score (i.e., a stochastic gradient). As is standard (e.g., creal2013generalized), the update involves a static learning-rate matrix $A$ and a dynamic scaling matrix $\mathcal{S}_{t-1}$, assumed known at time $t-1$. We assume their product $A \mathcal{S}_{t-1}$ to be positive definite, which is not automatic even if both $A$ and $\mathcal{S}_{t-1}$ are positive definite. The prediction step used to construct $\vartheta_{t+1|t}$ from $\vartheta_{t|t}$ is not essential at present; its discussion is postponed until after Definition (ref). To quantify the distributional closeness between $f_{t\vert t}$ and $p_t$, we propose to use the expected KL (EKL) divergence, where, in addition to the usual integral, the observation $y \in {\mathcal{Y}}$ used in the update $\vartheta_{t|t}(y)$ and thus in $f_{t\vert t}(\cdot) \equiv f(\cdot|\vartheta_{t|t}(y))$ is averaged out using the true density: \begin{align} \mathsf{EKL}(p_t \Vert f_{{t\vert t}}) := \int_{{\mathcal{Y}}} \int_{{\mathcal{Y}}} \log \!\left(\frac{p_t(x)}{f(x \vert \vartheta_{t|t}(y))} \right) p_t(x)\, p_t(y)\, \mathrm{d} x \,\mathrm{d} y. \end{align} The EKL divergence measure (ref) features a double integral over $x$ and $y$, admitting a natural two-sample interpretation: one draw $y$ updates the model density, and an independent redraw $x$ evaluates the updated model’s fidelity on new data. By averaging out the uncertainty in both samples, it provides a natural criterion for evaluating updating rules and combines the realized KL perspective of blasques2015information with the expectation-based measures of Gorgi2023 and Creal2024GMM (see related literature below). As Theorems (ref)--(ref) show, under various conditions on the score's derivative (the Hessian), any sufficiently small parameter change $\vartheta_{t\vert t}(Y_t) - \vartheta_{t\vert t-1}$ implies an EKL improvement, \begin{align} \mathsf{EKL}(p_t \Vert f_{{t\vert t}}) < \mathsf{EKL}(p_t \Vert f_{{t\vert t-1}}), \end{align} if and only if the expected parameter adjustment is aligned with the expected score, i.e., \begin{align} \mathbb{E}_{p_t} \!\big[ \vartheta_{t\vert t}(Y_t) - \vartheta_{t\vert t-1} \big]^\top \mathbb{E}_{p_t} \!\big[ {s}(X_t,\vartheta_{t\vert t-1}) \big] > 0, \end{align} where $\mathbb{E}_{p_t}[\cdot]$ denotes the expectation under the true density $p_t$. The equivalence between (ref) and (ref) provides an information-theoretic characterization of a family of updating rules. We refer to any (user-specified) updating rule $\vartheta_{t|t}(Y_t)$ satisfying (ref) as being score equivalent in expectations (SEE); we use the plural as (ref) involves two distinct expectations. Since $p_t$ is treated as unknown, we would ideally like (ref) to hold for broad classes of $p_t$. For SD updates (ref), in which $\mathbb{E}_{p_t}[\vartheta_{t\vert t}(Y_t)-\vartheta_{t\vert t-1}] = A \mathcal{S}_{t-1} \mathbb{E}_{p_t}[s(Y_t,\vartheta_{t\vert t-1})]$, condition (ref) simplifies to \begin{equation} \mathbb{E}_{p_t}[s(Y_t,\vartheta_{t\vert t-1})]^\top A \mathcal{S}_{t-1} \mathbb{E}_{p_t}[s(Y_t,\vartheta_{t\vert t-1})] > 0, \end{equation} which holds independently of the true law $p_t$ whenever the expected score is non-zero. Hence, sufficiently small SD updates with non-zero expectation are EKL reducing. Indeed, SD updates may be the practically most important class of SEE updates. Positive definiteness (and thus symmetry) of $A\mathcal{S}_{t-1}$, although not customary (e.g., gasperoni_score-driven_2023,d2024modeling), is essential in producing gradient-ascent directions and EKL improvement. The condition that the expected score at $\vartheta_{{t\vert t-1}}$ is nonzero can be interpreted as requiring that $\vartheta_{{t\vert t-1}}$ is not a stationary point of the map $\vartheta \mapsto \mathbb{E}_{p_t}[\log f(Y_t| \vartheta)]$; in particular, no improvement guarantees can be obtained if the prediction $\vartheta_{t\vert t-1}$ is already optimal. The broad scope of Theorems (ref)--(ref) stems from the intrinsic relationship between the EKL criterion (ref) and the driving mechanism in SD models, namely the score. The score is the derivative of the log likelihood, while the EKL criterion can be written in terms of the expected log-likelihood contribution; hence, they are intimately related. This connection allows our characterization results to extend to multivariate settings under mild conditions, accommodating misspecification, non-concave model log-densities, and general (positive-definite) matrix combinations $A \mathcal{S}_{t-1}$. As we show in Section (ref), other performance criteria for SD updates are less directly tied to the score and more restricted in scope. Our improvement guarantee in (ref) requires that the parameter adjustment be sufficiently small. To provide practical guidance on the admissible step size, Theorem (ref) further derives (non-infinitesimal) upper bounds on elements or eigenvalues of the matrix combination $A\mathcal{S}_{t-1}$ that still ensure EKL improvements. These bounds depend on the first two population moments of the score, sharpening the constant bound of Gorgi2023 and motivating the use of adaptive, moment-based learning rates for SD models, in analogy with adaptive methods in the optimization literature. \begin{table}[tb] \caption{Overview of different performance criteria considered in this paper. } \begin{footnotesize} \begin{threeparttable} \begin{tabular}{ll cc l} \toprule \multirow{2}{*}{Criterion} & \multirow{2}{*}{Reference} & Proper & Constructive & \multicolumn{1}{c}{Assumption related to} \\ & & criterion & “$\Longleftrightarrow$” & \multicolumn{1}{c}{expected Hessian}\\ \midrule $\mathsf{EKL}$ & this paper: Theorem (ref) & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & bounded ((ref)) \\ $\mathsf{EKL}$ & this paper: Theorem (ref) & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & locally bounded ((ref)) \\ \midrule $\mathsf{CEV}$ & Gorgi2023 & \color{mygreen} \ding{51} & \color{red} \ding{55} & negative definite ((ref))\\ $\mathsf{MSE}$ & Gorgi2023 & \color{mygreen} \ding{51} & \color{red} \ding{55} & negative definite ((ref)) \\ $\mathsf{EGMM}$ & Creal2024GMM & \color{mygreen} \ding{51} & \color{red} \ding{55} & bounded & negative definite ((ref)) \\ \midrule ${\mathsf{TKL}}$ & blasques2015information & \color{red} \ding{55} & \\ ${\mathsf{CKL}}$ & this paper: Appendix (ref) & \color{mygreen} \ding{51} & \color{red} \ding{55} & \\ \bottomrule \end{tabular} {NOTE: As the TKL measure is not a proper divergence measure, the remaining columns are left blank. For EGMM, we additionally need a uniform bound on the third derivative of the log-density. Table (ref) verifies these assumptions for specific model densities. } \end{threeparttable} \end{footnotesize} \end{table} Related literature. We compare the EKL measure with four related performance measures for SD models proposed in three closely related papers;\footnote{ These papers focus on providing improvement guarantees for a single time step. This differs from alternative approaches based on in-fill asymptotics (e.g., beutner2023consistency), which use the Kullback–Leibler divergence to characterize the pseudo-true path in the continuous-time limit.} see Table (ref) and Section (ref). First, Gorgi2023 show that SD updates move the parameter toward the pseudo-true value, as measured by the first and second moments of the update around the pseudo-true parameter, yielding their conditional expected variation (CEV) and mean squared error (MSE) criteria. Second, Creal2024GMM define an expected generalized method-of-moments (EGMM) loss using the expected score as a moment condition and show that SD updates reduce this loss under a scaling that requires a (practically infeasible) expectation under the true density. Third, blasques2015information propose a trimmed KL (TKL) measure which, despite known issues (e.g., blasques2018information), remains a standard reference for motivating SD models (e.g., holy2022modeling, delle2023modeling, gasperoni_score-driven_2023, and catania2026unobserved). For the first two related papers (i.e., Gorgi2023 and Creal2024GMM), we show in Sections (ref)--(ref) that their CEV, MSE, and EGMM performance guarantees for SD updates effectively require log-concave model densities, thereby excluding many fat-tailed distributions (e.g., the Student’s $t$ distribution with time-varying location) that fall within the scope of our EKL criterion. Moreover, neither Gorgi2023 nor Creal2024GMM fully characterize the model classes that improve their performance criteria; here, we generalize their results by establishing characterizations in a multivariate setting (Propositions (ref)--(ref)). In the multivariate case, we show that the approaches in Gorgi2023 and Creal2024GMM impose stringent restrictions on the matrix product $A\mathcal{S}_{t-1}$, often forcing it to be a scalar multiple of the identity. Moreover, the associated equivalence conditions are of limited use for model construction, as they effectively prescribe adjustment toward the \mbox{(pseudo-)}true parameter. The relatively greater complexity and narrower scope of these results highlight the appeal of our SEE condition (ref), which via (ref) naturally leads to the general class of (multivariate) SD updates. For the third related paper (i.e., blasques2015information), we show in Section (ref) that the TKL measure is based on a localized, trimmed KL divergence, which is known to be problematic in the scoring-rule literature (e.g., diks2011likelihood,gneiting_comparing_2011). Consequently, the associated equivalence condition is insensitive to the true density $p_t$ and thus unsuitable for gauging closeness to $p_t$. To address this within the localization framework of blasques2015information, we replace trimming with censoring, obtaining a censored KL (CKL) measure that preserves the fundamental properties of a KL divergence, but still yields an unhelpful equivalence condition (see Appendices (ref) and (ref) for details). \textbf{Notation.} Vectors are columns and $\|\cdot\|$ denotes the Euclidean norm for vectors and the spectral norm, that is, the operator norm induced by the Euclidean norm, for matrices. For a matrix $A$, $\operatorname{tr}(A)$ is its trace, and $\lambda_{\min}(A), \lambda_{\max}(A)$ its extremal eigenvalues. We use $O_k$ and $I_k$ for the $k \times k$ zero and identity matrix, respectively. The Loewner order for symmetric matrices is denoted $A \succeq (\succ)\, B$. Note that $-cI_k \preceq A \preceq cI_k$ if and only if $\|A\|\leq c$ for $c>0$. \section{Expected Kullback-Leibler reducing updates} After introducing notation and preliminaries in Section (ref), we establish our main result in Section (ref), whose assumptions are relaxed in Section (ref). \subsection{Preliminaries} We consider an outcome space ${\mathcal{Y}} \subseteq \mathbb{R}^l$ with $l \in \mathbb{N}$ and a stochastic process $\{Y_t: \Omega \to {\mathcal{Y}}\}_{t=1}^T$ on a complete probability space $(\Omega, {\mathcal{F}}, \mathbb{P})$ with a measurable space $({\mathcal{Y}}^T, {\mathcal{B}}({\mathcal{Y}}^T))$. We consider the flexible and user-chosen sigma-algebra (or information set) ${\mathcal{F}}_{t-1}$, which is such that $\sigma(Y_s: s \le t-1) \subseteq {\mathcal{F}}_{t-1} \subseteq {\mathcal{F}}$, but $Y_t$ is not ${\mathcal{F}}_{t-1}$-measurable. While we typically have ${\mathcal{F}}_{t-1} = \sigma(Y_s : s \le t-1)$ as, e.g., in Gorgi2023 and Creal2024GMM, allowing ${\mathcal{F}}_{t-1}$ to differ from $\sigma(Y_s : s \le t-1)$ is useful for incorporating external covariates or when the data are generated by a state-space model, in which case the true latent state at time $t$ can be included in ${\mathcal{F}}_{t-1}$ (for details, see Appendix (ref)). The true conditional law of $Y_t$ given ${\mathcal{F}}_{t-1}$ is $P_t \in {\mathcal{P}} \subseteq {\mathcal{P}}_0$, where ${\mathcal{P}}_0$ is the class of absolutely continuous distributions, with density $p_t$. The subclass ${\mathcal{P}}$ may vary across the results below. Probabilities and variances with respect to $p_t$ are denoted $\mathbb{P}_{p_t}(\cdot)$ and $\mathbb{V}_{p_t}(\cdot)$, respectively, and we occasionally write $Y_t \sim p_t$ or $p_t \in {\mathcal{P}}$. Statements involving random variables are meant to hold $P_t$-almost surely (a.s.) unless stated otherwise. The researcher’s (possibly misspecified) predictive density for $Y_t$ is $f_{t\vert t-1}(\cdot) \equiv f(\cdot|\vartheta_{t\vert t-1})$, where $\vartheta_{t\vert t-1} \in \Theta \subseteq \mathbb{R}^k$ is based on $\mathcal{F}_{t-1}$ and $\Theta$ is open and convex. A link function may be embedded in $f(\cdot|\vartheta_{t\vert t-1})$, mapping $\vartheta_{t\vert t-1}$ into some desired domain harvey2022score. Throughout, we assume that the support of $p_t$ is a subset of the support of $f(\cdot|\vartheta)$, for all $\vartheta \in \Theta$. Dependence on static parameters or exogenous covariates known at time $t-1$ is permitted but suppressed for readability. Given the realization $y_t$ of $Y_t$, the updated parameter is $\vartheta_{t\vert t} = \phi(y_t,\vartheta_{t\vert t-1})$, where $\phi:({\mathcal{Y}},\Theta)\to\Theta$ is the \emph{updating rule}. We write $\Delta\phi(y_t,\vartheta_{t\vert t-1}):=\vartheta_{t\vert t}-\vartheta_{t\vert t-1}$. To guarantee $\vartheta_{t\vert t}\in\Theta$ almost surely, we typically take $\Theta=\mathbb{R}^k$ and employ a suitable link function if needed. Following lange2022robust, we distinguish between the \emph{update} step, which yields $\vartheta_{t| t}$, and the \emph{prediction} step, which yields $\vartheta_{t+1 | t}$. The update $\vartheta_{t| t}$ is designed to improve the fit to the current density $p_t$ from which $y_t$ is drawn, whereas the prediction $\vartheta_{t+1| t}$ is aimed at the next-period density $p_{t+1}$, for which no observations have (yet) been collected. Ideally, the updated density $f_{{t\vert t}}(\cdot)\equiv f(\cdot|\vartheta_{t\vert t})$ improves upon $f_{t\vert t-1}(\cdot)$ by maximizing $\mathbb{E}_{p_t}[\log f(Y_t|\vartheta)]$ over $\vartheta \in \Theta$. However, this expectation is unobserved and cannot be estimated, as only a single realization $y_t$ of $Y_t$ is available. Stochastic-gradient methods therefore rely on the observed gradient, the \emph{score}, \[ {s}(y_t,\vartheta_{t\vert t-1}):=\left.\frac{\partial}{\partial\vartheta}\log f(y_t | \vartheta)\right|_{\vartheta_{t\vert t-1}}, \] leading to the class of score-driven (SD) filters (e.g., creal2013generalized, harvey2013dynamic). \begin{definition}[Score-driven update] The SD update is \begin{align} \vartheta_{t\vert t}^\mathrm{SD} = \phi_{\mathrm{SD}}(y_t,\vartheta_{t\vert t-1}) := \vartheta_{t\vert t-1} + A \mathcal{S}_{t-1} {s}(y_t,\vartheta_{t\vert t-1}), \end{align} with static learning-rate matrix $A$ and $\mathcal{F}_{t-1}$-measurable and observable scaling matrix $\mathcal{S}_{t-1}$, such that $A\mathcal{S}_{t-1} \succ O_k$. \end{definition} The scaling $\mathcal{S}_{t-1}$ is often based on powers of the Fisher information matrix (e.g., creal2013generalized; artemova2022score1), which is then assumed to be non-singular, i.e., \begin{equation} \mathcal{S}_{t-1}= \Big(\int_{{\mathcal{Y}}} s(y,\vartheta_{t\vert t-1})s(y,\vartheta_{t\vert t-1})^\top f(y|\vartheta_{t\vert t-1})\,\,\mathrm{d}y\Big)^{-\zeta}, \quad \zeta\in\{0,1/2,1\}. \end{equation} The scaling is based only on the postulated (not the true) density. For $\zeta=1/2$, we take the inverse of the (unique) positive-definite square root of the Fisher information, ensuring $\mathcal{S}_{t-1}\succ O_k$. However, positive definiteness of $A\mathcal{S}_{t-1}$ is not automatic even if both $A\succ O_k$ and $\mathcal{S}_{t-1}\succ O_k$ nicholson1979eigenvalue. In fact, the interpretation of SD updates as performing steepest ascent in the metric induced by $A\mathcal{S}_{t-1}$ requires this matrix to be positive definite (and therefore symmetric). Interestingly, much of the SD literature ignores this point.\footnote{E.g., creal2013generalized, gasperoni_score-driven_2023, and d2024modeling do not impose $A\mathcal{S}_{t-1}\succ O_k$. Symmetrized alternatives such as $\mathcal{S}_{t-1} A\mathcal{S}_{t-1}$ guarantee positive definiteness if $\mathcal{S}_{t-1}, A\succ O_k$, but deviate from the standard SD setup.} In practice, $A\mathcal{S}_{t-1}\succ O_k$ can be ensured if (i) $A$ and $\mathcal{S}_{t-1}$ are diagonal with positive entries, or (ii) one of them is positive definite, while the other is a positive multiple of the identity. After updating, predictions are constructed as $\vartheta_{t+1|t} = \omega + B \vartheta_{t|t}$, for some vector $\omega$ and matrix $B$. Combining update and prediction steps yields $\vartheta_{t+1|t} = \omega + B(\vartheta_{t|t-1} + A \mathcal{S}_{t-1} {s}(y_t,\vartheta_{t\vert t-1}))$, which up to reparameterization is the standard SD formulation. As our focus is on the update, however, we maintain a clean separation between both steps. \subsection{Main result: Only SD updates are EKL reducing} We now establish that only SD updates (and their equivalents) guarantee \emph{expected} Kullback--Leibler (EKL) improvement. To this end, we examine the EKL measure $ \mathbb{E}_{p_t}\!\left[\mathsf{KL}\big(p_t \Vert f_{t\vert t}\big)\right], $ where $f_{t\vert t}(\cdot) \equiv f(\cdot|\vartheta_{{t\vert t}}(Y_t))$ is evaluated at $Y_t \sim p_t$. More explicitly, recalling (ref), \begin{align} \mathsf{EKL}(p_t \Vert f_{{t\vert t}}) := \int_{{\mathcal{Y}}} \int_{{\mathcal{Y}}} \log \!\left(\frac{p_t(x)}{f(x | \vartheta_{t|t}(y))} \right) p_t(x)\, p_t(y)\, \mathrm{d} x \,\mathrm{d} y, \end{align} which we assume to be finite throughout. Here the dependence of $\vartheta_{t|t}(y) = \phi(y, \vartheta_{t\vert t-1})$ on the observation $y$ is made explicit. The assumed independence of the random variables underlying the observation $y$ (driving the update) and the hypothetical redraw $x$ (used to compute the KL divergence) is reflected in the product $p_t(x)\,p_t(y)$. Hence the EKL divergence (ref) is a “two-sample” criterion: one draw updates the parameter, and an independent draw evaluates the fidelity to the true distribution. EKL improvements therefore measure the expected gain under repeated sampling. By incorporating the uncertainty in both draws, the EKL divergence provides a natural criterion for assessing updating rules. \begin{figure}[tb] \begin{subfigure}[b]{0.45\textwidth} \caption{$\lambda_{t} < \vartheta_{t\vert t} \color{black} < \vartheta_{t\vert t-1}$} \end{subfigure} \begin{subfigure}[b]{0.45\textwidth} \caption{$\lambda_{t} > \vartheta_{t\vert t-1} > \vartheta_{t\vert t}$} \end{subfigure} \caption{SD location model in Example (ref) with two hypothetical true densities $p_t$.} \end{figure} \begin{example} Consider the true distribution $Y_t \sim p_t = {\mathcal{N}}(\lambda_t,1)$ with unknown time-varying mean $\lambda_t$, and the correctly specified model density $f_{t\vert t-1} = {\mathcal{N}}(\vartheta_{t\vert t-1},1)$. With unit scaling $\mathcal{S}_{t-1}=1$ and a (typically small) scalar learning rate $A=\alpha>0$, the SD update $\vartheta_{t\vert t}(y_t) = \vartheta_{t\vert t-1} + \alpha (y_t - \vartheta_{t\vert t-1})$ simply shifts the mean towards the observation $y_t$. Figure (ref) shows $f_{t\vert t-1}$ in red, $f_{t\vert t}$ in blue, and $p_t$ in black. In panel (a), where $y_t$ is typical under $p_t$, the update improves fit; in panel (b), where $y_t$ is an outlier, it worsens fit. As the setting of panel (a) is more likely than that of panel (b), KL improvements may hold \emph{in expectation}. Here, the expectation averages out $y_t$ in $\vartheta_{t|t}(y_t)$ with respect to the true law $p_t$, while $x$ represents an independent redraw, also from $p_t$, used to assess distributional closeness in (ref). \end{example} \begin{definition}[EKL difference] For any $(\vartheta_{t\vert t-1},p_t) \in \Theta \times {\mathcal{P}}$, the \emph{EKL difference} for an updating rule $\phi$ is \begin{align*} \Delta^{\mathsf{EKL}}(\phi) \equiv \Delta^{\mathsf{EKL}}(\phi| \vartheta_{t\vert t-1},p_t) := \mathsf{EKL}(p_{t} \Vert f_{t\vert t}) - \mathsf{EKL}(p_{t} \Vert f_{t\vert t-1}). \end{align*} \sloppy We say that an update $\phi$ is \emph{EKL reducing w.r.t.\ the class ${\mathcal{P}}$} if it guarantees an EKL reduction, $ \Delta^{\mathsf{EKL}}(\phi \vert \vartheta_{t\vert t-1}, p_t) < 0 $, for all $\vartheta_{t\vert t-1} \in \Theta$ and $p_t \in {\mathcal{P}}$ such that $\mathbb{E}_{p_t}\!\left[ \Delta\phi (Y_t,\vartheta_{t\vert t-1}) \right]^\top \mathbb{E}_{p_t}\!\left[ {s}(X_t,\vartheta_{t\vert t-1}) \right] \not= 0$. \end{definition} The condition $\mathbb{E}_{p_t}\!\left[ \Delta\phi (Y_t,\vartheta_{t\vert t-1}) \right]^\top \mathbb{E}_{p_t}\!\left[ {s}(X_t,\vartheta_{t\vert t-1}) \right] \neq 0$ excludes distributions $p_t$ for which $\vartheta_{t\vert t-1}$ coincides with the pseudo-true parameter; i.e., a point at which EKL improvements are impossible. The relevance of this condition will become clearer after Theorem (ref). Definition (ref) involves the difference of two EKL terms. Note, however, that $f_{t\vert t-1}$ does not depend on $Y_t$ and hence $ \mathsf{EKL}(p_{t}\Vert f_{t\vert t-1}) =\mathsf{KL}(p_{t} \Vert f_{t\vert t-1}). $ In order to analyze the resulting difference, we impose the following assumptions. For convenience, the model-implied Hessian is denoted by \[ {H}(x,\vartheta) \;:=\; \frac{\partial^2}{\partial\vartheta \,\partial\vartheta^\top} \log f(x | \vartheta). \] Our assumptions below ensure that local curvature (in $\vartheta$) of the expected log-likelihood can be bounded; they also justify (the integral form of) exact multivariate mean-value expansions, used in our proofs. \begin{assumption} \begin{enumerate}[label=(\roman*), itemsep=0.2em, topsep=0.5em] • $\Theta$ is open and convex, and $\vartheta_{t\vert t-1} \in \Theta$ and $\vartheta_{t\vert t}(Y_t) \in \Theta$. • $\log f(x|\vartheta)$ is twice continuously differentiable in $\vartheta$ for all $\vartheta \in \Theta$ and $x \in \mathcal{Y}$. • $\mathbb{E}_{p_t} \big[{s}(X_t,\vartheta_{t\vert t-1}) \big]$ and $\mathbb{E}_{p_t} \left[ \Vert \Delta \phi(Y_t, \vartheta_{t\vert t-1}) \Vert^2\right]$ are finite for all $\vartheta_{t\vert t-1} \in \Theta$ and $p_t \in \mathcal{P}$. \end{enumerate} \end{assumption} \begin{assumptionp}{$\mathcal{HB}$} There exists a constant $c<\infty$ such that $\sup_{\vartheta \in \Theta} \big\Vert \mathbb{E}_{p_t}[{H} (X_t, \vartheta)] \big\Vert \leq c$ for all $p_t \in {\mathcal{P}}$. \end{assumptionp} Assumption (ref) imposes standard smoothness and moment conditions that are mild and comparable to, or weaker than, those in blasques2015information, Gorgi2023, and Creal2024GMM. Our main substantive requirement is the \emph{Hessian boundedness} in expectation in Assumption (ref), which can also be written as $-c I_k \preceq \mathbb{E}_{p_t}[{H} (X_t, \vartheta)] \preceq c I_k$, uniformly for all $\vartheta \in \Theta$. For some model classes (see Section (ref)), this condition can be shown to hold under mild moment conditions on the class ${\mathcal{P}}$. For this condition to hold independently of $p_t$ (i.e., for all $p_t \in {\mathcal{P}}_0$), a sufficient condition is $\sup_{\vartheta \in \Theta} \Vert {H} (x, \vartheta) \Vert < \infty$ for all $x\in\mathcal{Y}$. For our second main result, we will relax Assumption (ref) to a localized version (Assumption (ref)). However, we note that Assumption (ref) is \emph{already weaker} than typical conditions in the literature (e.g., Gorgi2023; Creal2024GMM), which additionally demand negative definiteness of the expected Hessian, i.e., $\mathbb{E}_{p_t}[{H}(X_t,\vartheta)] \prec O_k$ for all $\vartheta \in \Theta$. This typical condition differs from Assumption (ref) because it additionally imposes a strict zero upper bound; as a result, it is much harder to achieve and effectively rules out model densities that fail to be log-concave in $\vartheta$ (see Section (ref) for examples). To characterize EKL-improving updates, we introduce a localization in the state space via \emph{linearly downscaled} updates. Given an updating rule $\vartheta_{t\vert t}(y) = \phi(y,\vartheta_{t\vert t-1})$, we define its downscaled version for a (typically small) value of $\kappa>0$ as \begin{equation} \vartheta^\kappa_{t|t}(y) := (1-\kappa)\vartheta_{t\vert t-1} + \kappa \vartheta_{t\vert t}(y),\quad 0 < \kappa\leq 1. \end{equation} Equivalently, \begin{equation} \Delta \phi_\kappa(y,\vartheta_{t\vert t-1}) = \vartheta^\kappa_{t\vert t}(y) - \vartheta_{t\vert t-1} = \kappa \Delta \phi(y,\vartheta_{t\vert t-1}), \end{equation} \sloppy so adjustments are downscaled by $\kappa$. For SD updates this yields $ \Delta \phi_{\mathrm{SD},\kappa}(y,\vartheta_{t\vert t-1}) = \kappa A \mathcal{S}_{t-1} {s}(y,\vartheta_{t\vert t-1}), $ where $\kappa$ directly modulates $A \mathcal{S}_{t-1}$. For general updating rules, downscaling provides linear control of the step size. \begin{theorem} Consider a class ${\mathcal{P}}$ and let Assumptions (ref) and (ref) hold. Then, for each $\vartheta_{t\vert t-1}\in \Theta$ and $p_t \in {\mathcal{P}}$ such that $\mathbb{E}_{p_t}\!\left[\Delta \phi(Y_t,\vartheta_{t\vert t-1})\right]^\top \mathbb{E}_{p_t}\!\big[{s}(X_t,\vartheta_{t\vert t-1}) \big]\neq 0$, we have: \begin{gather*} \mathbb{E}_{p_t} \big[ \Delta\phi (Y_t,\vartheta_{t\vert t-1}) \big]^\top \, \mathbb{E}_{p_t} \big[ {s}(X_t,\vartheta_{t\vert t-1}) \big] > 0 \\ \ \iff \quad \text{ there exists } \bar{\kappa} > 0 \text{ such that for all } \kappa \in (0, \bar{\kappa}] : \ \Delta^{\mathsf{EKL}}(\phi_\kappa| \vartheta_{t\vert t-1}, p_t) < 0. \end{gather*} \end{theorem} The proof of Theorem (ref) applies the path-integral version of the exact (multivariate) mean-value theorem to the EKL criterion at $\vartheta_{t\vert t-1}$, yielding \begin{align} \Delta^{\mathsf{EKL}}(\phi_\kappa| \vartheta_{t\vert t-1}, p_t) = -\kappa \, \mathbb{E}_{p_t}\!\left[ \Delta \phi(Y_t, \vartheta_{t\vert t-1}) \right]^\top \mathbb{E}_{p_t}\!\left[ {s}(X_t, \vartheta_{t\vert t-1}) \right] + \mathcal{O}(\kappa^2). \end{align} Thus, for sufficiently small updates (i.e., $\kappa>0$ small), EKL reductions require the $\mathcal{O}(\kappa)$ term to have the correct sign. The admissible update size depends on $p_t$ and $\vartheta_{t\vert t-1}$ and is characterized by an upper bound $\bar{\kappa}\equiv\bar{\kappa}(p_t,\vartheta_{t\vert t-1})$ on the downscaling parameter $\kappa$. \sloppy The condition $\mathbb{E}_{p_t}\left[\Delta \phi(Y_t,\vartheta_{t\vert t-1})\right]^\top \mathbb{E}_{p_t} \big[{s}(X_t,\vartheta_{t\vert t-1}) \big] \not= 0$ in Theorem (ref) assures that the $\mathcal{O}(\kappa)$ term in (ref) does not vanish. If the inner product of these vectors is zero (e.g., because one of them is zero or they are orthogonal), then the $\mathcal{O}(\kappa^2)$ term dominates; however, the $\mathcal{O}(\kappa^2)$ term may take either sign, such that no general improvement guarantees are then possible. We will use a similar approach in Section (ref) to analyze the criteria of Gorgi2023,Creal2024GMM,blasques2015information. For a given $\vartheta_{t\vert t-1} \in \Theta$, Theorem (ref) establishes EKL improvements for sufficiently small $\kappa > 0$ if and only if the expected update direction and the expected score are aligned, i.e., \begin{equation} \mathbb{E}_{p_t}\!\left[ \Delta\phi (Y_t,\vartheta_{t\vert t-1}) \right]^\top \mathbb{E}_{p_t}\!\left[ {s}(X_t,\vartheta_{t\vert t-1}) \right] > 0. \end{equation} \sloppy Updating rules that satisfy this condition for all $\vartheta_{t\vert t-1} \in \Theta$ and all $p_t \in {\mathcal{P}}$ with the nonzero inner product condition $\mathbb{E}_{p_t}\!\left[\Delta \phi(Y_t,\vartheta_{t\vert t-1})\right]^\top \mathbb{E}_{p_t}\!\big[{s}(X_t,\vartheta_{t\vert t-1})\big] \neq 0$ (recall Definition (ref)) are called \emph{score equivalent in expectations} (SEE) with respect to the class ${\mathcal{P}}$. Because the true density $p_t$ is unknown, we want the class $\mathcal P$ to be as broad as possible. The nonzero inner product condition typically excludes any $p_t$ for which $\vartheta_{t\vert t-1}$ is a stationary point, for then $\mathbb{E}_{p_t}\!\big[{s}(X_t,\vartheta_{t\vert t-1})\big]=0$. At such points, the EKL difference is zero to first order in $\kappa$, and the second-order term in (ref), whose sign is generally unknown, dominates. As our EKL criterion is formulated in terms of a double expectation, it is natural that our version (ref) of score equivalence also contains two expectations. For SD updates, we have $\mathbb{E}_{p_t}[ \Delta \phi_{\mathrm{SD}} (Y_t,\vartheta_{t\vert t-1}) ] = A \mathcal{S}_{t-1} \mathbb{E}_{p_t}[ {s}(Y_t,\vartheta_{t\vert t-1}) ]$, such that (ref) becomes \begin{align} \mathbb{E}_{p_t}[{s}(Y_t,\vartheta_{t\vert t-1})]^\top\, (A \mathcal{S}_{t-1})\, \mathbb{E}_{p_t}[{s}(X_t,\vartheta_{t\vert t-1})] \;>\; 0. \end{align} For this quadratic form to be positive, we merely require (i) a nonzero expected score (as discussed above), and (ii) $A \mathcal{S}_{t-1}\succ O_k$. Positive definiteness of $A \mathcal{S}_{t-1}$ is consistent with the discussion below Definition (ref).\footnote{Positivity in (ref) can also hold for non-symmetric $A\mathcal{S}_{t-1}$ if (and only if) its \emph{symmetric} part $(A\mathcal{S}_{t-1})+(A\mathcal{S}_{t-1})^\top$ is positive definite. Although non-symmetric versions of $A\mathcal{S}_{t-1}$ have been used (e.g., gasperoni_score-driven_2023,d2024modeling), we do not pursue them since they conflict with the SD philosophy that updates move in the direction of the score. Any non-symmetric matrix is an average of its symmetric and anti-symmetric parts: the anti-symmetric part $(A\mathcal{S}_{t-1})-(A\mathcal{S}_{t-1})^\top$, if non-zero, adds an update component $[(A\mathcal{S}_{t-1})-(A\mathcal{S}_{t-1})^\top]{s}(y_t,\vartheta_{t\vert t-1})$ that is \emph{perpendicular} to the score as ${s}(y_t,\vartheta_{t\vert t-1})^\top[(A\mathcal{S}_{t-1})-(A\mathcal{S}_{t-1})^\top]{s}(y_t,\vartheta_{t\vert t-1})=0$. Symmetric $A\mathcal{S}_{t-1}$ avoids this perpendicular component.} Theorem (ref) thus implies the following result for SD updates.\footnote{The informal arguments in blasques2021finite motivate a quantity we formalize as the EKL measure, but their simulation-based approach does not provide a theoretical characterization.} \begin{corollary} Consider a class ${\mathcal{P}}$ and let Assumptions (ref) and (ref) hold. Then, for any $\vartheta_{t\vert t-1} \in \Theta$ and $p_t \in {\mathcal{P}}$ with $ \mathbb{E}_{p_t} \big[{s}(X_t,\vartheta_{t\vert t-1}) \big]\neq 0$, there exists a $\bar{\kappa} >0$ such that $\Delta^{\mathsf{EKL}}(\phi_{\mathrm{SD}, \kappa} | \vartheta_{t\vert t-1}, p_t) < 0$ for all $\kappa \in (0, \bar{\kappa}]$, where $\Delta \phi_{\mathrm{SD}, \kappa} (y_t,\vartheta_{t\vert t-1}) = \kappa A \mathcal{S}_{t-1} {s}(y_t,\vartheta_{t\vert t-1})$. \end{corollary} Corollary (ref) shows that SD updates with generic learning-rate and scaling matrices are EKL reducing w.r.t.\ general classes ${\mathcal{P}}$ as long as $A\mathcal{S}_{t-1} \succ O_k$. This is in contrast with other methods that tend to require $A \mathcal{S}_{t-1}$ to be a scalar multiple of the identity (see Section (ref)). While Corollary (ref) invokes Assumption (ref) for simplicity, its result is valid under the weaker \emph{one-sided} condition, $-c I_k \preceq \mathbb{E}_{p_t} \big[ {H}(X_t,\vartheta) \big]$, uniformly for all $\vartheta \in \Theta$, for some $c<\infty$. The SEE condition (ref) offers an easily verifiable criterion for whether general updating rules are EKL reducing. For instance, an update that is usually score-driven but, with some probability less than one, sets $\vartheta_{t\vert t}=\vartheta_{t\vert t-1}$ remains SEE. Similarly, clipped SD updates (Proposition (ref) below) are SEE if the clipping constant is sufficiently large. The Kalman filter, the most widely used algorithm in state-space modeling durbin2012time, can be expressed as an SD update; hence, it is SEE and its downscaled version is EKL reducing (see Appendix (ref)). Implicit score-driven (ISD, lange2022robust) updates with sufficiently small learning rates are similarly EKL reducing (see Appendix (ref)). By contrast, quasi score-driven (QSD, blasques2023quasi) models use two densities: one to compute the score driving the dynamics and another to evaluate the log-likelihood (the model’s “fit”). Consequently, the SEE condition involves two expectations of distinct scores whose inner product need not be positive. Hence QSD updates with small learning rates are not always EKL reducing (see Appendix (ref)). The SEE condition (ref) can be contrasted with the (univariate) \emph{almost sure} score-equivalence condition in blasques2015information, i.e., $\operatorname{sign}(\Delta \phi(y_t, \vartheta_{t\vert t-1})) = \operatorname{sign}({s}(y_t, \vartheta_{t\vert t-1}))$ or its multivariate generalization, $\Delta \phi(y_t, \vartheta_{t\vert t-1})^\top {s}(y_t, \vartheta_{t\vert t-1}) > 0$. Interestingly, this pointwise condition neither implies nor is implied by the SEE condition (ref), which involves two expectations. \subsection{Extension to locally bounded Hessians in expectation} Assumption (ref) is violated by several widely used models; for example, a Gaussian stochastic volatility model with logarithmic variance (see Section (ref)). We therefore replace it with the weaker Assumption (ref), which requires the Hessian $H(X_t,\vartheta)$ to be merely \emph{locally bounded} ($\mathcal{HLB}$) in expectation. \begin{assumptionp}{$\mathcal{HLB}$} For any $\vartheta_{t\vert t-1} \in \Theta$, there exists a compact ball $\widetilde{\Theta}_r \equiv \widetilde{\Theta}_r(\vartheta_{t\vert t-1}) \subset \Theta$ around $\vartheta_{t\vert t-1}$ with radius $r >0$ such that $\underset{\vartheta \in \widetilde{\Theta}_r(\vartheta_{t\vert t-1})}{\sup} \big\Vert \mathbb{E}_{p_t} \big[ {H}(X_t,\vartheta) \big] \big\Vert < \infty$ for all $p_t \in {\mathcal{P}}$. \end{assumptionp} Assumption (ref) is mild and is satisfied by all examples in Section (ref): either for any $p_t \in {\mathcal{P}}_0$ or based on mild moment conditions on the class ${\mathcal{P}}$. For example, for volatility and dependence models based on the Gaussian distribution, the existence of a finite second moment of the data is sufficient for Assumption (ref). Under Assumption (ref), we obtain a characterization paralleling Theorem (ref), now restricted to \emph{bounded} updates. \begin{theorem} Consider a class ${\mathcal{P}}$ and let Assumptions (ref) and (ref) hold. Then, for any \emph{bounded update} $\phi$ with $\Vert \Delta\phi (Y_t,\vartheta_{t\vert t-1}) \Vert \le d(\vartheta_{t\vert t-1}) < \infty$ and any $\vartheta_{t\vert t-1} \in \Theta$ and $p_t \in {\mathcal{P}}$ such that $\mathbb{E}_{p_t}\!\left[\Delta \phi(Y_t,\vartheta_{t\vert t-1})\right]^\top \mathbb{E}_{p_t}\! \big[{s}(X_t,\vartheta_{t\vert t-1}) \big] \not= 0$, the following equivalence holds: \begin{gather*} \mathbb{E}_{p_t} \big[ \Delta\phi (Y_t,\vartheta_{t\vert t-1}) \big]^\top \, \mathbb{E}_{p_t} \big[ {s}(X_t,\vartheta_{t\vert t-1}) \big] > 0 \\ \ \iff \quad \text{ there exists } \bar{\kappa} > 0 \text{ such that for all } \kappa \in (0, \bar{\kappa}] : \ \Delta^{\mathsf{EKL}}(\phi_\kappa| \vartheta_{t\vert t-1}, p_t) < 0. \end{gather*} \end{theorem} The focus on bounded updates is consistent with results in optimization and machine learning boyd2004convex,nesterov2018lectures,mai2021stability, which emphasize that for optimization problems with merely locally (but not globally) bounded Hessians, step sizes must be capped. Applicability of Theorem (ref) can be guaranteed by a \emph{clipped} SD (CSD) rule, which restricts the update to some distance $d>0$ around $\vartheta_{t\vert t-1}$, \begin{align} \Delta \phi_{\mathrm{CSD}}^d(Y_t,\vartheta_{t\vert t-1}) := \min \left\{1, \frac{d}{\Vert \Delta\phi_{\mathrm{SD}}(Y_t,\vartheta_{t\vert t-1})\Vert} \right\} \Delta\phi_{\mathrm{SD}}(Y_t,\vartheta_{t\vert t-1}), \end{align} with $1/0:=\infty$, and where the constant $d$ may depend on $\vartheta_{t\vert t-1}$. Gradient clipping, standard in stochastic optimization and deep learning goodfellow2016deep, improves robustness by capping step sizes. Because clipping is triggered by large scores, it is not obvious that the SEE condition (ref) persists. The next proposition shows that, under general conditions on $p_t$, there exists a clipping constant $d>0$ such that CSD updates remain SEE. \begin{proposition} Fix $A\mathcal{S}_{t-1} \succ O_k$ and let Assumption (ref) hold. Moreover, let (i) $\mu_{p_t} := \mathbb{E}_{p_t} \big[{s}(Y_t,\vartheta_{t\vert t-1}) \big] \neq 0$ and (ii) $\textnormal{tr}( \Sigma_{p_t}) < \infty$ where $\Sigma_{p_t} := \mathbb{E}_{p_t} \left[{s}(Y_t,\vartheta_{t\vert t-1}){s}(Y_t,\vartheta_{t\vert t-1})^\top\right] - \mu_{p_t} \mu_{p_t}^\top$. If the constant $d>0$ is sufficiently large, such that \begin{align} \mathbb{P}_{p_t}\!\left( \big\Vert A \mathcal{S}_{t-1} {s}(Y_t,\vartheta_{t\vert t-1}) \big\Vert > d \right) \;<\; \frac{\|\mu_{p_t}\|_{A\mathcal{S}_{t-1}}^2}{\|\mu_{p_t}\|_{A\mathcal{S}_{t-1}}^2+\textnormal{tr} \big( A\mathcal{S}_{t-1}\,\Sigma_{p_t}\big)}, \end{align} where $\|\mu_{p_t}\|_{A\mathcal{S}_{t-1}}^2:= \mu_{p_t}^\top A\mathcal{S}_{t-1} \mu_{p_t}$, then the CSD update (ref) satisfies the SEE condition $\mathbb{E}_{p_t} \big[ \Delta \phi_{\mathrm{CSD}}^d (Y_t,\vartheta_{t\vert t-1}) \big]^\top\, \mathbb{E}_{p_t} \big[ {s}(X_t,\vartheta_{t\vert t-1}) \big] \; >\; 0$. \end{proposition} Inequality (ref) can be satisfied by choosing the (finite) clipping constant $d$ large enough, with the magnitude dependent on $p_t$. For example, if the prediction is inaccurate so that $\|\mu_{p_t}\|_{A {\mathcal{S}_{t-1}}}^2$ is large relative to $\mathrm{tr}\!\left(A {\mathcal{S}_{t-1}}\Sigma_{p_t}\right)$, condition (ref) is easily satisfied, as the right-hand side approaches one. Conversely, if the prediction is accurate and $\|\mu_{p_t}\|_{A {\mathcal{S}_{t-1}}}^2$ is close to zero, a larger $d$ may be required to keep clipping sufficiently rare. Combining Proposition (ref) with Theorem (ref), we find that for any prediction $\vartheta_{{t\vert t-1}}$ and any true density $p_t$, there exist thresholds $ \overline{\kappa} \equiv \overline{\kappa}(p_t,\vartheta_{{t\vert t-1}})>0$ and $\underline{d} \equiv \underline{d}(p_t,\vartheta_{{t\vert t-1}})<\infty$, such that for all $\kappa\in(0,\overline{\kappa}]$ and $d \in [\underline{d}, \infty)$, the resulting (C)SD updates improve our EKL criterion. In practice, $\overline{\kappa}$ and $\underline{d}$ are unknown and must be estimated from data (e.g., using empirical gradient moments). Theorem (ref) in the next section aims to offer practical guidance on how “large” SD updates can be. \begin{remark} As the proofs reveal, Theorems (ref) and (ref) extend to general distributions $P_t \ll \mu$ on a measure space $({\mathcal{Y}},\mathcal{G},\mu)$. Writing $p_t = \tfrac{\mathrm{d} P_t}{\mathrm{d} \mu}$ as the Radon–Nikodym derivative, all expectations can be interpreted as Lebesgue integrals with respect to $\mu$, which need not be the Lebesgue measure. For discrete distributions, $\mu$ is the counting measure (i.e., integrals reduce to sums), and mixed continuous–discrete cases are also covered. In particular, $p_t$ may be discrete even when $f$ is defined as a density with respect to the Lebesgue measure. \end{remark} \begin{remark} We have established guarantees for $\mathsf{EKL}(p_t \Vert f_{{t\vert t}})$ involving the current density $p_t$ based on an observation $Y_t \sim p_t$. We could also seek guarantees for $\mathsf{EKL}(p_{t+1} \Vert f_{{t\vert t}})$, involving the future density $p_{t+1}$, even though $f_{{t\vert t}}$ still depends only on $Y_t \sim p_t$. Adapting the proof of Theorem (ref) yields the following (necessary and sufficient) condition: \begin{align} \mathbb{E}_{p_t}[{s}(Y_t,\vartheta_{t\vert t-1})]^\top\, (A \mathcal{S}_{t-1})\, \mathbb{E}_{p_{t+1}}[{s}(X_t,\vartheta_{t\vert t-1})] \;>\; 0. \end{align} Because the two expectations are now taken with respect to different laws, this is not a quadratic form; however, (ref) can hold if \(p_t\) and \(p_{t+1}\) are sufficiently close. For SD models targeting future densities, condition (ref) characterizes when this is feasible. \end{remark} \section{Upper bounds for the learning-rate matrix} So far we asked whether there exists $\bar\kappa>0$ such that any downscaled update $\phi_\kappa$ with $\kappa\in(0,\bar\kappa]$ improves the EKL criterion. Here, for a given true density $p_t$, we instead quantify how large the SD matrix combination $\mathcal{A}_{t-1}:=A\mathcal{S}_{t-1}$ can be while still guaranteeing EKL improvement. If $\mathcal{A}_{t-1}$ is positive definite and sufficiently small, Corollary (ref) applies. The next result provides explicit upper bounds on the elements (or eigenvalues) of $\mathcal{A}_{t-1}$. When $\mathcal{S}_{t-1}=I_k$, these are bounds on the static learning-rate matrix $A$; more generally, they bound the combined matrix $\mathcal{A}_{t-1}$ multiplying the score, beyond the specific case $\mathcal{A}_{t-1}=A\mathcal{S}_{t-1}$. \begin{theorem} Consider a class ${\mathcal{P}}$ and let Assumptions (ref) and (ref) hold. For a given $\vartheta_{{t\vert t-1}} \in \Theta$, let $p_t \in {\mathcal{P}}$ be such that $\mu_{p_t} := \mathbb{E}_{p_t} \big[{s}(Y_t,\vartheta_{t\vert t-1}) \big]\neq 0$, with elements $\mu_1,\dots,\mu_k$. Moreover, let $\Sigma_{p_t} := \mathbb{E}_{p_t} \left[{s}(Y_t,\vartheta_{t\vert t-1}){s}(Y_t,\vartheta_{t\vert t-1})^\top\right] - \mu_{p_t} \mu_{p_t}^\top$ be finite with diagonal elements $\sigma_1^2, \dots,\sigma_k^2$. \sloppy Then, the condition $\Delta^{\mathsf{EKL}}(\phi_{\mathrm{SD}} | \vartheta_{t\vert t-1}, p_t) < 0$ holds either \color{black} \begin{enumerate}[label=(\Alph*), itemsep=0.2em, topsep=0.5em] • for SD updates with $\mathcal{A}_{t-1} = \alpha_{t-1} I_k \succ O_k$ if \begin{align} \alpha_{t-1} < \frac{2}{c} \, \frac{\|\mu_{p_t}\|^2 }{\|\mu_{p_t}\|^2+ \operatorname{tr}(\Sigma_{p_t})}; \qquad \text{or} \end{align} • for SD updates with a diagonal $\mathcal{A}_{t-1} = \operatorname{diag}(\alpha_{1, t-1}, \dots, \alpha_{k, t-1}) \succ O_k$ and non-zero elements $\mu_1,\dots,\mu_k$ if \begin{align} \alpha_{i,t-1} \le \frac{2}{c} \, \frac{\mu_{i}^2}{\mu_{i}^2+\sigma_{i}^2}, \qquad i=1,\ldots,k, \end{align} and the inequality (ref) is strict for at least one $i = 1,\dots, k$; or • for SD updates with a general $\mathcal{A}_{t-1} \succ O_k$ if \begin{align} \frac{\lambda_{\max}(\mathcal{A}_{t-1})^2}{\lambda_{\min}(\mathcal{A}_{t-1})} < \frac{2}{c} \, \frac{\Vert \mu_{p_t}\Vert^2}{\Vert \mu_{p_t} \Vert^2 + \operatorname{tr}(\Sigma_{p_t})}. \end{align} \end{enumerate} \end{theorem} The three formulations (A)–(C) of Theorem (ref) share the factor $2/c$ and the same “signal-to-noise” ratio based on its first two moments of the score, analyzed in detail below for the univariate case. In multivariate settings, (A) can be overly conservative, while (B) better adapts to heterogeneity in the cross-section. Its required score moments in each direction can be estimated from the data. E.g., if $p_t$ conditions on ${\mathcal{F}}_{t-1} = \sigma(Y_s: s \le t-1)$, one can use exponentially weighted moving averages, as is standard in the optimization literature (e.g., kingma2014adam). Part (C) allows a generic $\mathcal{A}_{t-1}\succ O_k$, with EKL improvements guaranteed if $\lambda_{\max}(\mathcal{A}_{t-1})\,\mathrm{cond}(\mathcal{A}_{t-1})$ is upper bounded, where $\mathrm{cond}(\cdot)=\lambda_{\max}(\cdot)/\lambda_{\min}(\cdot)$ is the condition number. Hence EKL improvements can be guaranteed for general learning-rate matrices $\mathcal{A}_{t-1}\succ O_k$ if we control both the largest eigenvalue and the condition number. Parts (B) and (C) are specific to our framework: they guide admissible choices of the combined matrix $\mathcal{A}_{t-1}=A\mathcal{S}_{t-1}$, allowing both diagonal $\mathcal{A}_{t-1}$ with heterogeneous entries and general positive definite $\mathcal{A}_{t-1}$. In contrast, the methods of Gorgi2023 (Sections (ref)--(ref)) and Creal2024GMM (Section (ref)) yield improvements only when $\mathcal{A}_{t-1}$ is proportional to the identity. To facilitate the comparison with Gorgi2023, we take $k=1$ and $\mathcal{S}_{t-1}=1$. Theorem (ref) guarantees EKL improvements for any SD update with learning rate $\alpha>0$ if \begin{align} \alpha < \bar{\alpha}_\mathsf{EKL} := \frac{2}{c} \;\frac{\left(\mathbb{E}_{p_t}[{s}(Y_t,\vartheta_{t\vert t-1})]\right)^2}{\left(\mathbb{E}_{p_t}[{s}(Y_t,\vartheta_{t\vert t-1})]\right)^2 + \mathbb{V}_{p_t}\!\left({s}(Y_t,\vartheta_{t\vert t-1})\right)}. \end{align} The bound splits into a static term $2/c$ and a dynamic “signal-to-noise” ratio formed from moments of the score. This ratio is always at most one, so $\bar{\alpha}_{\mathsf{EKL}}\le 2/c$. Both moments depend on the prediction $\vartheta_{{t\vert t-1}}$. If $\vartheta_{{t\vert t-1}}$ is close to the pseudo-true parameter \begin{align} \vartheta_t^\ast := \arg \max_{\vartheta \in \Theta} \mathbb{E}_{p_t}\big[\log f(Y_t | \vartheta)\big], \end{align} and, for simplicity, we assume $\vartheta_t^\ast \in \Theta$, then standard regularity conditions imply $\mathbb{E}_{p_t}\!\left[{s}(Y_t, \vartheta_t^\ast)\right] = 0$, so that $\bar{\alpha}_{\mathsf{EKL}}$ approaches zero as $\vartheta_{{t\vert t-1}}\to \vartheta_t^\ast$. In other words, the more accurate the prediction, the smaller the admissible learning rate. Gorgi2023 derive a simpler (and larger) bound in the univariate case. If $-c \le \mathbb{E}_{p_t} \big[ {H}(X_t, \vartheta) \big] < 0$ for all $\vartheta \in \Theta$, then any $\alpha < \bar{\alpha}_\mathsf{CEV} := 2/c$ ensures improvement under their conditional expected variation (CEV) criterion, \allowdisplaybreaks \begin{alignat}{2} \begin{aligned} \big\vert \vartheta_t^\ast - \mathbb{E}_{p_t}[\vartheta_{t\vert t}(Y_t)] \big\vert &< \big\vert \vartheta_t^\ast - \vartheta_{t\vert t-1} \big\vert \qquad &&\text{if } \vartheta_{t\vert t-1} \not= \vartheta_t^\ast, \\ \mathbb{E}_{p_t}[\vartheta_{t\vert t}(Y_t)] &= \vartheta_t^\ast \qquad &&\text{if } \vartheta_{t\vert t-1} = \vartheta_t^\ast. \end{aligned} \end{alignat} The CEV criterion depends solely on the distance of $\mathbb{E}_{p_t}[\vartheta_{t\vert t}(Y_t)]$ from $\vartheta_t^\ast$, disregarding further distributional differences between $f_{{t\vert t}}$ and $p_t$. We note that $\bar{\alpha}_\mathsf{EKL}\leq 2/c= \bar{\alpha}_\mathsf{CEV}$. Indeed, the EKL criterion demands decreasing learning rates as predictions become more accurate, but $\bar{\alpha}_\mathsf{CEV}$ remains fixed at $2/c$. Hence CEV may favor sizable learning rates even when the gradient is largely dominated by noise. The EKL criterion provides the sharper and more intuitive requirement that learning rates should shrink as predictions improve. \section{Comparison to related performance criteria} This section contrasts our EKL-based characterization of SD updates with four recent alternative performance measures. Sections (ref)--(ref) show that these approaches require stronger conditions (e.g., concave model log-densities) yet deliver weaker results: (i) their equivalence conditions are not directly useful for model construction, and (ii) non-SD updates may improve these measures, while SD updates are guaranteed to do so only under more restrictive assumptions. The performance measure discussed in Section (ref) is \emph{improper} in the statistical sense gneiting2007strictly: no conclusions on its basis can be drawn. \subsection{CEV of Gorgi2023} Gorgi2023 evaluate SD models using the univariate CEV criterion (ref). Our multivariate generalization of their CEV criterion can be written as \begin{align} \mathsf{CEV}(p_t \Vert f_{{t\vert t}}) := \Big( \mathbb{E}_{p_t} \big[ \vartheta_{t\vert t} (Y_t) \big] - \vartheta_t^\ast \Big)^\top \Omega_{t-1} \Big( \mathbb{E}_{p_t} \big[ \vartheta_{t\vert t} (Y_t) \big] - \vartheta_t^\ast \Big), \end{align} where $\Omega_{t-1}\succ O_k$ is an $\mathcal{F}_{t-1}$-measurable weighting matrix. In the univariate case, Gorgi2023 show that SD updates with sufficiently small learning rates reduce CEV losses, but leave open the question whether this property is unique to SD updates. Proposition (ref) below extends their analysis to the multivariate case, characterizing all CEV-improving updates by examining the CEV difference, defined as \begin{align*} \Delta^{\mathsf{CEV}}(\phi) \equiv \Delta^{\mathsf{CEV}}(\phi| \vartheta_{t\vert t-1},p_t) := \mathsf{CEV}(p_{t} \Vert f_{t\vert t}) - \mathsf{CEV}(p_{t} \Vert f_{t\vert t-1}). \end{align*} An update $\phi$ is \emph{CEV reducing w.r.t.\ ${\mathcal{P}}$} if $ \Delta^{\mathsf{CEV}}(\phi \vert \vartheta_{t\vert t-1}, p_t) < 0$ for all $\vartheta_{t\vert t-1} \in \Theta, p_t \in {\mathcal{P}}$ such that $\mathbb{E}_{p_t} \big[ \Delta\phi (Y_t,\vartheta_{t\vert t-1}) \big]^\top \Omega_{t-1} \big(\vartheta_t^\ast - \vartheta_{t\vert t-1} \big) \not=0$. This mirrors the exclusion in Definition (ref) and prohibits $\vartheta_{t\vert t-1}$ from being equal to $\vartheta_t^\ast$. The class ${\mathcal{P}}$ may differ from that in Theorem (ref). \begin{proposition} Consider a class ${\mathcal{P}}$ and let Assumption (ref) hold. Then, for $\phi_\kappa(y, \vartheta_{t\vert t-1})$ in (ref), the following equivalence holds for all $p_t \in {\mathcal{P}}$, $\vartheta_{t\vert t-1} \in \Theta$: \color{black} \begin{gather} \mathbb{E}_{p_t} \big[ \Delta\phi (Y_t,\vartheta_{t\vert t-1}) \big]^\top \Omega_{t-1} \big(\vartheta_t^\ast - \vartheta_{t\vert t-1} \big) > 0 \\ \ \iff \quad \text{ there exists } \bar{\kappa} > 0 \text{ such that for all } \kappa \in (0, \bar{\kappa}] : \ \Delta^{\mathsf{CEV}}(\phi_\kappa| \vartheta_{t\vert t-1}, p_t) < 0.\notag \end{gather} \end{proposition} Proposition (ref) shows that an updating rule $\phi$ is CEV reducing if and only if (ref) holds: on average, $\phi$ must move the parameter toward the pseudo-truth $\vartheta_t^\ast$, with admissible directions governed by $\Omega_{t-1}\succO_k$. Since $\vartheta_t^\ast$ is unknown, however, this condition is not directly usable for model construction. Moreover, it is unclear whether (ref) is satisfied only by SD updates. Next, we derive minimal conditions under which SD updates $\Delta \phi_{\mathrm{SD}}(Y_t,\vartheta_{t\vert t-1})=A \mathcal{S}_{t-1} {s}(Y_t,\vartheta_{t\vert t-1})$ guarantee CEV improvements. Using $\mathbb{E}_{p_t}\!\left[{s}(Y_t,\vartheta_t^\ast)\right]=0$, which holds under standard regularity conditions (see Gorgi2023), condition (ref) for CEV reductions becomes \allowdisplaybreaks \begin{align} &0<\mathbb{E}_{p_t} \left[ \Delta \phi_{\mathrm{SD}}(Y_t,\vartheta_{t\vert t-1})\right]^\top \Omega_{t-1} \big( \vartheta_t^\ast-\vartheta_{t\vert t-1} \big) \nonumber \\ &= \mathbb{E}_{p_t} \left[ {s}(Y_t,\vartheta_{t\vert t-1})^\top - {s}(Y_t,\vartheta^\ast_t)^\top \right] (A \mathcal{S}_{t-1}) \Omega_{t-1} \big( \vartheta_t^\ast-\vartheta_{t\vert t-1} \big) \nonumber \\ &=\big( \vartheta_t^\ast-\vartheta_{t\vert t-1} \big)^\top \left( \mathbb{E}_{p_t} \left[ - \int_0^1 {H} \big( Y_t, \vartheta_t^u \big) \mathrm{d}u \right] (A\mathcal{S}_{t-1}) \Omega_{t-1} \right) \big( \vartheta_t^\ast-\vartheta_{t\vert t-1} \big), \end{align} where we used the integral form of the mean-value theorem with ${\vartheta}_t^u:= u \vartheta_{t\vert t-1} + (1-u)\vartheta_t^\ast$ for $u\in[0,1]$. To ensure the positivity of the quadratic form (ref) for all $\vartheta_t^\ast-\vartheta_{t\vert t-1} \neq 0$, Gorgi2023 impose that the central term (in large parentheses) is positive. In their setting, this term is a scalar; since $A \mathcal{S}_{t-1} \Omega_{t-1} > 0$, it then suffices to assume $\mathbb{E}_{p_t}\left[-{H}\big(Y_t,\vartheta\big)\right] > 0$ for all $\vartheta$. In the multidimensional case, as considered here, the natural analogue is to require the \emph{symmetric} part of central matrix $ \mathbb{E}_{p_t}\!\left[-{H}\big(Y_t,\vartheta\big)\right]A \mathcal{S}_{t-1} \Omega_{t-1} $ to be positive definite: \begin{align} \mathbb{E}_{p_t}[-{H}(Y_t,\vartheta)] (A \mathcal{S}_{t-1}) \Omega_{t-1} + \Omega_{t-1} (A \mathcal{S}_{t-1}) \mathbb{E}_{p_t}[-{H}(Y_t,\vartheta)] \;\succ\; O_k, \quad \forall \vartheta \in \Theta. \end{align} However, ensuring that condition (ref) holds is difficult, as the symmetrized sum need not be positive definite even if all individual matrices are nicholson1979eigenvalue. The following assumption provides a sufficient condition that guarantees (ref). \begin{assumptionp}{$\mathcal{HN}$} (i) $\mathbb{E}_{p_t} \big[ {H} ( Y_t, \vartheta ) \big] \prec O_k$ for all $\vartheta\in \Theta$, $p_t \in {\mathcal{P}}$; and (ii) $A \mathcal{S}_{t-1} \Omega_{t-1} = d_{t-1} I_k$ for some $\mathcal{F}_{t-1}$-measurable $d_{t-1} > 0$. \end{assumptionp} Part (ii) of Assumption (ref) limits the generality of $A\mathcal{S}_{t-1}$ in the multivariate setting: in the canonical case $\Omega_{t-1}=I_k$, it effectively enforces a scalar learning rate and scaling function. Part (i) is also hard to ensure, even in one dimension, unless ${H}(y,\vartheta)$ is negative definite for all $y\in\mathcal{Y}$ and $\vartheta\in\Theta$, i.e., the model density is strictly log-concave. By contrast, the EKL-specific Assumption (ref) only requires boundedness of the expected Hessian, which often follows from mild moment conditions on $p_t$; see Section (ref). Assumption (ref) (i) reduces to $\mathbb{E}_{p_t}[{H}(Y_t,\vartheta)]<0$ in the scalar case. Relative to Gorgi2023, this is slightly more general because it does not impose a lower bound. By dropping the lower bound, we already cover the localized theory of Gorgi2023. An analogue of (ref) for general scaling matrices is given in Appendix (ref). Gorgi2023 relax the zero upper bound in Assumption (ref) (i) further, but only for small deviations $\|\vartheta_{t\vert t-1}-\vartheta_t^\ast\|$, which cannot be ensured in practice. In sum, (ref) and (ref) show that CEV reductions are guaranteed under Assumption (ref) but are otherwise hard to establish. Compared with Assumption (ref), which only requires a bounded expected Hessian, Assumption (ref) imposes a zero upper bound (negative definiteness of the expected Hessian), excluding, for example, harvey2014filtering's (harvey2014filtering) Student's $t$ location model (Section (ref)) and a bivariate Gaussian location-scale model (Section (ref)). The difference stems from the proof methods: for the EKL criterion, the Hessian appears in the $\mathcal{O}(\kappa^2)$ term of Theorem (ref) and only needs to be bounded, whereas for the CEV criterion it enters via (ref) in the $\mathcal{O}(\kappa)$ term of Proposition (ref), requiring the stronger sign restriction. Thus, we conclude that while non-SD rules may yield CEV reductions, generic SD updates with $A\mathcal{S}_{t-1} \succ O_k$ are not necessarily guaranteed to consistently do so. \subsection{MSE of Gorgi2023} Gorgi2023 study updates that reduce the mean squared error (MSE) relative to the pseudo-true parameter $\vartheta_t^\ast$ in (ref). A multivariate extension of their criterion is \begin{align} \mathsf{MSE}(p_t \Vert f_{{t\vert t}}) := \mathbb{E}_{Y_t \sim p_t} \!\left[ \big(\vartheta_{t\vert t}(Y_t) - \vartheta_t^\ast\big)^\top \Omega_{t-1} \big(\vartheta_{t\vert t}(Y_t) - \vartheta_t^\ast\big) \right], \end{align} where, as before, $\Omega_{t-1}\succ O_k$ is an $\mathcal{F}_{t-1}$-measurable weighting matrix. The dependence of $\mathsf{MSE}(p_t \Vert f_{{t\vert t}})$ on $\vartheta_{t\vert t}(Y_t)$ is captured through $f_{t\vert t}$ as $f_{t\vert t}(\cdot) = f(\cdot|\vartheta_{t\vert t}(Y_t))$. We define \begin{align*} \Delta^{\mathsf{MSE}}(\phi) \equiv \Delta^{\mathsf{MSE}}(\phi| \vartheta_{t\vert t-1},p_t) := \mathsf{MSE}(p_{t} \Vert f_{t\vert t}) - \mathsf{MSE}(p_{t} \Vert f_{t\vert t-1}), \end{align*} and an update $\phi$ is said to be \emph{MSE reducing w.r.t.\ ${\mathcal{P}}$} if $ \Delta^{\mathsf{MSE}}(\phi \vert \vartheta_{t\vert t-1}, p_t) < 0$, for all $\vartheta_{t\vert t-1} \in \Theta$ and $p_t \in {\mathcal{P}}$ such that $\mathbb{E}_{p_t} \big[ \Delta\phi (Y_t,\vartheta_{t\vert t-1}) \big]^\top \Omega_{t-1} \big(\vartheta_t^\ast - \vartheta_{t\vert t-1} \big) \not=0$, where the last condition prohibits $\vartheta_{t\vert t-1}$ from being identical to $\vartheta_t^\ast$ (similar to the condition in Definition (ref)). Gorgi2023 show that SD updates with sufficiently small learning rates reduce MSE losses, but leave open whether this property is unique to SD updates---a question addressed by our next result, which characterizes all MSE-reducing updates. \begin{proposition} Consider a class ${\mathcal{P}}$ and let Assumption (ref) hold. Then, for $\phi_\kappa(y, \vartheta_{t\vert t-1})$ in (ref), the following equivalence holds for all $p_t \in {\mathcal{P}}$, $\vartheta_{t\vert t-1} \in \Theta$: \color{black} \begin{gather} \mathbb{E}_{p_t} \big[ \Delta\phi (Y_t,\vartheta_{t\vert t-1}) \big]^\top \Omega_{t-1} \big(\vartheta_t^\ast - \vartheta_{t\vert t-1} \big) > 0 \\ \ \iff \quad \text{ there exists } \bar{\kappa} > 0 \text{ such that for all } \kappa \in (0, \bar{\kappa}] : \ \Delta^{\mathsf{MSE}}(\phi_\kappa| \vartheta_{t\vert t-1}, p_t) < 0. \notag \end{gather} \end{proposition} Somewhat surprisingly, the MSE improvement condition (ref) coincides with the CEV condition (ref), since both criteria share the same $\mathcal{O}(\kappa)$ term (as shown in the proofs). As discussed in Section (ref), this equivalence condition is not directly useful for model construction. Moreover, MSE-improvement guarantees for SD updates require the strong Assumption (ref), which imposes a negative-definite expected Hessian and effectively forces $A\mathcal{S}_{t-1}$ to be a scalar multiple of $\Omega_{t-1}^{-1}$, a substantial restriction in higher dimensions. \subsection{EGMM of Creal2024GMM} Creal2024GMM propose to update time-varying parameters using a GMM-type influence function based on an expected GMM objective, which resembles an SD update when the moment condition uses the score. In our notation, their proposed updating rule is:\footnote{The last equation on p. 3 of Creal2024GMM is missing a minus sign.} \begin{align} \Delta \phi_{\mathrm{GMM}}(y_t, \vartheta_{t\vert t-1}) = A \, \big(-\mathbb{E}_{p_t} \big[{H}(X_t,\vartheta_{t\vert t-1}) \big] \big)^{-1} \, {s}(y_t, \vartheta_{t\vert t-1}). \end{align} This update is motivated in Creal2024GMM as being similar to an SD step with \emph{inverse Fisher scaling}. However, the proposed scaling matrix \begin{align} \mathcal{S}_{t-1}^{\mathrm{GMM}} = \Big( -\mathbb{E}_{p_t}\big[{H}(X_t,\vartheta_{t\vert t-1})\big] \Big)^{-1} \end{align} is infeasible in practice due to the dependence on the unknown true density $p_t$ and hence does not follow from the standard Fisher scaling formula (ref). In general, $\mathbb{E}_{p_t}\!\big[-{H}(X_t,\vartheta_{t\vert t-1})\big]$ should \emph{not} be interpreted as an information matrix (e.g., van1998asymptotic). An information matrix arises only under correct specification; namely, when (i) the model density $f(\cdot| \vartheta_{t\vert t-1})$ has the same functional form as the true density $p_t = p(\cdot| \lambda_t)$ for some true parameter $\lambda_t$, and (ii) the parameters coincide, $\vartheta_{t\vert t-1} = \lambda_t$. Neither condition is guaranteed in (ref). Consequently, $\mathbb{E}_{p_t}\![-{H}(X_t,\vartheta_{t\vert t-1})]$ may even be indefinite, so that gradient updates can move in directions that are poorly aligned with (or even opposite to) the score. This phenomenon can occur, for instance, in the Student's $t$ location model and the Gaussian location-scale model, discussed in Appendix (ref) and Section (ref), respectively. Creal2024GMM show that infinitesimal updates of the form (ref) improve their objective. However, their result relies on the infeasible scaling in (ref) and does hence not apply to feasible SD models. To investigate whether feasible SD models (and only SD models) are EGMM reducing, we consider general downscaled updates $\phi_\kappa$ as defined in (ref). Specifically, when the score is used as the moment condition, the expected GMM (EGMM) objective in Creal2024GMM takes the form: \begin{align} \begin{aligned} &\mathsf{EGMM} \big(p_t \Vert f_{t\vert t} \big) = \mathbb{E}_{Y_t \sim p_t} \left[ \mathbb{E}_{X_t \sim p_t} \left[{s}(X_t, \vartheta_{t\vert t}(Y_t)) \right]^\top \Omega_{t-1} \mathbb{E}_{X_t \sim p_t} \left[{s}(X_t, \vartheta_{t\vert t}(Y_t)) \right] \right], \end{aligned} \end{align} where, as before, $\Omega_{t-1}\succ O_k$ is an $\mathcal{F}_{t-1}$-measurable weighting matrix. Proposition (ref) below extends the analysis in Creal2024GMM by fully characterizing which updates $\phi$ guarantee EGMM improvements in the sense that the loss difference $\Delta^{\mathsf{EGMM}}(\phi)$ is negative: \begin{align} \Delta^{\mathsf{EGMM}}(\phi) \equiv \Delta^{\mathsf{EGMM}}(\phi| \vartheta_{t\vert t-1},p_t) := \mathsf{EGMM}(p_{t} \Vert f_{t\vert t}) - \mathsf{EGMM}(p_{t} \Vert f_{t\vert t-1}). \end{align} An update $\phi$ is \emph{EGMM reducing w.r.t.\ ${\mathcal{P}}$} if $ \Delta^{\mathsf{EGMM}}(\phi \vert \vartheta_{t\vert t-1}, p_t) < 0$ for all $\vartheta_{t\vert t-1} \in \Theta$ and $p_t \in {\mathcal{P}}$ such that $\mathbb{E}_{p_t} \left[\Delta \phi(X_t, \vartheta_{{t\vert t-1}}) \right]^\top \mathbb{E}_{p_t} \left[ - {H}(X_t, \vartheta_{t\vert t-1}) \right] \Omega_{t-1} \mathbb{E}_{p_t} \left[{s}(X_t, \vartheta_{{t\vert t-1}}) \right] \not= 0$. Here, the last condition is again analogous to (though distinct from) the one in Definition (ref) and typically prohibits that the prediction $\vartheta_{{t\vert t-1}}$ coincides with the pseudo-true parameter. \begin{proposition} Consider a class ${\mathcal{P}}$, let Assumptions (ref) and (ref) hold and let $\sup_{\vartheta \in \Theta} \big\Vert (\partial/\partial \vartheta_j) \, \mathbb{E}_{p_t} \big[ {H}(X_t, \vartheta) \big] \big\Vert < \infty$ for all $j=1,\dots,k$. Then, for $\phi_\kappa(y, \vartheta_{t\vert t-1})$ given in (ref), the following equivalence holds for each $\vartheta_{t\vert t-1} \in \Theta$ and $p_t \in {\mathcal{P}}$, for which $\mathbb{E}_{p_t} \left[\Delta \phi(X_t, \vartheta_{{t\vert t-1}}) \right]^\top \mathbb{E}_{p_t} \left[ - {H}(X_t, \vartheta_{t\vert t-1}) \right] \Omega_{t-1} \mathbb{E}_{p_t} \left[{s}(X_t, \vartheta_{{t\vert t-1}}) \right] \not= 0$: \color{black} \begin{gather} \mathbb{E}_{p_t} \left[\Delta \phi(X_t, \vartheta_{{t\vert t-1}}) \right]^\top \mathbb{E}_{p_t} \left[ - {H}(X_t, \vartheta_{t\vert t-1}) \right] \Omega_{t-1} \mathbb{E}_{p_t} \left[{s}(X_t, \vartheta_{{t\vert t-1}}) \right] > 0 \\ \ \iff \quad \text{ there exists } \bar{\kappa} > 0 \text{ such that for all } \kappa \in (0, \bar{\kappa}] : \ \Delta^{\mathsf{EGMM}}(\phi_\kappa| \vartheta_{t\vert t-1}, p_t) < 0. \notag \end{gather} \end{proposition} The assumptions of Proposition (ref) closely parallel those of Proposition 3 in Creal2024GMM. Our proof builds on arguments in their online supplement and incorporates a path-integral version of the exact multivariate mean-value theorem, yielding \begin{align*} &\Delta^{\mathsf{EGMM}}(\phi_\kappa)\notag = - 2 \kappa \mathbb{E}_{p_t} \left[\Delta \phi(X_t, \vartheta_{{t\vert t-1}}) \right]^\top \mathbb{E}_{p_t} \!\left[ - {H}(X_t, \vartheta_{t\vert t-1}) \right] \Omega_{t-1} \mathbb{E}_{p_t} \!\left[{s}(X_t, \vartheta_{{t\vert t-1}}) \right] + \mathcal{O}(\kappa^2). \end{align*} By analyzing the sign of the $\mathcal{O}(\kappa)$ term, which determines the sign of $\Delta^{\mathsf{EGMM}}(\phi_\kappa)$ for sufficiently small $\kappa$, we obtain the equivalence condition (ref) in Proposition (ref). However, (ref) is hard to interpret because it involves a nontrivial matrix product with $\mathbb{E}_{p_t}\!\left[-{H}(X_t,\vartheta_{t\vert t-1})\right]$, so Proposition (ref) provides limited guidance for model construction. To derive practically verifiable conditions, we next examine minimal conditions under which SD updates guarantee EGMM improvements. Substituting $\Delta \phi_{\mathrm{SD}}(Y_t,\vartheta_{t\vert t-1}) = A \mathcal{S}_{t-1} {s}(Y_t,\vartheta_{t\vert t-1})$ into (ref), we get \begin{equation} \mathbb{E}_{p_t} \!\left[{s}(X_t, \vartheta_{{t\vert t-1}}) \right]^\top \Big(A \mathcal{S}_{t-1} \mathbb{E}_{p_t} \!\left[ - {H}(X_t, \vartheta_{t\vert t-1}) \right] \Omega_{t-1} \Big) \mathbb{E}_{p_t} \!\left[{s}(X_t, \vartheta_{{t\vert t-1}}) \right] > 0. \end{equation} For this quadratic form to be positive for all $\mathbb{E}_{p_t}[{s}(X_t,\vartheta_{t\vert t-1})] \neq 0$, the symmetric part of the central matrix (in large brackets) must be positive definite for each $\vartheta_{t\vert t-1}$, i.e., we need \begin{equation} (A \mathcal{S}_{t-1}) \mathbb{E}_{p_t} \!\left[ - {H}(X_t, \vartheta) \right] \Omega_{t-1} + \Omega_{t-1} \mathbb{E}_{p_t}\!\left[ - {H}(X_t, \vartheta) \right] (A \mathcal{S}_{t-1}) \succ O_k, \quad \forall \vartheta \in \Theta. \end{equation} Creal2024GMM consider the infeasible scaling (ref), which is such that the expected Hessian in condition (ref) drops out. For the resulting expression to be positive definite, $A$ should still be a positive scalar multiple of the identity. However, for general feasible scalings (i.e., not depending on $p_t$), the expected Hessian in condition (ref) cannot be removed. A sufficient condition to guarantee condition (ref) is to impose $\mathbb{E}_{p_t}\!\left[{H}(X_t,\vartheta)\right]\precO_k$ and take $A\mathcal{S}_{t-1}$ as a positive scalar multiple of $\Omega_{t-1}\succ O_k$. Then (ref) reduces to $2\,\Omega_{t-1}\,\mathbb{E}_{p_t}\!\left[-{H}(X_t,\vartheta)\right]\Omega_{t-1}\succO_k$, which holds due to symmetry. To guarantee that SD updates yield EGMM improvements, we thus impose the following conditions, combining those of Proposition (ref) with those ensuring that (ref) holds. \begin{assumptionp}{$\mathcal{HBNT}$} \sloppy For all $p_t \in {\mathcal{P}}$, we have (i) $\sup_{\vartheta \in \Theta} \Vert \mathbb{E}_{p_t}[{H} (X_t, \vartheta)] \Vert \leq c < \infty$, (ii) $\mathbb{E}_{p_t} \big[ {H} ( X_t, \vartheta ) \big] \prec O_k$ for all $\vartheta\in \Theta$, (iii) $\sup_{\vartheta \in \Theta} \big\Vert (\partial/\partial \vartheta_j) \, \mathbb{E}_{p_t} \big[ {H}(X_t, \vartheta) \big] \big\Vert < \infty$, for all $j=1,\dots,k$, and (iv) $A \mathcal{S}_{t-1} = d_{t-1} \Omega_{t-1}$ for some $\mathcal{F}_{t-1}$-measurable $d_{t-1} > 0$. \end{assumptionp} Part (i) of Assumption (ref) requires the Hessian to be \emph{bounded} (denoted $\mathcal{B}$). Part (ii) imposes the expected Hessian to be \emph{negative definite} (denoted $\mathcal{N}$), while part (iii) demands the uniform boundedness of the \emph{third} derivatives (denoted $\mathcal{T}$) as in Creal2024GMM. Part (iv) is critical in the multivariate case to ensure that (ref) holds, while in the univariate case it is automatically satisfied. Unfortunately, even Assumption (ref) can be hard to verify. In particular, the expected Hessian in part (ii) is typically not negative definite unless ${H}(y,\vartheta)$ itself is negative definite for all $y\in\mathcal{Y}$ and $\vartheta\in\Theta$. If ${H}(y,\vartheta)$ fails to be negative definite on parts of $\mathcal{Y}$, one would need to restrict $p_t$ from placing too much mass there; since $p_t$ is uncontrolled, in practice this leaves few alternatives to imposing ${H}(y,\vartheta)\prec O_k$ for all $y\in\mathcal{Y},\vartheta\in\Theta$. Thus, EGMM improvements can be straightforwardly guaranteed only for densities that are log concave. For general densities and generic SD updates (with $A\mathcal{S}_{t-1} \succ O_k$, however small), EGMM improvements are not ensured. This underscores that the equivalence condition (ref) is not directly useful for model design, unlike the SEE condition (ref). \subsection{TKL of blasques2015information} blasques2015information were the first to ask what performance guarantees SD updates can deliver. Rather than taking expectations as in Theorem (ref), they study a localized KL divergence. In the univariate case ($l=k=1$), they introduce a KL-type measure that restricts $x$ to the vicinity of $y_t$ by \emph{trimming} outcomes: \begin{align} {\mathsf{TKL}}_{B}(p_{t} \Vert f_{t\vert t}) := \int_{B} \log \left( \frac{p_t(x)}{f(x | \vartheta_{t\vert t})} \right)p_t(x) \mathrm{d} x. \end{align} Here, the integration is restricted to the neighborhood ${B} \equiv {B_{\delta}(y_t)} := \{x\in{\mathcal{Y}}:\lvert x-y_t\rvert \le \delta\}$ around the realization $y_t$ for some (small) $\delta>0$, so outcomes outside this ball are trimmed. blasques2015information call an updating scheme $\phi$ (locally) “optimal” if it necessarily reduces the TKL divergence for any $y_t\in{\mathcal{Y}}$, i.e., if \begin{align} \Delta^{{\mathsf{TKL}}}_{\delta}(\phi) \equiv \Delta^{{\mathsf{TKL}}}_{\delta} (\phi|y_t, \vartheta_{t\vert t-1}, p_{t}) := {\mathsf{TKL}}_B \big(p_{t} \Vert f_{t\vert t}\big) - {\mathsf{TKL}}_B \big(p_{t} \Vert f_{t\vert t-1} \big) < 0. \end{align} The criterion $\Delta^{{\mathsf{TKL}}}_{\delta}(\phi)<0$, which we call \emph{TKL reducing}, requires that updating the model density from $f_{{t\vert t-1}}$ to $f_{{t\vert t}}$ decreases its discrepancy from the true density $p_t$, at least locally around the observation $y_t$. Roughly, blasques2015information show that SD updates are unique in guaranteeing $\Delta^{{\mathsf{TKL}}}_{\delta}(\phi)<0$ for all $y_t\in{\mathcal{Y}}$. This seems to offer an appealing theoretical property that distinguishes SD updates from other updating rules. However, while blasques2015information are formally correct, criterion (ref) relative to which they establish improvement turns out to be vacuous. To illustrate, suppose $f_{{t\vert t}}(x)>f_{{t\vert t-1}}(x)$ for all $x\in B$. Then a straightforward calculation shows that, for any $p_t$, \begin{align*} &{\mathsf{TKL}}_{B}(p_{t} \Vert f_{t\vert t}) - {\mathsf{TKL}}_{B}(p_{t} \Vert f_{t\vert t-1}) \\ &= \int_{{B_{\delta}(y_t)}} \log \left( \frac{p_t(x)}{f(x | \vartheta_{t\vert t})} \right)p_t(x) \mathrm{d} x - \int_{{B_{\delta}(y_t)}} \log \left( \frac{p_t(x)}{f(x | \vartheta_{t\vert t-1})} \right)p_t(x) \mathrm{d} x \\ &= \int_{{B_{\delta}(y_t)}} \left[ \log \big( f(x | \vartheta_{t\vert t-1}) \big) - \log \big(f(x | \vartheta_{t\vert t})\big) \right] p_t(x) \mathrm{d} x \quad <\; 0, \end{align*} where the negativity follows as $f_{t\vert t}(x) > f_{t\vert t-1}(x)$ for all $x \in B$.\footnote{ SD models, which update according to the derivative of the model density, are designed to improve the model log-density at $y_t$, such that for smooth densities, they also improve the model density for all $x \in {B_{\delta}(y_t)}$ for $\delta$ small enough. However, since the improper TKL measure is disconnected from $p_t$ such updates cannot be interpreted as genuine improvements toward the true density $p_t$.} Because the negativity holds independently of the truth $p_t$, we obtain the puzzling implication \begin{align} {\mathsf{TKL}}_B(p_t\Vert f_{t\vert t}) < {\mathsf{TKL}}_B(p_t\Vert f_{t\vert t-1}) \quad \text{ whenever } \quad f_{t\vert t}(x) > f_{t\vert t-1}(x) \;\; \text{for all } x \in B. \end{align} This shows that trimming-based localization yields a criterion that is disconnected from the true density and hence uninformative. For instance, even if $f_{t\vert t-1}=p_t$, the TKL criterion would still favor moving away from the true density $p_t$ whenever $f_{t\vert t}>f_{t\vert t-1}$ on $B$. This is also evident in Example (ref) and panel (b) of Figure (ref), where the updated (blue) density is not closer to $p_t$ near $y_t$ than the original (red) density, despite the TKL improvement. The puzzling implication in (ref) (i.e., that improvements are always possible irrespective of $p_t$) has long been known in the literature on localized scoring rules. Since the KL divergence corresponds to the expected logarithmic score, TKL relates to the trimmed logarithmic score of amisano2007comparing. Subsequent work has shown that trimming yields an improper scoring rule, and that localization should instead use censoring to preserve (strict) propriety; see diks2011likelihood, gneiting_comparing_2011, and depunder2023localzing. Following this suggestion of censoring, in Appendix (ref) we show that sufficiently small score-driven updates are censored KL (CKL) reducing if and only if $p_t(y_t)>f(y_t| \vartheta_{t| t-1})$. While the dependence on the true density $p_t$ is thus restored, this condition is not verifiable in practice and provides little guidance for model construction. Having shown that the TKL measure is an improper criterion, we do not consider it further. \section{Examples} Here we take several popular model densities and investigate whether the respective guarantees for SD updates being EKL, CEV, MSE or EGMM reducing can be applied. Section (ref) presents results for eleven models with a univariate time-varying parameter and Section (ref) considers a Gaussian model with a bivariate location-scale time-varying parameter. To interpret the results below, recall that the performance measures' applicability to SD models primarily differs in their requirements on the expected Hessian. EKL improvements require only (localized) boundedness (Assumptions (ref), (ref)), whereas CEV, MSE, and EGMM improvements (additionally) demand negative definiteness, which is typically possible only if the realized Hessian itself is almost surely negative definite. In the multivariate case, Assumptions (ref) and Assumption (ref) further impose (differing) restrictions on the learning-rate and scaling matrices. \subsection{Applicability to eleven univariate model densities} Table (ref) illustrates the broad applicability of our EKL reduction guarantees of SD updates across eleven model densities with a univariate time-varying parameter. Nine of these models are adapted from koopman2016predicting, complemented by two conditional location (local-level) models with Gaussian and Student’s $t$ distributions. The collection is deliberately diverse, covering models with time-varying parameters for level, scale, dependence, count, intensity, and duration. These densities---or probability functions, in the case of discrete observations (see Remark (ref))---are postulated by the researcher, while the true data-generating density remains unknown. In Appendix (ref), we provide full details of the postulated models in Table (ref) and discuss three models in detail, related to (i) intensities, (ii) negative binomial counts, and (iii) the Student’s $t$ location model. For each of the eleven models, Table (ref) reports the range of the expected Hessians, together with the moment conditions on the class ${\mathcal{P}}$ of true distribution $p_t$ required for these results. The ranges are then used to assess whether the corresponding assumptions are satisfied---indicated by check or cross marks---for each performance criterion. The required moment conditions are mild, involving only first and second moments, and ensure that the random variables are not zero almost surely under the true distribution. \begin{table}[t] \caption{Applicability of performance guarantees to eleven postulated densities. } \begin{footnotesize} \begin{threeparttable} \begin{tabular}{l@l@c@c@c@c@c@c@c@c} \toprule \multicolumn{2}{l}{\multirow{1}{*}{\bf Postulated model}} & \multicolumn{2}{c}{\multirow{1}{*}{\bf Properties of expected Hessian}} & \multicolumn{4}{c}{\multirow{1}{*}{\bf Performance criteria}} \\ \cmidrule(r{5pt}l{5pt}){1-2} \cmidrule(r{5pt}l{5pt}){3-4} \cmidrule(r{5pt}l{5pt}){5-8} \multirow{2}{*} & \multirow{2}{*} & \multirow{2}{*} & \multirow{2}{*} & EKL & EKL & CEV/MSE & EGMM \\ & & Moment condition & Range of & Thm. (ref) & Thm. (ref) & Pos. (ref) & Pos. (ref) \\ Type & Distribution & (restriction on $\mathcal{P}$) & expected Hessian & (ref) & (ref) & (ref) & (ref) \\ \cmidrule(r{5pt}l{5pt}){1-2} \cmidrule(r{5pt}l{5pt}){3-4} \cmidrule(r{5pt}l{5pt}){5-8} Count & Poisson & \multicolumn{1}{c}{$-$} & \multicolumn{1}{c}{$(-\infty, 0)$} & \color{red} \ding{55} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & \color{red} \ding{55} \\ \addlinespace Count & Neg. Binomial & \multicolumn{1}{l}{$\mathbb{E}_{p_t}[Y_t]<\infty$} & \multicolumn{1}{c}{$\left[- \tfrac{\xi + \mathbb{E}_{p_t}[Y_t]}{4}, 0 \right)$} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} \\ \addlinespace Intensity & Exponential & \multicolumn{1}{l}{$0 < \mathbb{E}_{p_t}[Y_t]<\infty$} & \multicolumn{1}{c}{$(-\infty, 0)$} & \color{red} \ding{55} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & \color{red} \ding{55} \\ \addlinespace Duration & Gamma & \multicolumn{1}{l}{$0 < \mathbb{E}_{p_t}[Y_t]<\infty$} & \multicolumn{1}{c}{$(-\infty, 0)$} & \color{red} \ding{55} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & \color{red} \ding{55} \\ \addlinespace Duration & Weibull & \multicolumn{1}{l}{$0 < \mathbb{E}_{p_t} [Y_t^\xi] <\infty$} & \multicolumn{1}{c}{$(-\infty, 0)$} & \color{red} \ding{55} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & \color{red} \ding{55} \\ \addlinespace Volatility & Gaussian & \multicolumn{1}{l}{$0 < \mathbb{E}_{p_t}[Y_t^2]<\infty$} & \multicolumn{1}{c}{$(-\infty, 0)$} & \color{red} \ding{55} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & \color{red} \ding{55} \\ \addlinespace Volatility & Student's \emph{t} & \multicolumn{1}{c}{$0<\mathbb{E}_{p_t}[Y_t^2]\qquad\;\;\;$} & \multicolumn{1}{c}{$[-\tfrac{\nu+1}{8},0)$} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} \\ \addlinespace Dependence & Gaussian & \multicolumn{1}{l}{$\mathbb{E}_{p_t} [Y^2_{jt}]<\infty$} & \multicolumn{1}{c}{$\left(-\infty,\tfrac{1}{4}\right]$} & \color{red} \ding{55} & \color{mygreen} \ding{51} & \color{red} \ding{55} & \color{red} \ding{55} \\ \addlinespace Dependence & Student's \emph{t} & \multicolumn{1}{c}{$-$} & \multicolumn{1}{c}{$[-\tfrac{\nu+1}{4},\tfrac{1}{4}]$} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & \color{red} \ding{55} & \color{red} \ding{55} \\ \addlinespace Level & Gaussian & \multicolumn{1}{c}{$-$} & \multicolumn{1}{c}{$\left[-\sigma^{-2},-\sigma^{-2}\right]$} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} \\ \addlinespace Level & Student's \emph{t} & \multicolumn{1}{c}{$-$} & \multicolumn{1}{c}{$\left[-\tfrac{\nu+1}{\nu \, \sigma^2},\tfrac{\nu+1}{8\,\nu \, \sigma^2}\right]$} & \color{mygreen} \ding{51} & \color{mygreen} \ding{51} & \color{red} \ding{55} & \color{red} \ding{55}\\ \bottomrule \end{tabular} {\textbf{NOTE:} The column \emph{Postulated model} specifies the model type and distribution. The stated range for the expected Hessian, $\mathbb{E}_{p_t}[{H}(Y_t,\vartheta)]$, holds under the moment conditions in the column \emph{Moment condition} (with $j\in\{1,2\}$), which restricts the class ${\mathcal{P}}$ of permitted true distributions; these conditions are verified in Table (ref) in Appendix (ref). The lower bounds in the moment conditions exclude $Y_t=0$ a.s.\ and are only required to satisfy Assumption (ref). In \emph{Performance criteria}, line 1 gives the criterion, line 2 the corresponding result/condition, and line 3 the relevant assumption (in addition to Assumption (ref)). Check and cross marks in the table indicate which assumptions hold. } \end{threeparttable} \end{footnotesize} \end{table} From the ranges and the expected Hessians reported in Table (ref), we observe that Assumption (ref) is satisfied for five of the eleven models, thereby guaranteeing EKL improvements for classes ${\mathcal{P}}$ that are only restricted by the given moment conditions. For the remaining six models, the expected Hessian is unbounded in the time-varying parameter. This issue is addressed by Assumption (ref), which applies to all eleven models. Under this assumption, Theorem (ref) and Proposition (ref) ensure that clipped SD updates, with sufficiently large clipping constants, are EKL reducing. In contrast, the CEV and MSE criteria of Gorgi2023 rely on a negative, though potentially unbounded, expected Hessian as specified in Assumption (ref), which is satisfied by eight of the models. For the three practically relevant cases where this assumption fails, it appears impossible to impose reasonable conditions on ${\mathcal{P}}$ that would restore Assumption (ref); see Appendix (ref) for details on the local-level model with a Student’s $t$ distribution discussed in harvey2014filtering. Finally, guarantees for the EGMM criterion hold only in three cases where the strongest Assumption (ref) is met. In sum, Table (ref) shows that the EKL criterion, under the mild Assumption (ref), provides the only formal result currently applicable across the full class of models considered. \subsection{Bivariate dynamic parameter: Gaussian location-scale model} We now illustrate in a simple bivariate setting that only the EKL result remains applicable. Consider the Gaussian location-scale model $\mathcal{N}\big(\mu_{{t\vert t-1}},\,\exp(\lambda_{t\vert t-1})\big)$ with time-varying mean $\mu_{{t\vert t-1}}\in \mathbb{R}$ and variance $\exp(\lambda_{{t\vert t-1}})>0$, which we collect as a bivariate time-varying parameter $\vartheta_{{t\vert t-1}} = (\mu_{{t\vert t-1}}, \lambda_{{t\vert t-1}})^\top \in {\mathbb{R}}^2 = \Theta$. The exponential link guarantees positivity for any $\lambda_{t\vert t-1}\in \mathbb{R}$. For $\vartheta=(\mu, \lambda)^\top \in \mathbb{R}^2$, a direct calculation for the Hessian matrix yields \[ {H}(x,\vartheta) = \begin{pmatrix} -\,\mathrm{e}^{-\lambda} & (\mu-x)\,\mathrm{e}^{-\lambda}\\[2pt] (\mu-x)\,\mathrm{e}^{-\lambda} & -\tfrac{1}{2}(\mu-x)^2 \mathrm{e}^{-\lambda} \end{pmatrix}. \] As the top-left element is negative, while the determinant $-\tfrac{1}{2}(\mu-x)^2 \mathrm{e}^{-2\lambda}$ is non-positive, the Hessian is indefinite (unless $x=\mu$, where the Hessian is negative semi-definite). Next, we consider the expected Hessian. Suppose $X_t \mid {\mathcal{F}}_{t-1} \sim p_t \in {\mathcal{P}}$, the class restricted to finite means and variances, denoted $\mu_{p_t}$ and $\sigma_{p_t}^2$. Then, for any $\mu \in \mathbb{R}$, we have $\mathbb{E}_{p_t}[(\mu-X_t)^2] = (\mu_{p_t}-\mu)^2+\sigma_{p_t}^2$ and the expected Hessian becomes \begin{equation} \mathbb{E}_{p_t}[{H}(X_t,\vartheta)] = \begin{pmatrix} -\,\mathrm{e}^{-\lambda} & (\mu-\mu_{p_t})\,\mathrm{e}^{-\lambda}\\[2pt] (\mu-\mu_{p_t})\,\mathrm{e}^{-\lambda} & -\tfrac{1}{2}\big((\mu_{p_t}-\mu)^2+\sigma_{p_t}^2\big)\,\mathrm{e}^{-\lambda} \end{pmatrix}. \end{equation} As this matrix is bounded for $(\mu,\lambda)^\top$ in compact subsets of $\mathbb{R}^2$, Assumption (ref) holds. By Theorem (ref) and Proposition (ref), therefore, clipped SD updates with sufficiently large clipping constants and small learning rates are EKL reducing w.r.t.\ ${\mathcal{P}}$. For the CEV, MSE, and EGMM criteria with $\Omega_{t-1}=I_k$, Assumptions (ref) (ii) and (ref) (iv) force $A\mathcal{S}_{t-1}$ to be a scalar multiple of the identity and require a negative definite expected Hessian. In (ref), the expected Hessian has negative diagonal entries, but its determinant is $\frac{1}{2}\big(\sigma_{p_t}^2-(\mu-\mu_{p_t})^2\big)\mathrm{e}^{-2\lambda}$. When $\mu$ is sufficiently far from $\mu_{p_t}$, this determinant is negative, so the expected Hessian is \emph{indefinite} (i.e., it has both positive and negative eigenvalues). Negative definiteness of the expected Hessian in Assumptions (ref) and (ref) fails; i.e., no improvement guarantees for CEV, MSE, or EGMM are available. In sum, even in this simple bivariate case, a Gaussian distribution with time-varying mean and variance (e.g., as in ARMA--GARCH-type specifications), Hessian negative definiteness fails almost surely as well as in expectation. Improvement guarantees arise only under the EKL criterion, as proposed here, under a localized expected-Hessian boundedness condition that is considerably less restrictive than alternative versions in the literature. \section{Conclusion} We have characterized score-driven (SD) updates via the expected Kullback-Leibler (EKL) measure. We showed that EKL improvements occur if and only if the expected parameter adjustment aligns with the expected score; i.e., their inner product should be positive. This equivalence continues to hold under merely locally bounded Hessians when updates are clipped. The resulting conditions are interpretable and constructive, directly guiding update design. By deriving explicit upper bounds on admissible learning-rate matrices in terms of score moments, we also connected SD models to adaptive optimization methods. Relative to recent alternative performance measures, our approach is advantageous in (i) yielding equivalence conditions directly useful for model construction, (ii) requiring the mildest Hessian conditions, as illustrated by examples, (iii) extending naturally to the multivariate case without restrictive choices of weighting, scaling, or learning-rate matrices, and (iv) providing intuitive, constructive upper bounds for learning-rate matrices. Together, these findings provide a rigorous justification for SD updates and establish the EKL divergence as their natural information-theoretic foundation. \section*{Acknowledgments} We thank Peter Boswijk, Cees Diks, Dick van Dijk, Simon Donker van Heel, Andrew Harvey, Yi He, Frank Kleibergen, Roger Laeven, Alessandra Luati, André Lucas, Bram van Os, Andrew Patton, Phyllis Wan and Chen Zhou for their valuable comments. We also thank participants at seminars at Heidelberg University and Goethe University Frankfurt, and at the 2024 ISF (Dijon), the 2024 Bernoulli-IMS World Congress (Bochum), the 2025 QFFE Conference (Marseille), the 2025 IAAE Annual Conference (Turin), and the 2025 CFE-CMStatistics Conference (London). T. Dimitriadis gratefully acknowledges support from the German Research Foundation (DFG) under project number 502572912. \singlespacing \spacingset{1} {2pt} \putbib[Bibliography-LSPS-v2]