EconBase
← Back to paper

Tweedie Calculus

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

91,244 characters · 14 sections · 55 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Tweedie Calculus

titlepage\begin{abstract} Tweedie’s formula is central to measurement-error analysis and empirical Bayes. Under Gaussian noise, the formula identifies the posterior mean directly from the observed-data density, bypassing nonparametric deconvolution. Beyond a few classical examples, however, no general theory explains when analogous identities hold, how they are structured, or how to derive them for non-Gaussian noise and for posterior functionals other than the mean. This paper develops such a framework for additive-noise models. I characterize when conditional expectations of an unobserved latent variable, given the observed signal, admit direct expressions in terms of the observed density---identities I call Tweedie representations---and show that they are governed by a linear map, the Tweedie functional. Under general conditions, I prove that this functional exists, is unique, and is continuous. I also provide a constructive method for deriving it by extending the inverse Fourier transform of an explicit tempered distribution. This recasts the search for Tweedie-type formulas as a problem in the calculus of tempered distributions. The framework recovers the classical Gaussian formula and yields new representations for posterior means under non-Gaussian noise. I apply the method to construct unbiased representations of nonlinear functionals of latent variables and to derive Tweedie formulas for the product-Laplace mechanism used in differential privacy. Finally, I show that the approach extends beyond the standard additive model. In the heteroskedastic Gaussian sequence model, where the noise covariance is itself random, a change of variables restores the required additive-noise structure conditionally, yielding Tweedie representations without additional restrictions on the joint law of the latent parameter and noise covariance. \end{abstract} \setcounter{page}{0} \thispagestyle{empty}

\doparttoc \faketableofcontents

Introduction

Statistical analysis often seeks to recover latent heterogeneity from noisy measurements. In compound decision, measurement-error, and empirical Bayes problems, objects of interest such as posterior means and decision rules depend on the unknown distribution of latent effects. This naturally shifts attention to recovering that distribution from the data, a task known as deconvolution. Deconvolution, however, is notoriously difficult: it is an ill-posed inverse problem, and the object to be recovered is itself infinite-dimensional.

Tweedie’s formula shows that deconvolution can sometimes be sidestepped altogether. Suppose that $Y \sim \mathcal{N}(X, \sigma^2)$, where $X \sim P_X$ is an unknown latent signal. robbins1956Tweedie credits Maurice Tweedie with the identity

equation[equation omitted — 100 chars of source]

where $f_Y$ is the marginal density of the observable $Y$. The key feature of this identity is that it represents the posterior mean directly as a function of the observable density, without requiring recovery of the latent distribution $P_X$.

Accordingly, Tweedie's formula offers a different route to estimation and inference. Instead of estimating the latent signal distribution, one can express posterior functionals directly as maps of the observed marginal density, which can itself be estimated from the data.\footnote{In empirical Bayes terminology, the former approach is \(g\)-modeling and the latter is \(f\)-modeling efron2014two,efron2019bayes.} This perspective has further implications beyond estimation. When the posterior mean is used as a decision rule, its Tweedie representation gives it a shrinkage interpretation: the posterior mean equals the raw observation plus a correction determined by the observed marginal density. In the Gaussian location model, this correction moves \(y\) toward regions of higher observed density, with magnitude governed by the noise variance and the score \(f_Y'(y)/f_Y(y)\).\footnote{This shrinkage view includes the James-Stein estimator james1961estimation as an instance: in the Gaussian sequence model with a Gaussian latent distribution, the empirical Bayes plug-in version of Tweedie's formula yields James--Stein shrinkage up to degrees-of-freedom corrections efron2016empirical. } In some models, Tweedie-type representations also imply shape restrictions, such as monotonicity, which can be used to improve estimation houwelingen1983monotone. They also shed light on the statistical difficulty of estimation: in the Gaussian posterior-mean case, the difficulty is that of estimating the score of the observed marginal density, a pointwise derivative-estimation problem with known minimax rates under standard smoothness conditions stone1983optimal,pinkse2023estimates.

Although formulas of this form are available for particular models and particular estimands, there is no general framework that explains when they exist, what structure they share, or how they can be derived. This paper provides such a framework. I consider general additive-noise models: triples \((Y,X,V)\) of \(\mathbb{R}^d\)-valued random variables satisfying \(Y=X+V\). For a broad class of functions \(g\) and noise laws \(V\), I study when posterior expectations\footnote{I use the term \textquote{posterior} rather than \textquote{conditional} following a Bayesian or empirical Bayes interpretation: \(X\) is the latent signal, \(Y\) is the observation, and \(\mathbb E[g(X)\mid Y=y]\) is the posterior expectation of \(g(X)\) after observing \(Y=y\).} admit a representation of the form

equation[equation omitted — 123 chars of source]

where \(\mathcal{T}_{g,V,y}\) is a linear functional on a function space containing all possible observed densities. It depends on the estimand \(g\), the noise law of \(V\), and the evaluation point \(y\), but not on the latent distribution \(P_X\). For example, in the one-dimensional Gaussian case with \(g(x)=x\), the functional in question is \[ \mathcal{T}_{g,V,y}[f_Y] = y f_Y(y) + \sigma^2 f_Y'(y), \] which recovers (ref). I call \(\mathcal{T}_{g,V,y}\) the Tweedie functional, and I call any expression of the form (ref) a Tweedie representation. I characterize when such functionals exist and provide a general strategy to compute them in closed form.

Several functionals fit into this notation. The examples below are common in the literature and will be used throughout the paper.

itemize• Posterior moments. For a scalar latent variable, posterior moments correspond to polynomial choices of \(g\). The most used is the posterior mean, which is the Bayes rule for estimating the latent variable under squared-error loss. The conditional risk of this rule is the posterior variance. Since the posterior variance is determined by the first two posterior moments, it also admits a Tweedie representation whenever those moments do. More generally, posterior risks under polynomial loss functions can be written in terms of finitely many posterior moments, so they inherit Tweedie representations from the corresponding moment representations. • Posterior cumulative distribution function. For \(t\in\mathbb R\), the posterior cumulative distribution function corresponds to the indicator \(g_t(x)=\mathbbm{1}\{x\le t\}\). Thus \(F_{X\mid Y}(t\mid y)=\mathbb E[\mathbbm{1}\{X\le t\}\mid Y=y]\). This functional encompasses several quantities of interest. In large-scale studies, a common target is the posterior probability that an effect is positive, \(\Pr(X>0\mid Y=y)=1-F_{X\mid Y}(0\mid y)\). Such sign probabilities arise when assessing directional evidence across many units and have been studied in both theoretical and applied work; see es2005asymptotic,dattner2011deconvolution,efron2016empirical,greenshtein2018application,brennan2020estimating. A related quantity is the local false sign rate, \(\min\{F_{X\mid Y}(0\mid y),1-F_{X\mid Y}(0\mid y)\}\), which measures the smaller posterior probability assigned to the two possible signs stephens2017false. • Posterior moment generating function. For \(t\in\mathbb R^d\), the posterior moment generating function corresponds to \(g_t(x)=\exp(t^\top x)\). When finite in a neighborhood of zero, the posterior moment generating function characterizes the posterior distribution and can be used to recover posterior moments and other distributional features. Moment generating functions also provide a standard tool for deriving tail bounds; see wainwright2019high.

The preceding examples motivate three questions about Tweedie representations that organize the paper:

enumerate[label=(\roman*)] • Existence and regularity of the Tweedie functional. Section (ref) asks when a posterior functional admits a Tweedie representation of the form (ref). The main result, Theorem (ref), gives sufficient conditions under which the Tweedie functional exists, is unique, and is continuous in a natural topology. The theorem applies to a broad class of additive-noise models, including cases where the law of \(V\) need not be symmetric or have full support. Although the result is not constructive, it gives a general structural characterization of such representations. • Strategy for computing Tweedie representations. Section (ref) builds on the existence result and develops a general method for deriving Tweedie representations through Fourier analysis. The main tool is the theory of tempered distributions, developed by schwartz1966theorie and reviewed in Section (ref). This language is needed because the relevant inverse Fourier transforms need not exist as ordinary functions. More importantly, it also provides a Fourier theory for linear functionals: Fourier inversion applies not only to functions, but also to linear maps, yielding a calculus for functionals. The main result, Theorem (ref), shows that, under suitable conditions, the Tweedie functional can be computed from the inverse Fourier transform of a known, computable linear functional. A major practical advantage of this viewpoint is that it reduces the derivation of Tweedie representations to the computation of inverse Fourier transforms. In this sense, the problem becomes one of calculus. As such, the relevant inverse transform can be computed directly or identified from standard transform tables, much as Laplace-transform tables are used in the study of differential equations. Useful references for this purpose include bateman_1954 and kammler2007first. • Applications. Section (ref) illustrates the scope of the framework through several applications. First, it uses the Fourier-based strategy of Section (ref) to derive posterior-mean representations for a broad class of noise distributions. Table (ref) summarizes these formulas. Except for the Gaussian and Laplace cases, these representations appear to be new. Second, the section identifies cases in which the Tweedie functional itself can be written as an expectation of an explicit function of the observed variable. This connects the framework to the classical problem of unbiased estimation under additive noise, going back to kolmogorov1950unbiased in the Gaussian case. Third, Section (ref) applies the framework to the conventional Laplace mechanism from differential privacy. In particular, the paper derives Tweedie representations for several posterior functionals under this noise structure. Table (ref) summarizes selected cases. Section (ref) studies the heteroskedastic Gaussian sequence model. This model is especially important in economics and the social sciences, where researchers often observe noisy estimates of means or effects across many units or populations.\footnote{walters2024empirical provides a review of common applications of this model in economics.} In this model, \(Y\) is a noisy estimate of a latent parameter \(X\), and the noise covariance matrix \(\Sigma\) may vary across units. It may also be random and correlated with \(X\). The key observation is that, after conditioning on \(\Sigma=\Sigma_0\) and standardizing, the model reduces to an additive Gaussian model with fixed covariance. The general framework therefore applies conditionally on \(\Sigma\), yielding Tweedie representations in terms of the conditional density \(f_{Y\mid\Sigma}(\cdot\mid\Sigma_0)\). This conditional formulation gives explicit representations for posterior means and for other empirical Bayes functionals, including posterior covariances, moments, risk functions, distribution functions, and moment generating functions.

\paragraph{Relation to the literature.}

The paper by raphan2011least is closest to my approach. They show that, for several additive-noise models, the posterior mean can be written as the ratio of a linear operator applied to the observed marginal density and the density itself. They further note that, once such an operator is available, iterating it recovers posterior moments and, more generally, posterior expectations of polynomial functions. This paper formalizes and extends that observation. Rather than build outward from the posterior-mean case, I study when, for a general measurable function \(g\), the conditional expectation \(\mathbb E[g(X)\mid Y=y]\) admits an analogous representation. I also establish formal properties of the representing map, including existence, uniqueness, and continuity under the relevant topology. The framework also extends beyond the basic additive-noise setting, for example to heteroskedastic Gaussian sequence models, where the additive-noise structure holds only conditionally.

On the computational side, the paper's main innovation is to introduce tempered distributions as a tool for deriving Tweedie representations. This serves two purposes. First, it makes rigorous the Fourier calculations that underlie many Tweedie-type identities for posterior moments, including cases in which the relevant inverse Fourier transforms exist only as distributions. Second, it yields a systematic procedure for deriving Tweedie representations for posterior functionals beyond moments, including nonsmooth and non-polynomial targets. This approach complements derivations that exploit special algebraic structure in the sampling model. This approach complements derivations that exploit special algebraic structure in the sampling model. For example, in exponential families, efron2011tweedie observes that posterior-moment representations can be obtained by differentiating the log ratio of the marginal density to the base-measure density defining the family. Related identities, including Stein’s lemma and Stein’s unbiased risk estimate, are often derived by integration by parts stein1981estimation.

The paper also provides several new Tweedie representations. For instance, it derives a Tweedie representation for the posterior cumulative distribution function under Gaussian noise. It also derives analogous representations for posterior means under several non-Gaussian noises, including Cauchy, Gumbel, and Gamma noise. In this respect, the paper is related to the literature on deriving new Tweedie-type identities in Bayesian, empirical Bayes, and measurement-error problems. Historically, the Gaussian identity appears in print in dyson1926method, where Dyson attributes it to Arthur Eddington. Later, robbins1956Tweedie reported the formula and credited Maurice Kenneth Tweedie. Much of the subsequent literature develops model-specific analogues or extensions of this identity. west1982aspects derives a scale-parameter analogue of Tweedie's formula for models with known location and unknown scale. pericchi1992exact and pericchi1993posterior derive exact relationships for posterior moments and cumulants in normal-location and exponential-family settings. Polson1991 obtains an exact representation of the posterior mean for location models with normal scale-mixture priors, and polson2019bayesian gives a Gaussian linear-regression analogue. robbins1982estimating derives a chi-square analogue in the empirical Bayes problem of estimating many variances. shi2015nonlinear develop Laplace-noise-based Tweedie-type formulae.

Finally, the paper relates to applications of Tweedie's formula. These methods are widely used in empirical Bayes problems efron2011tweedie,efron2014two, including large-scale hypothesis testing efron2012large. More recently, Tweedie formulas have been used in diffusion-based sampling and score-based generative modeling montanari2023sampling,song2020score,sohl2015deep. Related denoising problems also arise in differential privacy, where Gaussian or Laplace noise is often added to statistics to limit the information revealed about any single data point. For Laplace noise, hong2003simple propose a correction based on the Tweedie representer for nonlinear method-of-moments models with measurement error, and shi2019estimation apply Laplace-based Tweedie-type formulas to measurement-error and regression problems. Tweedie's formula has also been used for forecasting in dynamic panel models Liu2020. Recent work by liang2025distributional,liang2025distributionalII uses identities closely related to Tweedie's formula to construct shrinkage maps for recovering the latent signal distribution. More broadly, the framework is connected to the deconvolution literature fan1991optimal,butucea2009adaptive,pensky2017minimax,hall2007ridge and to additive measurement-error models carroll2006measurement,meister2009,schennach2004estimation,schennach2004nonparametric,schennach_2013.

Setup and basic notions in Fourier Analysis

This section introduces the additive-noise model studied throughout the paper, states the standing assumptions under which the theory is developed, and records several immediate consequences. The framework is intentionally broad: it includes the classical Gaussian model underlying Tweedie's formula and accommodates more general deconvolution problems in arbitrary dimensions.

Setup

Let $\mathcal{P}(\mathbb{R}^d)$ denote the set of probability measures on $(\mathbb{R}^d,\mathcal{B}(\mathbb{R}^d))$, where $\mathcal{B}(\mathbb{R}^d)$ is the Borel $\sigma$-algebra on $\mathbb{R}^d$. For a random vector \(Z\), let \(P_Z\) denote its law. Consider a triplet of $d-$dimensional random vectors $(Y,X,V)$ defined on the same probability space $(\Omega,\mathscr{A}, \mathbb{P} )$. Here $Y \sim P_Y$ is the observed vector, $X \sim P_X$ is the latent signal of interest, and $V \sim P_V$ is the noise.

For any $P \in \mathcal{P}(\mathbb{R}^d)$, let $\varphi_P : \mathbb{R}^d \to \mathbb{C}$ denote its characteristic function, defined by \[ \varphi_P(\omega) \coloneq \int_{\mathbb{R}^d} e^{i\omega^\top x}\,dP(x) = \mathbb{E}_{Z\sim P}\!\left[e^{i\omega^\top Z}\right], \qquad \omega \in \mathbb{R}^d. \] When $Z \sim P_Z$, I use the shorthand notation $\varphi_{Z} \coloneq \varphi_{P_Z}$. I now state the main assumptions used throughout the paper.

assumption[Additivity and independence] The triple $(Y,X,V)$ is related by addition $Y=X+V$, and \(X\) and \(V\) are independent.

Under this assumption, the observed characteristic function factors as \(\varphi_Y(\omega)=\varphi_X(\omega)\varphi_V(\omega)\) for all \(\omega\in\mathbb R^d\).\footnote{schennach2019convolution emphasizes that such factorization can hold even without independence of \(X\) and \(V\). In this paper, however, full independence is needed to connect deconvolution to conditional expectations. It is also the standard assumption in additive-noise applications.} This factorization is what makes closed-form Tweedie representations possible: it allows terms involving the unknown latent law \(P_X\), encoded by \(\varphi_X\), to be rewritten in terms of the observed marginal law and the known noise law, characterized by \(\varphi_Y\) and \(\varphi_V\).

assumption[Identification] The noise characteristic function is nonzero almost everywhere: \[ \varphi_V(\omega)\neq 0 \quad\text{for Lebesgue-a.e. }\omega\in\mathbb{R}^d. \]

Assumption (ref) is the basic nondegeneracy condition behind Fourier deconvolution. Informally, recovering $P_X$ from the relation $\varphi_Y(\omega)=\varphi_X(\omega)\varphi_V(\omega)$ requires division by $\varphi_V(\omega)$. Therefore, this assumption rules out, except on a null set, frequency regions at which inversion breaks down. It is satisfied by the standard Gaussian, Laplace, logistic, and many other common noise laws. It also holds, for example, when \(V\) has compact support.\footnote{A simple sufficient condition is the existence of an exponential moment: \(\mathbb E [e^{r\|V\|_2}]<\infty\) for some \(r>0\). Then \(\varphi_V\) extends holomorphically to a complex neighborhood of \(\mathbb R^d\). Since \(\operatorname{Re}(\varphi_V)(0)=1\), its real part is a nonzero real-analytic function. Accordingly, the zero set of \(\operatorname{Re}\varphi_V\) has Lebesgue measure zero, and so does the zero set of \(\varphi_V\) mityagin2020zero.}

assumption[Noise regularity] The noise distribution $P_V$ admits a Lebesgue density $f_V$.

Assumption (ref), together with the convolution structure imposed by Assumption (ref), has two immediate consequences. First, the observed law \(P_Y\) is absolutely continuous with respect to Lebesgue measure, so the observed density \(f_Y\) appearing in Tweedie representations is well defined. Second, as Proposition (ref) shows, this density is obtained by convolving the latent law \(P_X\) with the noise density \(f_V\). This convolution formula will be essential for characterizing the Tweedie functional.

propositionUnder assumptions (ref) and (ref), the law $P_Y$ is absolutely continuous with respect to Lebesgue measure. A version of its density is \[ f_Y(y)=\int_{\mathbb{R}^d} f_V(y-x)\,dP_X(x), \qquad y\in\mathbb{R}^d. \]

Some important notions in Fourier analysis

This subsection fixes the Fourier analysis conventions used throughout the paper. It first defines the ordinary Fourier transform for integrable functions and then extends the transform to tempered distributions.

For \(1\le p\le\infty\) and \(\mathbb K\in\{\mathbb R,\mathbb C\}\), let \(L^p(\mathbb R^d;\mathbb K)\) denote the usual Lebesgue space, with norm \[ \|f\|_p =

cases\left(\int_{\mathbb R^d}|f(x)|^p\,dx\right)^{1/p}, & 1\le p<\infty,\\[1ex] \operatorname*{ess\,sup} \limits_{x\in\mathbb R^d}|f(x)|, & p=\infty.

\] When \(\mathbb K\) is omitted, it is understood to be \(\mathbb R\). The symbol \(i\) denotes the imaginary unit, and \(\overline z\) denotes the complex conjugate of \(z\in\mathbb C\). Let \(C_0(\mathbb R^d;\mathbb K)\) denote the space of continuous \(\mathbb K\)-valued functions on \(\mathbb R^d\) that vanish at infinity.

For \(\psi\in L^1(\mathbb R^d;\mathbb C)\), define its Fourier transform by \[ \widetilde\psi(\omega) \coloneq \int_{\mathbb R^d} e^{i\omega^\top x}\psi(x)\,dx, \qquad \omega\in\mathbb R^d. \] For \(\psi\in L^1(\mathbb R^d;\mathbb C)\), let \(\psi^\sharp\) denote the inverse Fourier integral \[ \psi^\sharp(x) \coloneq \frac{1}{(2\pi)^d} \int_{\mathbb R^d} e^{-i\omega^\top x}\psi(\omega)\,d\omega, \qquad x\in\mathbb R^d, \] whenever this integral is well defined. Thus \(\psi\mapsto\psi^\sharp\) denotes the inverse Fourier operation under the convention above. If \(\psi,\widetilde\psi\in L^1(\mathbb R^d;\mathbb C)\), Fourier inversion gives \[ \psi(x) = \widetilde\psi^\sharp(x) = \frac{1}{(2\pi)^d} \int_{\mathbb R^d} e^{-i\omega^\top x}\widetilde\psi(\omega)\,d\omega \] for Lebesgue-a.e. \(x\in\mathbb R^d\).

For the present paper, this classical framework is too narrow. Many objects that arise in the computation of Tweedie representations are not integrable, so their inverse Fourier transforms need not exist as functions. More importantly, Tweedie functionals are linear maps acting on the observed density. Consequently, it is advantageous to use a Fourier theory that can transform and invert linear functionals, rather than one limited to ordinary functions. I use the framework of tempered distributions, introduced by schwartz1966theorie, for this purpose. This framework extends Fourier analysis from functions to continuous linear functionals and thus provides a natural language for deconvolution problems.

The starting point is the Schwartz space on \(\mathbb R^d\), denoted by \(\mathscr S(\mathbb R^d)\). It is the space of smooth test functions against which tempered distributions will be evaluated. Formally, \[ \mathscr{S}(\mathbb{R}^d) \coloneq \left\{ \psi\in C^\infty(\mathbb{R}^d;\mathbb{C}) : \sup_{x\in\mathbb{R}^d} \bigl|x^{\alpha} \partial^{\beta} \psi(x)\bigr|<\infty \text{ for all multi-indices }\alpha,\beta\in\mathbb{N}_0^d \right\}. \] Here \(x^{\alpha} \coloneq x_1^{\alpha_1}\cdots x_d^{\alpha_d}\), \(\partial^{\beta} \coloneq \partial^{|\beta|}/ \partial x_1^{\beta_1}\cdots \partial x_d^{\beta_d}\), and \(|\beta|=\beta_1+\cdots+\beta_d\). Thus \(\mathscr S(\mathbb R^d)\) consists of smooth functions whose derivatives decay faster than any inverse polynomial. It contains, for example, compactly supported smooth functions and Gaussian densities.

The Fourier transform is especially well behaved on \(\mathscr S(\mathbb R^d)\). The map \(\psi\mapsto\widetilde\psi\) is a continuous automorphism of \(\mathscr S(\mathbb R^d)\), with inverse \(\psi\mapsto\psi^\sharp\). Thus every Schwartz function has a Fourier transform, and every Schwartz function is the inverse Fourier transform of another Schwartz function. Continuity is understood with respect to the standard Schwartz topology, under which \(\mathscr S(\mathbb R^d)\) is a Fréchet space.\footnote{That is, a complete metrizable locally convex vector space. For a textbook construction of the topology, see treves2006topological.} Appendix (ref) provides additional details on this construction.

A tempered distribution is a continuous linear functional on \(\mathscr S(\mathbb R^d)\). I write $\mathscr S'(\mathbb R^d) \coloneq \bigl(\mathscr S(\mathbb R^d)\bigr)'$ for the space of tempered distributions. Critically, many ordinary functions define tempered distributions when their integration functionals are continuous on \(\mathscr S(\mathbb R^d)\). Specifically, if \(Q:\mathbb R^d\to\mathbb C\) is measurable and \[ \mathcal I_Q[\psi] \coloneq \int_{\mathbb R^d}Q(x)\psi(x)\,dx, \qquad \psi\in\mathscr S(\mathbb R^d), \] defines a continuous linear functional on \(\mathscr S(\mathbb R^d)\), then \(\mathcal I_Q\in\mathscr S'(\mathbb R^d)\).\footnote{Lemma (ref) gives a sufficient condition on $Q$ such that it defines a tempered distribution.} The space \(\mathscr S'(\mathbb R^d)\) also contains singular distributions that are not induced by ordinary functions. For example, evaluation at a point \(a\in\mathbb R^d\), \[ \operatorname{ev}_a[\psi]\coloneq \psi(a), \qquad \psi\in\mathscr S(\mathbb R^d), \] is a tempered distribution.

The Fourier transform extends from Schwartz functions to tempered distributions by duality. With the convention above, define operators \(\mathcal F,\mathcal F^{-1}:\mathscr S'(\mathbb R^d)\to \mathscr S'(\mathbb R^d)\) by \[ \mathcal F\{T\}[\psi] \coloneq T[\widetilde\psi], \qquad \mathcal F^{-1}\{T\}[\psi] \coloneq T[\psi^\sharp], \qquad \psi\in\mathscr S(\mathbb R^d). \] These definitions are well posed because the maps \(\psi\mapsto\widetilde\psi\) and \(\psi\mapsto\psi^\sharp\) map \(\mathscr S(\mathbb R^d)\) continuously to itself. Under the weak-\(\star\) topology on \(\mathscr S'(\mathbb R^d)\), the operators \(\mathcal F\) and \(\mathcal F^{-1}\) are continuous automorphisms; see Appendix (ref).

Existence of the Tweedie functional

This section develops the existence theory for Tweedie representations. In the additive-noise model, the problem reduces to the convolution operator induced by the noise. I characterize when this operator is invertible on its range and use the resulting inverse to equip the class of admissible observed densities with a natural topology.

To formalize the representation problem, let \(\lambda:\mathbb{R}^d\to\mathbb{C}\) be measurable and define \[ \Lambda_\lambda(P_X) \coloneq \int_{\mathbb{R}^d} \lambda(x)\,dP_X(x), \] whenever the integral is well defined. The question is whether this latent functional depends on \(P_X\) only through the observed marginal density \(f_Y\). Given a triple \((Y,X,V)\) satisfying Assumptions (ref)--(ref), I identify a class of integrands \(\lambda\) for which there exists a unique linear functional \(\mathcal{T}_{\lambda,V}\) satisfying $\Lambda_\lambda(P_X)=\mathcal{T}_{\lambda,V}[f_Y]$. For these integrands, the latent integral is therefore representable in terms of the observed density alone.

To state this result, let $\mathcal{M}(\mathbb{R}^d)$ denote the space of finite signed Radon measures on $\mathbb{R}^d$. For $\mu\in\mathcal{M}(\mathbb{R}^d)$, define its total variation norm by \[ \|\mu\|_{TV} \coloneq \sup\left\{ \left|\int_{\mathbb{R}^d} h(x)\,d\mu(x)\right| : h\in C_0(\mathbb{R}^d),\ \|h\|_\infty\le 1 \right\}. \] It is well known that $\bigl(\mathcal{M}(\mathbb{R}^d),\|\cdot\|_{TV}\bigr)$ is a Banach space rudin1987real. Since $\mathcal{P}(\mathbb{R}^d)\subseteq \mathcal{M}(\mathbb{R}^d)$, the results below apply in particular to probability measures. I nevertheless work on $\mathcal{M}(\mathbb{R}^d)$ because its vector space structure will be useful throughout.

Given the noise law $P_V$ with density $f_V$, define the convolution operator $\mathcal{K}_V:\mathcal{M}(\mathbb{R}^d)\to L^1(\mathbb{R}^d;\mathbb{C})$ by \[ \mathcal{K}_V[\mu](y) \coloneq (f_V * \mu)(y) = \int_{\mathbb{R}^d} f_V(y-x)\,d\mu(x), \qquad y\in\mathbb{R}^d. \] The operator $\mathcal{K}_V$ is linear and continuous from $\bigl(\mathcal{M}(\mathbb{R}^d),\|\cdot\|_{TV}\bigr)$ into $\bigl(L^1(\mathbb{R}^d;\mathbb{C}),\|\cdot\|_1\bigr)$. Indeed, it is immediate from the defintion that $\|\mathcal{K}_V[\mu]\|_1 \le \|\mu\|_{TV}$.

The range of $\mathcal{K}_V$ consists of all $L^1$ functions obtainable by convolving the noise density $f_V$ against a finite signed latent measure. I call this range the class of $V$-admissible mixtures and write \[ \mathcal{A}_V(\mathbb{R}^d) \coloneq \{f=\mathcal{K}_V[\mu]=f_V * \mu : \mu\in\mathcal{M}(\mathbb{R}^d)\} \subseteq L^1(\mathbb{R}^d;\mathbb{C}). \] By Proposition (ref), whenever the observed density $f_Y$ exists, it belongs to $\mathcal{A}_V(\mathbb{R}^d)$.

I next establish that $\mathcal{K}_V$ is a bijection between latent measures and $V$-admissible mixtures. Surjectivity is immediate from the definition of $\mathcal{A}_V(\mathbb{R}^d)$. The remaining issue is injectivity. Assumption (ref) provides it by ruling out distinct latent measures that generate the same mixture.

propositionUnder Assumption (ref), the operator $\mathcal{K}_V$ is a bijection between $\mathcal{M}(\mathbb{R}^d)$ and $\mathcal{A}_V(\mathbb{R}^d)$.

Proposition (ref) implies the inverse $\mathcal{K}_V^{-1} : \mathcal{A}_V(\mathbb{R}^d) \to \mathcal{M}(\mathbb{R}^d)$ well defined. I use this inverse to transfer the total variation norm from $\mathcal{M}(\mathbb{R}^d)$ to $\mathcal{A}_V(\mathbb{R}^d)$. Specifically, if $f \in \mathcal{A}_V(\mathbb{R}^d)$ and $f = f_V \ast \mu$ for some $\mu \in \mathcal{M}(\mathbb{R}^d)$, define \[ \|f\|_{\mathcal{A}_V} \coloneq \|\mathcal{K}_V^{-1}[f]\|_{TV} = \|\mu\|_{TV}. \]

This definition is unambiguous because Proposition (ref) gives a unique representing measure $\mu$. Furthermore, with this norm, $\mathcal{K}_V : \bigl(\mathcal{M}(\mathbb{R}^d),\|\cdot\|_{TV}\bigr) \to \bigl(\mathcal{A}_V(\mathbb{R}^d),\|\cdot\|_{\mathcal{A}_V}\bigr)$ is an isometric isomorphism. Hence $\mathcal{A}_V(\mathbb{R}^d)$ inherits the Banach space structure of $\mathcal{M}(\mathbb{R}^d)$. Let $\mathcal{A}_V'(\mathbb{R}^d) \coloneq \bigl(\mathcal{A}_V(\mathbb{R}^d)\bigr)'$ denote the Banach dual of $\mathcal{A}_V(\mathbb{R}^d)$, that is, the space of continuous linear functionals on $\mathcal{A}_V(\mathbb{R}^d)$. The next theorem shows that the density-based representation maps sought here are precisely elements of this dual space.

theorem[Existence and uniqueness of the representing functional] Let $(Y,X,V)$ satisfy Assumptions (ref) and (ref). For every $\lambda\in C_0(\mathbb{R}^d;\mathbb{C})$, there exists a unique functional $\mathcal{T}_{\lambda,V}\in \mathcal{A}'_V(\mathbb{R}^d)$ such that \[ \mathcal{T}_{\lambda,V}\!\left[\mathcal{K}_V[\mu]\right] = \int_{\mathbb{R}^d}\lambda(x)\,d\mu(x) \qquad \text{for all }\mu\in\mathcal{M}(\mathbb{R}^d). \] In particular, under Assumption (ref), $f_Y = \mathcal{K}_V[P_X]$, and so \[ \mathcal{T}_{\lambda,V}[f_Y] = \int_{\mathbb{R}^d} \lambda(x)\, dP_X(x). \]

Theorem (ref) establishes existence and uniqueness of the representation map. Specifically, for any $\lambda\in C_0(\mathbb{R}^d)$, the theorem gives a unique continuous linear functional $\mathcal{T}_{\lambda,V}\in \mathcal{A}_V'(\mathbb{R}^d)$ such that $\mathcal{T}_{\lambda,V}[\mathcal{K}_V[\mu]] = \int_{\mathbb{R}^d}\lambda(x)\,d\mu(x)$ for every $\mu\in\mathcal{M}(\mathbb{R}^d)$. Applying this identity to the observed density $f_Y=\mathcal{K}_V[P_X]$ shows that the latent integral $\Lambda_\lambda(P_X)$ can be recovered as the value of a continuous linear functional of $f_Y$.

An immediate consequence of Theorem (ref) is the existence and uniqueness of the Tweedie representation. The key point is that, after applying Bayes' formula to $\theta_V^g$, the numerator is an integral of a known function with respect to the latent law $P_X$. Indeed, under Assumptions (ref) and (ref), for any $y$ with $f_Y(y)>0$,

equation[equation omitted — 209 chars of source]

Thus, whenever $\lambda_{g,V,y}\in C_0(\mathbb{R}^d)$, the numerator in (ref) falls under Theorem (ref). This gives the following corollary.

corollary[Existence and uniqueness of the Tweedie functional] Fix $y\in\mathbb R^d$ and let $g:\mathbb R^d\to\mathbb R$ be measurable. Under assumptions (ref), (ref), and (ref), if $\lambda_{g,V,y}(x)=g(x)f_V(y-x) \in C_0(\mathbb{R}^d;\mathbb{C})$, then there exists a unique $\mathcal T_{g,V,y}\in \mathcal A_V'(\mathbb R^d)$ such that, for every $f_Y=\mathcal K_V[P_X]\in\mathcal A_V$ with $f_Y(y)>0$, the posterior expectation of $g(X)$ at $Y=y$ is finite and satisfies \[ \theta_V^g(y)= \mathbb{E}[g(X) \mid Y = y] = \frac{\mathcal{T}_{g,V,y}[f_Y]}{f_Y(y)} \]
remarkThe representation is defined only at points $y$ such that $f_Y(y)>0$. This restriction is harmless for most statements about $Y$, since the representation is defined $P_Y$-almost surely: \[ P_Y(\{y:f_Y(y)=0\}) = \int_{\{y:f_Y(y)=0\}} f_Y(y)\,dy = 0. \]
remarkA limitation of the preceding corollary is that it applies only when the integrand \(\lambda_{g,V,y}\) is continuous and vanishes at infinity. This excludes important cases, such as indicator functions \(g\), which are needed to represent posterior cumulative distribution functions. Section (ref) later removes this restriction by showing that the corresponding Tweedie functionals arise as limits of Tweedie functionals associated with continuous approximations to \(g\).
remarkCorollary (ref) provides a pointwise identity for each fixed \(y\in\mathbb{R}^d\). In many applications, however, the object of interest is the conditional mean as a function of the observation, namely \(\mathbb{E}[g(X)\mid Y]\) viewed as an integrable \(\sigma(Y)\)-measurable random variable. Proposition (ref) in Appendix (ref) gives an operator-level version of the pointwise representation. Concretely, it establishes under further regularity that the pointwise maps \(\mathcal{T}_{g,V,y}\) are obtained by evaluating a single linear operator \(\mathscr{T}_{g,V}\): \[ \mathscr{T}_{g,V}[f_Y](y) = \mathcal{T}_{g,V,y}[f_Y] \qquad \text{for every } y \text{ with } f_Y(y)>0. \] Consequently, the conditional mean admits the global representation \[ \mathbb{E}[g(X)\mid Y] = \frac{\mathscr{T}_{g,V}[f_Y](Y)}{f_Y(Y)} \qquad \mathbb{P}\text{-almost surely}. \] I call \(\mathscr{T}_{g,V}\) the Tweedie operator. This formulation separates the dependence on the observed density from the dependence on the evaluation point: the density \(f_Y\) is first mapped to the function \(\mathscr{T}_{g,V}[f_Y]\), and the pointwise numerator is then obtained by evaluation. Equivalently, it factors the Tweedie functional as $\mathcal{T}_{g,V,y}=\operatorname{ev}_y \;\circ \; \mathscr{T}_{g,V}$. Because the image of \(\mathscr{T}_{g,V}\) can be characterized, this operator-level representation provides a way to study global regularity properties of the posterior mean. In particular, Corollary (ref) uses this structure to establish differentiability of \(\mathbb{E}[g(X)\mid Y]\). A related operator-level formulation for posterior means appears in raphan2011least.

Computing Tweedie representations

For each \(y\in\mathbb R^d\), Section (ref) identified the Tweedie functional as the unique linear map \(\mathcal T_{g,V,y}\) satisfying

\[ \theta_V^g(y) = \frac{1}{f_Y(y)}\,\mathcal T_{g,V,y}[f_Y], \qquad \mathcal T_{g,V,y}[f_Y] = \int_{\mathbb R^d}\lambda_{g,V,y}(x)\,dP_X(x), \qquad \mathcal T_{g,V,y}\in\mathcal A'_V(\mathbb R^d). \]

This section turns that representation into a computational method. The method has three steps. First, associate the posterior functional of interest with a tempered distribution. Second, compute its inverse Fourier transform in the sense of distributions. Third, use the resulting expression to define and extend the corresponding linear functional on the space of $V-$admissible mixtures. The second step is central. Once the relevant tempered distribution is identified in step one, deriving the Tweedie representation becomes a problem in Fourier calculus. This reduction makes the method systematic and applicable across a broad class of posterior functionals.

To carry out this strategy, however, I need to impose additional regularity on the noise density. This regularity is inherited by \(V\)-mixtures: if the noise density has \(k\) bounded continuous derivatives that vanish at infinity, then every mixture generated by this noise has the same degree of smoothness. This assumption seems natural in the present problem. Indeed, in the Gaussian case, Tweedie's formula already requires differentiability of the observed marginal density.

The smoothness assumption lets us work in an ambient Banach space whose topology is weaker than the original \(\mathcal A_V\)-topology used for existence. Crucially, Schwartz functions are dense in this weaker topology. Since the Schwartz space is the natural setting for Fourier inversion of tempered distributions, we can perform the deconvolution and inverse Fourier calculation there, and then extend the resulting functional continuously to all \(V\)-mixtures. This is how the abstract functional from Section (ref) becomes computable.

I now make this construction explicit. For \(k\in\mathbb N_0\), define \[ C_0^k(\mathbb R^d;\mathbb C) \coloneq \left\{ \psi\in C^k(\mathbb R^d;\mathbb C) : \partial^\alpha \psi\in C_0(\mathbb R^d;\mathbb C) \text{ for all } |\alpha|\le k \right\}, \] Thus \(C_0^k(\mathbb R^d;\mathbb C)\) is the space of \(k\)-times continuously differentiable functions whose derivatives up to order \(k\) vanish at infinity. Now define \[ \Xi^k(\mathbb R^d;\mathbb C) \coloneq C_0^k(\mathbb R^d;\mathbb C)\cap L^1(\mathbb R^d;\mathbb C), \qquad \|\psi\|_{\Xi^k} \coloneq \max_{|\alpha|\le k}\|\partial^\alpha \psi\|_\infty+\|\psi\|_1. \]

assumption[Noise smoothness of order \(k\)] For some \(k\in\mathbb N_0\), the noise density \(f_V\) belongs to \(C_0^k(\mathbb R^d;\mathbb R)\).

Assumption (ref) says that the noise density has continuous derivatives up to order \(k\), all of which vanish at infinity. Since \(f_V\) is a density, the assumption naturally implies that \(f_V\in\Xi^k(\mathbb R^d;\mathbb R)\). This assumption encompasses many common noise laws, for suitable choices of \(k \geq 0\), such as Gaussian, logistic, Gumbel, Laplace, and Cauchy distributions.

The next lemma records the basic facts about this space used to establish the main theorem. The space \(\Xi^k\) is a Banach space; the space of \(V\)-mixtures embeds continuously into \(\Xi^k\); and Schwartz functions are dense in \(\Xi^k\).

lemma[Mixture smoothness and approximation] Assume \(f_V\) satisfies Assumption (ref) with smoothness order \(k\). Then: \begin{itemize} • Completeness. \((\Xi^k(\mathbb R^d;\mathbb C),\|\cdot\|_{\Xi^k})\) is a Banach space. • \(V\)-mixtures inherit noise regularity. The inclusion $\mathcal A_V(\mathbb R^d)\hookrightarrow \Xi^k(\mathbb R^d;\mathbb C)$ is continuous. • Schwartz approximation. The inclusion $\mathscr S(\mathbb R^d)\hookrightarrow \Xi^k(\mathbb R^d;\mathbb C)$ is continuous, and \(\mathscr S(\mathbb R^d)\) is dense in \((\Xi^k(\mathbb R^d;\mathbb C),\|\cdot\|_{\Xi^k})\). \end{itemize}

A general strategy for computing Tweedie representations

I now state the main result of this section, and the main computational result of the paper.

theorem[Computation of Tweedie functionals] Fix \(y\in\mathbb R^d\) and let \(g:\mathbb R^d\to\mathbb R\) be measurable. Suppose that \((Y,X,V)\) satisfies Assumptions (ref), (ref), (ref), and (ref) for some \(k\ge 0\). Let $\lambda_{g,V,y}(x)\coloneq g(x)f_V(y-x)$, $\lambda_{g,V,y}\in \Xi^0(\mathbb R^d;\mathbb C)$, $\widetilde{\lambda}_{g,V,y}\in L^1(\mathbb R^d;\mathbb C)$, and define \[ \mathcal Q_{g,V,y}(\omega)\coloneq \begin{cases} \dfrac{\widetilde{\lambda}_{g,V,y}(\omega)}{\varphi_V(-\omega)}, & \varphi_V(-\omega)\neq 0,\\[1ex] 0, & \text{otherwise}. \end{cases} \] Assume further that \begin{enumerate} • Temperedness. There exists \(N\in\mathbb N_0\) such that \[ \int_{\mathbb R^d} \frac{|\mathcal Q_{g,V,y}(\omega)|}{(1+\|\omega\|_2)^N}\,d\omega <\infty. \] Thus \(\mathcal I_{\mathcal Q_{g,V,y}}[\psi] \coloneq \int_{\mathbb R^d}\mathcal Q_{g,V,y}(\omega)\psi(\omega)\,d\omega\) defines a tempered distribution on \(\mathscr S(\mathbb R^d)\). • \(\Xi^k\)-continuity. There exists \(C>0\) such that, for every \(\psi\in\mathscr S(\mathbb R^d)\), \[ \bigl|\mathcal F^{-1}\{\mathcal I_{\mathcal Q_{g,V,y}}\}[\psi]\bigr| \le C\|\psi\|_{\Xi^k}. \] \end{enumerate} Then \(\mathcal F^{-1}\{\mathcal I_{\mathcal Q_{g,V,y}}\}\) extends uniquely to a continuous linear functional on \(\Xi^k(\mathbb R^d;\mathbb C)\), and its restriction to \(\mathcal A_V(\mathbb R^d)\) is the Tweedie functional \(\mathcal T_{g,V,y}\). Equivalently, for every \(f\in \mathcal A_V(\mathbb R^d)\) and every sequence \((f_n)_{n=1}^\infty\subset \mathscr S(\mathbb R^d)\) with \(\|f_n-f\|_{\Xi^k}\to 0\), \[ \mathcal T_{g,V,y}[f] = \lim_{n\to\infty} \mathcal F^{-1}\{\mathcal I_{\mathcal Q_{g,V,y}}\}[f_n], \] and the limit is independent of the approximating sequence.

Theorem (ref) turns the abstract existence of the Tweedie functional into a closed-form formula. Relative to the existence result, the theorem adds the Fourier-inversion conditions \(\lambda_{g,V,y},\widetilde{\lambda}_{g,V,y}\in L^1(\mathbb R^d;\mathbb C)\), which need to be checked case by case. For example, under Gaussian noise, the exponential decay of the noise density and its Fourier transform usually allows \(g\) to have polynomial growth of arbitrary degree.

The construction then proceeds by identifying a complex-valued function \(\mathcal Q_{g,V,y}\), which I call the Fourier-domain representer. The key insight is to treat \(\mathcal Q_{g,V,y}\), when possible, as the integration-defined tempered distribution \(\mathcal I_{\mathcal Q_{g,V,y}}\), rather than only as a function. Its inverse Fourier transform is then well defined in the sense of tempered distributions and is again a linear functional, now on the Schwartz space. This is where the ambient space \(\Xi^k(\mathbb R^d;\mathbb C)\) enters: if the resulting linear functional is continuous with respect to the \(\Xi^k\)-norm, then it extends continuously to \(\Xi^k(\mathbb R^d;\mathbb C)\), and its restriction to the mixture space is the Tweedie functional. Thus \(\mathcal T_{g,V,y}\) can be computed by evaluating the inverse Fourier transform of \(\mathcal I_{\mathcal Q_{g,V,y}}\) on Schwartz approximations to the observed marginal density and passing to the \(\Xi^k\)-limit.

Pragmatically, the theorem gives the following three-step strategy:

enumerate[label=Step \arabic*., leftmargin=*, align=left] • Compute the Fourier-domain representer \(\mathcal Q_{g,V,y}(\omega)\) and verify it defines a tempered distribution by integration. • Compute the inverse Fourier transform \(\mathcal F^{-1}\{\mathcal I_{\mathcal Q_{g,V,y}}\}\) on Schwartz space \(\mathscr S(\mathbb R^d)\). • Prove \(\Xi^k\)-continuity and compute the Tweedie functional evaluated at a $V-$mixture by taking the \(\Xi^k\)-limit along a Schwartz approximating sequence.

I now apply this strategy to recover the classical Tweedie formula.

\paragraph{Example: Original Tweedie's formula.} Suppose \(d=1\), \(Y=X+V\), \(X\perp\!\!\!\!\perp V\), and \(V\sim \mathcal N(0,\sigma^2)\). Let \(g(x)=x\). Since the Gaussian density belongs to \(C_0^\infty(\mathbb R)\), the smoothness assumption holds for every \(k\); I take \(k=1\). Fix \(y\in\mathbb R\). The first step is to compute the Fourier-domain representer. Here \[ \lambda_{g,V,y}(x)=x f_V(y-x), \qquad \widetilde{\lambda}_{g,V,y}(\omega) = \bigl(y+i\sigma^2\omega\bigr)e^{i\omega y}e^{-\sigma^2\omega^2/2}. \] Since \(\varphi_V(-\omega)=e^{-\sigma^2\omega^2/2}\), $\mathcal Q_{g,V,y}(\omega) = \bigl(y+i\sigma^2\omega\bigr)e^{i\omega y}$.

The second step is to compute the inverse Fourier transform of the integration functional induced by \(\mathcal Q_{g,V,y}\). Since \(\mathcal Q_{g,V,y}\) grows at most linearly, it defines a tempered distribution by integration. By the standard transform identities kammler2007first, \[ \mathcal F^{-1}\{\mathcal I_{\mathcal Q_{g,V,y}}\}[\psi] = y\psi(y)+\sigma^2\psi'(y), \qquad \psi\in\mathscr S(\mathbb R). \]

The third step is to verify \(\Xi^1\)-continuity. Indeed, \[ \left| y\psi(y)+\sigma^2\psi'(y) \right| \le \bigl(|y|+\sigma^2\bigr)\|\psi\|_{\Xi^1}. \] Thus the functional extends continuously to \(\Xi^1(\mathbb R;\mathbb C)\). By Theorem (ref), for any sequence \((f_n)_{n\ge1}\subset\mathscr S(\mathbb R)\) with \(\|f_n-f_Y\|_{\Xi^1}\to0\), \[ \mathcal T_{g,V,y}[f_Y] = \lim_{n\to\infty} \left\{ y f_n(y)+\sigma^2 f_n'(y) \right\} = y f_Y(y)+\sigma^2 f_Y'(y), \] since \(f_n\to f_Y\) in \(\Xi^1\), implies both \(f_n(y)\to f_Y(y)\) and \(f_n'(y)\to f_Y'(y)\).\footnote{From this expression, we also observe that the Tweedie operator for the original Tweedie formula is \[(\mathscr{T}_{g,V})[f_Y](y)=y f_Y(y)+\sigma^2 f_Y'(y)\] } Dividing by \(f_Y(y)\), whenever \(f_Y(y)>0\), gives the classical Tweedie formula: \[ \mathbb E[X\mid Y=y] = \frac{\mathcal T_{g,V,y}[f_Y]}{f_Y(y)} = y+\sigma^2\frac{f_Y'(y)}{f_Y(y)}. \]

Tweedie formulas via approximation by continuous functions

Theorem (ref) gives a systematic way to compute Tweedie representations when the Fourier-domain representer \(\mathcal Q_{g,V,y}\) defines a tempered distribution. Some functionals fall outside this setting. For example, the function \(g\) defining the estimand may be discontinuous. In such cases, we can still apply the strategy by approximating the original problem with instances for which the tools apply. The next result gives one such approximation argument.

proposition[Tweedie functionals via approximation] Let $(Y,X,V)$ satisfy Assumptions (ref) and (ref). Fix \(y\in\mathbb R^d\) and a measurable function \(g:\mathbb R^d\to\mathbb R\). Let \((g_n)_{n\ge1}\) be a sequence of measurable functions. Define $\lambda_{n,g,V,y}(x)=g_n(x)f_V(y-x)$ and $\lambda_{g,V,y}(x)=g(x)f_V(y-x)$. Assume that $\lambda_{n,g,V,y} \in C_0(\mathbb{R}^d)$ for every $n \geq 1$, that $\lambda_{n,g,V,y}(x) \to \lambda_{g,V,y}(x)$ for all $x\in\mathbb R^d$, and that $\sup_{n\ge1}\|\lambda_{n,g,V,y}\|_\infty<\infty$. Then, for each \(n\ge1\), there exists a unique functional $\mathcal T_{n,g,V,y}\in \mathcal A_V'(\mathbb R^d)$ such that \[ \mathcal T_{n,g,V,y}[f_V*\mu] = \int_{\mathbb R^d}\lambda_{n,g,V,y}(x)\,d\mu(x), \qquad \mu\in\mathcal M(\mathbb R^d). \] Moreover, there exists a unique continuous linear functional $\mathcal T_{g,V,y}\in \mathcal A_V'(\mathbb R^d)$ such that \[ \mathcal T_{g,V,y}[f_V*\mu] = \int_{\mathbb R^d}\lambda_{g,V,y}(x)\,d\mu(x), \qquad \mu\in\mathcal M(\mathbb R^d). \] In fact, for every \(f\in\mathcal A_V(\mathbb R^d)\), this map is given as a pointwise limit \[ \mathcal T_{g,V,y}[f] = \lim_{n\to\infty} \mathcal T_{n,g,V,y}[f]. \] In particular, if \((Y,X,V)\) also satisfies Assumption (ref), then \(f_Y=f_V*P_X\), and whenever \(f_Y(y)>0\), \[ \mathbb E[g(X)\mid Y=y] = \frac{1}{f_Y(y)} \lim_{n\to\infty} \mathcal T_{n,g,V,y}[f_Y]. \]

The next example illustrates how this approximation argument produces a Tweedie representation beyond the direct scope of Theorem (ref)

\paragraph{Example -- Posterior distribution function under Gaussian noise.} Suppose \(d=1\), \(Y=X+V\), \(X\perp\!\!\!\!\perp V\), and \(V\sim \mathcal N(0,\sigma^2)\). We seek a Tweedie-type representation for the posterior distribution function, $\mathbb P[X\le x\mid Y=y]$, which corresponds to the measurable function $g_x(u)=\mathbbm 1\{u\le x\}$, $x\in\mathbb R$. For \(n\ge1\), define the smooth approximation \[ g_{n,x}(u) = \Phi\left(\frac{x+n^{-1/2}-u}{n^{-1}}\right), \qquad u\in\mathbb R. \] It is straightforward to check that \((g_{n,x})_{n\ge1}\) satisfies the conditions of Proposition (ref). Moreover, using Theorem (ref) to each \(g_{n,x}\) gives \[ \mathcal T_{n,g_x,V,y}[f] = \sum_{k=0}^{\infty} \frac{(-1)^k}{k!} \left( \frac{\sigma^2}{\sqrt{\sigma^2+n^{-2}}} \right)^k \Phi^{(k)} \left( \frac{x+n^{-1/2}-y}{\sqrt{\sigma^2+n^{-2}}} \right) f^{(k)}(y), \quad f \in \mathcal{A}_V(\mathbb{R}). \] where $\Phi$ is the cumulative distribution function of a standard normal variable. For each fixed \(n\), this series converges absolutely.\footnote{Details of this calculation are given in Lemma (ref).} Therefore,

align*[align* omitted — 361 chars of source]

To the best of my knowledge, this representation is new. It also has a simple interpretation. The first term, \(\Phi((x-y)/\sigma)\), is the posterior distribution function obtained under a flat, or locally uninformative, prior on \(X\). The remaining terms correct this flat-prior approximation by using information in the marginal density \(f_Y\) and its derivatives.

The formula also shows why posterior distribution functions are harder to recover than posterior means. In the Gaussian case, the posterior mean depends only on the first derivative of \(f_Y\), whereas the posterior cdf involves derivatives of \(f_Y\) of all orders. In special cases, moreover, the expression simplifies. For example, if \(X\sim\mathcal N(m,\tau^2)\), then $X\mid Y=y \sim \mathcal{N}\left( \frac{\sigma^2m+\tau^2y}{\sigma^2+\tau^2}, \frac{\sigma^2\tau^2}{\sigma^2+\tau^2} \right)$. In this case, the derivatives of \(f_Y\) are controlled, and the limit in \(n\) can be interchanged with the series. Mehler's formula mehler1866ueber then confirms the closed-form series expansion \[ \frac{1}{f_Y(y)} \sum_{k=0}^{\infty} \frac{(-1)^k \sigma^k}{k!} \Phi^{(k)} \left( \frac{x-y}{\sigma} \right) f_Y^{(k)}(y) = \Phi\left( \frac{ x-\frac{\sigma^2m+\tau^2y}{\sigma^2+\tau^2} }{ \sqrt{\frac{\sigma^2\tau^2}{\sigma^2+\tau^2}} } \right)=\mathbb P[X\le x\mid Y=y]. \]

Applications of Tweedie Calculus

This section illustrates the range of Tweedie calculus through four applications. First, I derive posterior mean formulas for several non-Gaussian additive-noise models. Second, I identify cases in which Tweedie functionals can be written as expectations of explicit functions of the observed variable. Third, I study posterior functionals under the conventional Laplace mechanism used in differential privacy. Fourth, I extend the calculus to the heteroskedastic Gaussian sequence model by conditioning on the potentially random noise covariance structure. The detailed calculations supporting these applications are relegated to Appendix (ref).

Posterior means beyond Gaussian noise

The first application concerns posterior means under additive noise laws other than the Gaussian. As noted before, the three-step strategy from Section (ref) can be applied case by case. Alternatively, the same tools can also identify entire families of Tweedie functionals from the structure of their Fourier-domain representers. The next proposition gives one such characterization. It generalizes the original Gaussian calculation by showing that certain factorizations of \(\mathcal Q_{g,V,y}\) yield weighted-derivative representations.

proposition[Tweedie representation via factorization of measures] Fix \(y\in\mathbb R^d\) and let \(g:\mathbb R^d\to\mathbb R\) be measurable. Suppose that \((Y,X,V)\) satisfies Assumptions (ref), (ref), (ref), and (ref) with smoothness order \(k\in\mathbb N_0\). Assume \[ \lambda_{g,V,y}(x)\coloneq g(x)f_V(y-x)\in \Xi^0(\mathbb R^d;\mathbb C), \qquad \widetilde{\lambda}_{g,V,y}\in L^1(\mathbb R^d;\mathbb C), \] Suppose there exists a family of finite signed Radon measures \(\{\mu_\alpha:|\alpha|\le k\}\subset \mathcal M(\mathbb R^d)\) such that \[ \mathcal Q_{g,V,y}(\omega) = \sum_{|\alpha|\le k}(i\omega)^\alpha\varphi_{\mu_\alpha}(\omega), \qquad \varphi_\mu(\omega) \coloneq \int_{\mathbb{R}^d} e^{i\omega^\top x}\,d\mu(x), \qquad\text{for Lebesgue-a.e. }\omega\in\mathbb R^d. \] Then, for every \(f\in\mathcal A_V(\mathbb R^d)\), \[ \mathcal T_{g,V,y}[f] = \sum_{|\alpha|\le k} \int_{\mathbb R^d}\partial^\alpha f(z)\,d\mu_\alpha(z). \] In particular, whenever \(f_Y(y)>0\), \[ \mathbb E[g(X)\mid Y=y] = \frac{1}{f_Y(y)} \sum_{|\alpha|\le k} \int_{\mathbb R^d}\partial^\alpha f_Y(z)\,d\mu_\alpha(z). \]

Proposition (ref) covers the case in which the Fourier-domain representer \(\mathcal Q_{g,V,y}\) factors into powers of \(\omega\) whose coefficients are Fourier--Stieltjes transforms of finitely many measures. In this case, the Tweedie functional is a finite linear combination of weighted derivatives of \(f_Y\). The original Tweedie formula is one instance of this factorization. We now turn to a non-Gaussian example and show that the same method applies just as naturally under Laplace noise.

\paragraph{Example: Posterior mean under Laplace noise.} Suppose \(d=1\), \(Y=X+V\), \(X\perp\!\!\!\!\perp V\), and \(V\) has Laplace density \[ f_V(u)=\frac{1}{2b}e^{-|u|/b}, \qquad b>0. \] Let \(g(x)=x\). Since \(f_V\in C_0(\mathbb R)\), the smoothness assumption holds with \(k=0\). In this setting, \[ \lambda_{g,V,y}(x)=x f_V(y-x), \qquad \widetilde{\lambda}_{g,V,y}(\omega) = e^{i\omega y} \left( \frac{y}{1+b^2\omega^2} + \frac{2ib^2\omega}{(1+b^2\omega^2)^2} \right). \] Since \(\varphi_V(-\omega)=(1+b^2\omega^2)^{-1}\), we obtain \[ \mathcal Q_{g,V,y}(\omega) = \left( y+\frac{2ib^2\omega}{1+b^2\omega^2} \right)e^{i\omega y}. \]

Now \(e^{i\omega y}=\varphi_{\delta_y}(\omega)\). Let \(\nu_y\) be the finite signed measure with density $d\nu_y(z) = \operatorname{sgn}(z-y)e^{-|z-y|/b}\,dz.$ A direct calculation then gives \[ \varphi_{\nu_y}(\omega) = \frac{2ib^2\omega}{1+b^2\omega^2}e^{i\omega y}. \] Thus \[ \mathcal Q_{g,V,y}(\omega) = y\,\varphi_{\delta_y}(\omega)+\varphi_{\nu_y}(\omega)=\varphi_{y\delta_y+\nu_y}(\omega). \] Proposition (ref) applies with \(k=0\) and \(\mu_0=y\delta_y+\nu_y\). Therefore, whenever \(f_Y(y)>0\), \[ \mathbb E[X\mid Y=y] = y+ \frac{1}{f_Y(y)} \int_{\mathbb R} \operatorname{sgn}(z-y)e^{-|z-y|/b}f_Y(z)\,dz. \]

Other families, such as Proposition (ref), are developed in Appendix (ref) using the same techniques. I conclude this subsection by illustrating the flexibility of the method across several noise distributions. Table (ref) reports posterior mean representations for these noise laws, with the corresponding derivations collected in Appendix (ref).

Several comments are in order. First, despite their analytic diversity, the resulting decision rules share a common shrinkage form. Each posterior mean is the observation, after correcting for the noise location, plus a normalized adjustment depending only on the observed marginal density \(f_Y\). Thus the unknown prior distribution enters only through \(f_Y\), and the adjustment can be interpreted as an empirical Bayes shrinkage term. The form of this shrinkage term, however, depends sharply on the noise law.

Second, the table shows that statistical difficulty is distribution-specific. For Gaussian noise, the correction is local: it depends only on \(f_Y(y)\) and \(f_Y'(y)\). This locality appears to be exceptional. The other entries involve nonlocal functionals of \(f_Y\).

Finally, the table illustrates the value of recasting the problem as Fourier calculus. The Cauchy entry is especially instructive. Its posterior mean involves the Hilbert transform of the observed density, a representation that would be hard to guess without the language of tempered distributions. In the present approach, it follows from the standard identity identifying the Hilbert transform with the inverse Fourier transform of the integration map defined by \(\omega\mapsto -i\operatorname{sgn}(\omega)\) grafakos2008classical.

landscape\begin{table}[ht] \scriptsize {3.5pt} \begin{threeparttable} \caption{Posterior mean Tweedie representations for several additive-noise models.} \begin{tabular}{@lcc>{$\displaystyle}l<{$}@} \toprule Noise law & $f_V(v)$ & $\varphi_V(\omega)$ & \multicolumn{1}{c@}{Tweedie representation of $\mathbb E[X\mid Y=y]$} \\ \midrule \multicolumn{4}{@l}{Two-sided laws on $\mathbb R$}\\[1pt] \begin{tabular}[t]{@l@} $\mathrm{Normal}(\mu,\sigma^2)$\\[-1pt] $\mu\in\mathbb R,\ \sigma>0$ \end{tabular} & $\displaystyle \frac{1}{\sqrt{2\pi}\sigma}\exp\!\left(-\frac{(v-\mu)^2}{2\sigma^2}\right)$ & $\displaystyle e^{i\mu\omega-\sigma^2\omega^2/2}$ & y-\mu+\sigma^2\frac{f_Y'(y)}{f_Y(y)} \\ \addlinespace[\smallskipamount] \begin{tabular}[t]{@l@} $\mathrm{Generalized\ Laplace}(\mu,b,\lambda)$\\[-1pt] $\mu\in\mathbb R,\ b>0,\ \lambda>\frac12$ \end{tabular} & $\displaystyle \frac{1}{\sqrt{\pi}\,\Gamma(\lambda)\,b}\Bigl(\frac{|v-\mu|}{2b}\Bigr)^{\lambda-\frac12} K_{\lambda-\frac12}\!\Bigl(\frac{|v-\mu|}{b}\Bigr)$ & $\displaystyle e^{i\mu\omega}(1+b^2\omega^2)^{-\lambda}$ & \begin{aligned}[t] y-\mu &+\frac{\lambda}{f_Y(y)} \left[ \int_{\mathbb R}\operatorname{sgn}(z-y)e^{-|z-y|/b}f_Y(z)\,dz \right] \end{aligned} \\ \addlinespace[\smallskipamount] \begin{tabular}[t]{@l@} $\mathrm{Laplace}(\mu,b)$\\[-1pt] $\mu\in\mathbb R,\ b>0$ \end{tabular} & $\displaystyle \frac{1}{2b}e^{-|v-\mu|/b}$ & $\displaystyle \frac{e^{i\mu\omega}}{1+b^2\omega^2}$ & \begin{aligned}[t] y-\mu &+\frac{1}{f_Y(y)} \left[ \int_{\mathbb R}\operatorname{sgn}(z-y)e^{-|z-y|/b}f_Y(z)\,dz \right] \end{aligned} \\ \addlinespace[\smallskipamount] \begin{tabular}[t]{@l@} $\mathrm{Asymmetric\ Laplace}(\mu,b_-,b_+)$\\[-1pt] $\mu\in\mathbb R,\ b_->0,\ b_+>0$ \end{tabular} & $\displaystyle \frac{1}{b_-+b_+} \left(\exp\!\Bigl(\frac{v-\mu}{b_-}\Bigr)\mathbf{1}_{\{v<\mu\}} +\exp\!\Bigl(-\frac{v-\mu}{b_+}\Bigr)\mathbf{1}_{\{v\ge \mu\}}\right)$ & $\displaystyle \frac{e^{i\mu\omega}}{(1+i b_-\omega)(1-i b_+\omega)}$ & y-\mu+ \frac{1}{f_Y(y)} \left[ \int_y^\infty e^{-(z-y)/b_-}f_Y(z)\,dz - \int_{-\infty}^y e^{(z-y)/b_+}f_Y(z)\,dz \right] \\ \addlinespace[\smallskipamount] \begin{tabular}[t]{@l@} $\mathrm{Logistic}(\mu,s)$\\[-1pt] $\mu\in\mathbb R,\ s>0$ \end{tabular} & $\displaystyle \frac{e^{-(v-\mu)/s}}{s\bigl(1+e^{-(v-\mu)/s}\bigr)^2}$ & $\displaystyle e^{i\mu\omega}\frac{\pi s\omega}{\sinh(\pi s\omega)}$ & \begin{aligned}[t] y-\mu &+\frac{1}{f_Y(y)} \left[ \int_0^\infty \frac{f_Y(y+t)-f_Y(y-t)}{e^{t/s}-1}\,dt \right] \end{aligned} \\ \addlinespace[\smallskipamount] \begin{tabular}[t]{@l@} $\mathrm{Gumbel}(\mu,\beta)$\\[-1pt] $\mu\in\mathbb R,\ \beta>0$ \end{tabular} & $\displaystyle \frac{1}{\beta}\exp\!\left[-\frac{v-\mu}{\beta}-e^{-(v-\mu)/\beta}\right]$ & $\displaystyle e^{i\mu\omega}\Gamma(1-i\beta\omega)$ & \begin{aligned}[t] y-\mu-\beta\gamma_E &+\frac{1}{f_Y(y)} \left[ \int_0^\infty \frac{f_Y(y)-f_Y(y-u)}{e^{u/\beta}-1}\,du \right] \end{aligned} \\ \addlinespace[\smallskipamount] \begin{tabular}[t]{@l@} $\mathrm{Cauchy}(\mu,\gamma)$\\[-1pt] $\mu\in\mathbb R,\ \gamma>0$ \end{tabular} & $\displaystyle \frac{1}{\pi}\frac{\gamma}{(v-\mu)^2+\gamma^2}$ & $\displaystyle e^{i\mu\omega-\gamma|\omega|}$ & y-\mu-\gamma\,\frac{\mathcal H[f_Y](y)}{f_Y(y)} \\ \addlinespace[\smallskipamount] \begin{tabular}[t]{@l@} $\mathrm{Hyperbolic\ Secant}(\mu,s)$\\[-1pt] $\mu\in\mathbb R,\ s>0$ \end{tabular} & $\displaystyle \frac{1}{2s}\operatorname{sech}\!\left(\frac{\pi(v-\mu)}{2s}\right)$ & $\displaystyle e^{i\mu\omega}\operatorname{sech}(s\omega)$ & \begin{aligned}[t] y-\mu &+\frac{1}{f_Y(y)} \left[ \int_0^\infty \frac{f_Y(y+t)-f_Y(y-t)} {2\sinh\!\left(\frac{\pi t}{2s}\right)}\,dt \right] \end{aligned} \\ \addlinespace[\smallskipamount] \addlinespace[2pt] \multicolumn{4}{@l}{One-sided laws on $(0,\infty)$}\\[1pt] \begin{tabular}[t]{@l@} $\mathrm{Gamma}(\alpha,\theta)$\\[-1pt] $\alpha>1,\ \theta>0$ \end{tabular} & $\displaystyle \frac{1}{\Gamma(\alpha)\theta^\alpha}v^{\alpha-1}e^{-v/\theta}\mathbf{1}_{\{v>0\}}$ & $\displaystyle (1-i\theta\omega)^{-\alpha}$ & \begin{aligned}[t] y &-\frac{\alpha}{f_Y(y)} \left[ \int_{-\infty}^y e^{(z-y)/\theta}f_Y(z)\,dz \right] \end{aligned} \\ \addlinespace[\smallskipamount] \begin{tabular}[t]{@l@} $\mathrm{Noncentral}\ \chi^2_\nu(\delta)$\\[-1pt] $\nu>2,\ \delta\ge 0$ \end{tabular} & $\displaystyle \frac{1}{2}e^{-(v+\delta)/2} \left(\frac{v}{\delta}\right)^{\nu/4-1/2} I_{\nu/2-1}\!\bigl(\sqrt{\delta v}\bigr)\mathbf{1}_{\{v>0\}}$ & $\displaystyle (1-2i\omega)^{-\nu/2} \exp\!\Bigl(\frac{i\delta\omega}{1-2i\omega}\Bigr)$ & y- \frac{1}{f_Y(y)} \left[ \frac{\nu}{2}\int_{-\infty}^y e^{(z-y)/2}f_Y(z)\,dz + \frac{\delta}{4}\int_{-\infty}^y (y-z)e^{(z-y)/2}f_Y(z)\,dz \right] \\[2pt] \addlinespace[\smallskipamount] \begin{tabular}[t]{@l@} $\mathrm{Inverse\ Gaussian}(\mu,\lambda)$\\[-1pt] $\mu>0,\ \lambda>0$ \end{tabular} & $\displaystyle \left(\frac{\lambda}{2\pi v^3}\right)^{1/2} \exp\!\left(-\frac{\lambda(v-\mu)^2}{2\mu^2v}\right)\mathbf{1}_{\{v>0\}}$ & $\displaystyle \exp\!\left[\frac{\lambda}{\mu}\left(1-\sqrt{1-\frac{2i\mu^2\omega}{\lambda}}\right)\right]$ & y- \frac{1}{f_Y(y)} \left[ \sqrt{\frac{\lambda}{2\pi}} \int_{-\infty}^y (y-z)^{-1/2} \exp\!\left(-\frac{\lambda(y-z)}{2\mu^2}\right) f_Y(z)\,dz \right] \\[2pt] \bottomrule \end{tabular} \begin{tablenotes}[flushleft] • Notes: Parameter restrictions are displayed below each distribution and indicate the conditions under which the displayed Tweedie representation is derived in the appendix. Here $\Gamma$ denotes the Euler gamma function, $I_\nu$ and $K_\nu$ the modified Bessel functions of the first and second kind, respectively, $\mathcal H$ the Hilbert transform, $\gamma_E$ the Euler--Mascheroni constant, $\operatorname{sgn}(x)$ the sign function, and $\operatorname{sech}(x)=2/(e^x+e^{-x})$. \end{tablenotes} \end{threeparttable} \end{table}

Tweedie functionals given as an expectation

This subsection considers Tweedie representations when the Fourier-domain representer is integrable: $\mathcal Q_{g,V,y}\in L^1(\mathbb R^d;\mathbb C).$ In this regime, the Tweedie functional takes the simple form of an expectation of the inverse Fourier transform of \(\mathcal Q_{g,V,y}\).

This expectation representation connects the present framework to the classical literature on unbiased estimation of functions of latent variables under additive noise. In the Gaussian case, this problem goes back to kolmogorov1950unbiased. A key feature of the present result is that the representation is obtained in closed form from the Fourier-domain representer. Other contributions, notably stefanski1989unbiased and voinov2012unbiased, derive related formulas through power-series expansions of analytic functions.

proposition[Tweedie representation via expectations] Fix \(y\in\mathbb R^d\) and let \(g:\mathbb R^d\to\mathbb R\) be measurable. Suppose that \((Y,X,V)\) satisfies Assumptions (ref), (ref), and (ref). Assume \[ \lambda_{g,V,y}(x)\coloneq g(x)f_V(y-x)\in \Xi^0(\mathbb R^d;\mathbb C), \qquad \widetilde{\lambda}_{g,V,y}\in L^1(\mathbb R^d;\mathbb C), \] and let \(\mathcal Q_{g,V,y}\) be the Fourier-domain representer. If $\mathcal Q_{g,V,y}\in L^1(\mathbb R^d;\mathbb C)$, then for every \(f\in\mathcal A_V(\mathbb R^d)\), \[ \mathcal T_{g,V,y}[f] = \int_{\mathbb R^d}\mathcal Q_{g,V,y}^{\sharp}(z)f(z)\,dz, \] where \(Q_{g,V,y}^\sharp\) is real-valued. In particular, whenever \(f_Y(y)>0\), \[ \mathbb E[g(X)\mid Y=y] =\frac{1 }{f_Y(y)}\left[\int_{\mathbb R^d}\mathcal Q_{g,V,y}^{\sharp}(z)f_Y(z)\,dz\right]= \frac{1 }{f_Y(y)}\mathbb E\!\left[\mathcal Q_{g,V,y}^{\sharp}(Y)\right]. \]

Proposition (ref) shows that, whenever the inverse Fourier image \(\mathcal Q_{g,V,y}^{\sharp}\) is sufficiently regular, the Tweedie functional can be represented as an expectation of an explicit function of the observed variable. This is especially useful in practice, since it yields unbiased estimators directly from the observed sample. The following example illustrates this mechanism.

\paragraph{Example: Tweedie functional as an expectation.} Suppose \(Y=X+V\), \(X\perp\!\!\!\!\perp V\), and \(V\sim\mathcal N(0,1)\). For \(a>1\), let \[ g(x) = \frac{1}{a} \exp\!\left(\frac{a^2-1}{2a^2}x^2\right). \] Since \(f_V(u)=(2\pi)^{-1/2}e^{-u^2/2}\), a direct calculation gives \[ \lambda_{g,V,y}(x) = e^{\frac{(a^2-1)y^2}{2}} \frac{1}{\sqrt{2\pi}\,a} \exp\!\left(-\frac{(x-a^2y)^2}{2a^2}\right),\quad \widetilde{\lambda}_{g,V,y}(\omega) = \exp\!\left( \frac{(a^2-1)y^2}{2} +i\omega a^2 y -\frac{a^2\omega^2}{2} \right). \]

and hence, \[ \mathcal Q_{g,V,y}(\omega) = \exp\!\left( \frac{(a^2-1)y^2}{2} +i\omega a^2 y -\frac{(a^2-1)\omega^2}{2} \right). \] Thus \(\mathcal Q_{g,V,y}\in L^1(\mathbb R)\), and Proposition (ref) yields \[ \mathcal T_{g,V,y}[f_Y] = \mathbb E\!\left[\mathcal{Q}_{g,V,y}^{\sharp}(Y)\right], \quad \mathcal{Q}_{g,V,y}^{\sharp}(z)= e^{\frac{(a^2-1)y^2}{2}} \frac{1}{\sqrt{2\pi(a^2-1)}} \exp\!\left(-\frac{(z-a^2y)^2}{2(a^2-1)}\right). \]

Functionals under the conventional Laplace mechanism

Differential privacy is a standard framework for protecting individual privacy in statistical and machine-learning applications. Its canonical additive-noise mechanism is the Laplace mechanism. Given a vector-valued query, such as Census counts across geographic or demographic cells, the released value is obtained by adding independent one-dimensional Laplace noise to each coordinate. This is, in particular, the conventional Laplace mechanism of dwork2014algorithmic, which gives the standard differential privacy guarantee in dwork2014algorithmic. This subsection studies Tweedie representations for posterior functionals after such a release.

This subsection studies Tweedie representations for posterior functionals after such a release. Motivated by the Laplace mechanism, consider the model \[ Y=X+V, \qquad X\perp\!\!\!\!\perp V, \qquad V=(V_1,\ldots,V_d), \qquad V_j\overset{\mathrm{i.i.d.}}{\sim}\operatorname{Lap}(0,b). \]

Table (ref) collects several Tweedie representations for the product-Laplace model in dimension one. Appendix (ref) provides the derivations and extends the posterior mean, posterior covariance matrix, and posterior moment generating function formulas to arbitrary dimensions.

Tweedie representations in the heteroskedastic Gaussian sequence model

The final application considers the heteroskedastic Gaussian sequence model. Its distinctive feature is that the noise covariance varies across units and may be statistically related to the latent parameter. This setting arises naturally in empirical Bayes problems with many noisy estimates, especially in economics and related social sciences.\footnote{Examples include value-added modeling einav2025producing, the evaluation of place-based policies chetty2018impacts, the measurement of discrimination kline2022systemic, policy targeting moon2025optimal, and research production abadie2026estimatingvalueevidencebaseddecision.} A common interpretation in these settings is that \(Y\) is an estimator of a latent parameter \(X\), constructed from a sample mean or, more generally, from an asymptotically linear statistic. A central limit theorem then motivates a Gaussian approximation to the estimation error. Heteroskedasticity arises when sample sizes, sampling designs, or underlying population variability differ across units, making the covariance matrix unit-specific. It is therefore natural to model \((X,\Sigma)\) as drawn from a joint superpopulation distribution \(P_{X,\Sigma}\).

\newcolumntype{L}[1]{>{\arraybackslash}m{#1}} \newcolumntype{C}[1]{>{\arraybackslash}m{#1}}

landscape\begin{table}[ht] \scriptsize {4pt} \begin{threeparttable} \caption{Examples of Tweedie representations under Laplace noise, \(d=1\).} \begin{tabular}{@L{0.18\linewidth}C{0.18\linewidth}C{0.60\linewidth}@} \toprule Name of functional & \(g(x)\) & \(\mathbb E[g(X)\mid Y=y]\) \\ \midrule \multicolumn{3}{@l}{Panel A. Tweedie representations for conditional expectations}\\ Posterior mean & \(\displaystyle x\) & \makebox[\linewidth][c]{\(\displaystyle y+ \frac{1}{f_Y(y)} \int_{\mathbb R} \operatorname{sgn}(u)e^{-|u|/b}f_Y(y+u)\,du \)} \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior second moment & \(\displaystyle x^2\) & \makebox[\linewidth][c]{\(\displaystyle y^2 + \frac{2y}{f_Y(y)} \int_{\mathbb R} \operatorname{sgn}(u)e^{-|u|/b}f_Y(y+u)\,du + \frac{1}{f_Y(y)} \int_{\mathbb R} (2|u|-b)e^{-|u|/b}f_Y(y+u)\,du \)} \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior moment generating function & \(\displaystyle e^{tx},\quad |t|<1/b\) & \makebox[\linewidth][c]{\(\displaystyle e^{ty} +\frac{e^{ty}}{f_Y(y)} \left[ \int_{\mathbb R} e^{tu} \left\{ t\operatorname{sgn}(u)-\frac{bt^2}{2} \right\} e^{-|u|/b}f_Y(y+u)\,du \right] \)} \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior distribution function at \(a\) & \(\displaystyle \mathbf 1\{x\le a\},\quad a\in\mathbb R\) & \makebox[\linewidth][c]{\(\displaystyle \begin{cases} \displaystyle \frac{ e^{-(y-a)/b} \left\{ f_Y(a)-bD_+f_Y(a) \right\} }{ 2f_Y(y) }, & a\le y, \\[1.1em] \displaystyle 1- \frac{ e^{-(a-y)/b} \left\{ f_Y(a)+bD_+f_Y(a) \right\} }{ 2f_Y(y) }, & a>y. \end{cases} \)} \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior squared risk around \(a\) & \(\displaystyle (x-a)^2,\quad a\in\mathbb R\) & \makebox[\linewidth][c]{\(\displaystyle (y-a)^2 + \frac{2(y-a)}{f_Y(y)} \int_{\mathbb R} \operatorname{sgn}(u)e^{-|u|/b}f_Y(y+u)\,du + \frac{1}{f_Y(y)} \int_{\mathbb R} (2|u|-b)e^{-|u|/b}f_Y(y+u)\,du \)} \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior hinge loss around \(a\) & \(\displaystyle (x-a)_+,\quad a\in\mathbb R\) & \makebox[\linewidth][c]{\(\displaystyle (y-a)_+ + \frac{1}{f_Y(y)} \int_{0}^{\infty} e^{-|u|/b}f_Y(y+u)\,du - \frac{1}{f_Y(y)} \int_{\min(0,a-y)}^{\max(0,a-y)} e^{-|u|/b}f_Y(y+u)\,du - \frac{b e^{-|y-a|/b}f_Y(a)}{2f_Y(y)} \)} \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior pinball loss around \(a\) & \(\displaystyle \rho_\tau(x-a), \; a\in\mathbb R,\ \tau\in(0,1)\) & \(\displaystyle \begin{aligned} & \tau(y-a)_+ + (1-\tau)(a-y)_+ + \frac{\tau}{f_Y(y)} \int_{0}^{\infty} e^{-|u|/b}f_Y(y+u)\,du + \frac{1-\tau}{f_Y(y)} \int_{-\infty}^{0} e^{-|u|/b}f_Y(y+u)\,du \\ &\quad - \frac{1}{f_Y(y)} \int_{\min(0,a-y)}^{\max(0,a-y)} e^{-|u|/b}f_Y(y+u)\,du - \frac{b e^{-|y-a|/b}f_Y(a)}{2f_Y(y)} \end{aligned} \) \\ \addlinespace[0.75em] \addlinespace[0.75em] \multicolumn{3}{@l}{Panel B. Tweedie representations for transformations}\\ Posterior variance & \(\displaystyle \text{--}\) & \makebox[\linewidth][c]{\(\displaystyle \frac{1}{f_Y(y)} \int_{\mathbb R} (2|u|-b)e^{-|u|/b}f_Y(y+u)\,du - \left[ \frac{1}{f_Y(y)} \int_{\mathbb R} \operatorname{sgn}(u)e^{-|u|/b}f_Y(y+u)\,du \right]^2 \)} \\ \addlinespace[0.75em] \bottomrule \end{tabular} \begin{tablenotes} • Notes: Throughout, \(Y=X+V\), \(X\perp\!\!\!\!\perp V\), \(V\sim\operatorname{Laplace}(0,b)\), \(b>0\), and \(f_Y(y)>0\). The operator \(D_+\) is the right derivative of the marginal density with respect to its scalar argument: \(D_+f_Y(a)=\lim_{h\downarrow0}\{f_Y(a+h)-f_Y(a)\}/h\). Here \((z)_+=\max\{z,0\}\), and \(\rho_\tau(z)=\tau z_+ +(1-\tau)(-z)_+\) for \(\tau\in(0,1)\). \end{tablenotes} \end{threeparttable} \end{table}

Formally, let \(\mathbb S_{++}^d\) denote the set of symmetric positive definite \(d\times d\) matrices. The model is \[ Y=X+\Sigma^{1/2}V, \] where \(V\sim\mathcal N(0,I_d)\), \((X,\Sigma)\sim P_{X,\Sigma}\), and \((X,\Sigma)\perp\!\!\!\!\perp V\). It is common practice to consider $\Sigma$ to be known and equal to a sample-based estimate.

At first glance, heteroskedasticity seems to put the model outside the scope of the preceding sections, since there is no single fixed noise distribution to deconvolve. The key observation is that, after conditioning on \(\Sigma=\Sigma_0\) and standardizing by \(\Sigma_0^{-1/2}\), the model becomes an additive Gaussian model with identity noise covariance. The results from Sections (ref) and (ref) can therefore be applied conditionally, yielding Tweedie representations in terms of the conditional density \(f_{Y\mid\Sigma}(\cdot\mid\Sigma_0)\).

Critically, this conditional approach does not require an explicit model for the dependence between the latent parameter \(X\) and the noise covariance \(\Sigma\). This distinguishes the approach from the existing empirical Bayes literature on heteroskedastic normal means, which imposes parametric structure, or other restrictions, on the joint distribution \(P_{X,\Sigma}\) chamberlain1984,jiang2010empirical,efron2016empirical,Weinstein2018,ignatiadis2019covariate,kline2024discrimination,chen2026empirical.

For a measurable function \(g:\mathbb R^d\to\mathbb R\), the objects of interest in this section are conditional posterior functionals of the form \[ \mathbb E\!\left[g(X)\mid Y=y,\Sigma=\Sigma_0\right]. \] If \(\Sigma\) were degenerate at \(\Sigma_0\in\mathbb S_{++}^d\), this would reduce to the usual posterior quantity under Gaussian noise with covariance \(\Sigma_0\). In the general case, however, \(\Sigma\) is random, so the model is not initially of the additive-noise form studied in the previous sections. Conditioning on \(\Sigma=\Sigma_0\) removes this difficulty. To see this, define \[ Z\coloneq \Sigma^{-1/2}Y, \qquad W\coloneq \Sigma^{-1/2}X. \]

Then, conditional on \(\Sigma=\Sigma_0\), we have $Z=W+V$, $W\perp\!\!\!\!\perp V$, $V\sim\mathcal N(0,I_d)$. In particular, \[ \varphi_{Z\mid \Sigma}(\omega\mid \Sigma_0) = \varphi_{W\mid \Sigma}(\omega\mid \Sigma_0)\varphi_V(\omega). \]

Thus, conditional on \(\Sigma=\Sigma_0\), the standardized model satisfies the same structural conditions as the additive Gaussian model studied in Sections (ref) and (ref). The results developed there therefore apply conditional on \(\Sigma\), and hence pointwise in \(\Sigma_0\). The next corollary records this implication formally.

proposition[Conditional Tweedie representation in the heteroskedastic Gaussian sequence model] Fix \(y\in\mathbb R^d\), and for \(\Sigma_0\in\mathbb S_{++}^d\), set $z_0\coloneq \Sigma_0^{-1/2}y$. Let \(g:\mathbb R^d\to\mathbb R\) be measurable. Then, for \(P_\Sigma\)-almost every \(\Sigma_0\), the following statements hold. \begin{enumerate}[label=(\alph*)] • If the expectations below are well defined, then \(f_{Z\mid \Sigma}(z_0\mid \Sigma_0)>0\) and \[ \mathbb E[g(X)\mid Y=y,\Sigma=\Sigma_0] = \frac{ \displaystyle \int_{\mathbb R^d} g(\Sigma_0^{1/2}w)\phi(z_0-w)\, dP_{W\mid\Sigma}(w\mid\Sigma_0) }{ f_{Z\mid\Sigma}(z_0\mid\Sigma_0) }. \] • Define $\lambda_{g,y,\Sigma_0}(w) \coloneq g(\Sigma_0^{1/2}w)\phi(z_0-w)$. If \(\lambda_{g,y,\Sigma_0}\in C_0(\mathbb R^d;\mathbb C)\), then there exists a unique continuous linear functional $\mathcal T_{g,y,\Sigma_0}\in \mathcal A_V'(\mathbb R^d)$, such that \[ \mathcal T_{g,y,\Sigma_0}[f] = \int_{\mathbb R^d} \lambda_{g,y,\Sigma_0}(w)\,d\mu(w) \qquad \text{for every } f=\phi*\mu \in \mathcal{A}_V(\mathbb R^d). \] In particular, \[ \mathcal T_{g,y,\Sigma_0} \!\left[f_{Z\mid\Sigma}(\cdot\mid\Sigma_0)\right] = \int_{\mathbb R^d} \lambda_{g,y,\Sigma_0}(w)\, dP_{W\mid\Sigma}(w\mid\Sigma_0), \] and therefore \[ \mathbb E[g(X)\mid Y=y,\Sigma=\Sigma_0] = \frac{ \mathcal T_{g,y,\Sigma_0} \!\left[f_{Z\mid\Sigma}(\cdot\mid\Sigma_0)\right] }{ f_{Z\mid\Sigma}(z_0\mid\Sigma_0) }. \] • Suppose, in addition, that $\lambda_{g,y,\Sigma_0}\in \Xi^0(\mathbb R^d;\mathbb R)$, and $\widetilde{\lambda}_{g,y,\Sigma_0}\in L^1(\mathbb R^d;\mathbb C)$. Define \[ \mathcal Q_{g,y,\Sigma_0}(\omega) \coloneq \frac{\widetilde{\lambda}_{g,y,\Sigma_0}(\omega)} {\varphi_V(-\omega)} = e^{\|\omega\|_2^2/2} \widetilde{\lambda}_{g,y,\Sigma_0}(\omega), \qquad \omega\in\mathbb R^d. \] If \(\mathcal Q_{g,y,\Sigma_0}\) satisfies conditions \((i)\) and \((ii)\) of Theorem (ref) for some \(N\ge 0\) and \(k\ge 0\), then for every \(f\in\mathcal A_V(\mathbb R^d)\) and every sequence \((f_n)_{n\ge1}\subset\mathscr S(\mathbb R^d)\) such that \(\|f_n-f\|_{\Xi^k}\to0\), \[ \mathcal T_{g,y,\Sigma_0}[f] = \lim_{n\to\infty} \mathcal F^{-1} \{\mathcal I_{\mathcal Q_{g,y,\Sigma_0}}\}[f_n], \] and the limit is independent of the approximating sequence. \end{enumerate}

The proposition implies that, once we condition on \(\Sigma=\Sigma_0\), the heteroskedastic problem can be handled with the same computational strategy developed in Section (ref). In particular, Tweedie formulas can be obtained by working with the Fourier-domain representer \(\mathcal Q_{g,y,\Sigma_0}\), exactly as in the additive Gaussian model after the appropriate standardization.

I apply this conditional Tweedie calculus to a range of empirical Bayes estimands. Table (ref) summarizes the one-dimensional formulas. Appendix (ref) provides the derivations and extends the formulas for the posterior mean, posterior covariance, and posterior moment generating function to arbitrary dimensions.

landscape\begin{table}[ht] \scriptsize {4pt} \begin{threeparttable} \caption{ Examples of conditional Tweedie representations in the heteroskedastic Gaussian sequence model, \(d=1\).} \begin{tabular}{@ >{\arraybackslash}p{0.225\linewidth} >{\arraybackslash}p{0.135\linewidth} >{\arraybackslash}p{0.615\linewidth} @} \toprule Name of functional & \(g(x)\) & \(\mathbb E[g(X)\mid Y=y,\Sigma=\sigma^2]\) \\ \midrule \multicolumn{3}{@l}{Panel A. Tweedie representations for conditional expectations}\\ Posterior mean & \(\displaystyle x\) & \makebox[\linewidth][c]{\(\displaystyle y+\sigma^2 \frac{\partial_y f_{Y\mid\Sigma}(y\mid\sigma^2)} {f_{Y\mid\Sigma}(y\mid\sigma^2)} \)} \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior second moment & \(\displaystyle x^2\) & \makebox[\linewidth][c]{\(\displaystyle y^2+\sigma^2 +2y\sigma^2 \frac{\partial_y f_{Y\mid\Sigma}(y\mid\sigma^2)} {f_{Y\mid\Sigma}(y\mid\sigma^2)} +\sigma^4 \frac{\partial_y^2 f_{Y\mid\Sigma}(y\mid\sigma^2)} {f_{Y\mid\Sigma}(y\mid\sigma^2)} \)} \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior moment generating function & \(\displaystyle e^{tx},\quad t\in\mathbb R\) & \makebox[\linewidth][c]{\(\displaystyle \exp\!\left(ty+\frac12\sigma^2t^2\right) \frac{f_{Y\mid\Sigma}(y+\sigma^2t\mid\sigma^2)} {f_{Y\mid\Sigma}(y\mid\sigma^2)} \)} \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior \(k\)-th raw moment & \(\displaystyle x^k,\quad k\in\mathbb N\) & \makebox[\linewidth][c]{\(\displaystyle \sum_{r=0}^{k} \binom{k}{r}\sigma^{2r} \left[ \sum_{j=0}^{\lfloor (k-r)/2\rfloor} \frac{(k-r)!}{2^j j!(k-r-2j)!} \sigma^{2j}y^{k-r-2j} \right] \frac{\partial_y^r f_{Y\mid\Sigma}(y\mid\sigma^2)} {f_{Y\mid\Sigma}(y\mid\sigma^2)} \)} \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior \(2m\)-risk around \(a\) & \(\displaystyle |x-a|^{2m},\quad a\in\mathbb R,\ m\in\mathbb N\) & \makebox[\linewidth][c]{\(\displaystyle \sum_{r=0}^{2m} \binom{2m}{r}\sigma^{2r} \left[ \sum_{j=0}^{\lfloor (2m-r)/2\rfloor} \frac{(2m-r)!}{2^j j!(2m-r-2j)!} \sigma^{2j}(y-a)^{2m-r-2j} \right] \frac{\partial_y^r f_{Y\mid\Sigma}(y\mid\sigma^2)} {f_{Y\mid\Sigma}(y\mid\sigma^2)} \)} \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior distribution function at \(a\) & \(\displaystyle \mathbf 1\{x\le a\},\quad a\in\mathbb R\) & \makebox[\linewidth][c]{\(\displaystyle \Phi\!\left(\frac{a-y}{\sigma}\right) + \frac{1}{f_{Y\mid\Sigma}(y\mid\sigma^2)} \lim_{n\to\infty} \sum_{\ell=1}^{\infty} \frac{(-1)^\ell}{\ell!} \left( \frac{\sigma^2}{\sqrt{\sigma^2+n^{-2}}} \right)^\ell \Phi^{(\ell)} \!\left( \frac{a+n^{-1/2}-y}{\sqrt{\sigma^2+n^{-2}}} \right) \partial_y^\ell f_{Y\mid\Sigma}(y\mid\sigma^2) \)} \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior hinge loss around \(a\) & \(\displaystyle (x-a)_+,\quad a\in\mathbb R\) & \(\displaystyle \begin{aligned} & \sigma\phi\!\left(\frac{a-y}{\sigma}\right) + (y-a) \left\{ 1-\Phi\!\left(\frac{a-y}{\sigma}\right) \right\} + \sigma^2 \left\{ 1-\Phi\!\left(\frac{a-y}{\sigma}\right) \right\} \frac{\partial_y f_{Y\mid\Sigma}(y\mid\sigma^2)} {f_{Y\mid\Sigma}(y\mid\sigma^2)} \\ &\quad+ \frac{1}{f_{Y\mid\Sigma}(y\mid\sigma^2)} \lim_{n\to\infty} \sum_{\ell=2}^{\infty} \frac{(-1)^\ell}{\ell!} \left( \frac{\sigma^2}{\sqrt{\sigma^2+n^{-2}}} \right)^\ell \sqrt{\sigma^2+n^{-2}} \Phi^{(\ell-1)} \!\left( \frac{a+n^{-1/2}-y}{\sqrt{\sigma^2+n^{-2}}} \right) \partial_y^\ell f_{Y\mid\Sigma}(y\mid\sigma^2) \end{aligned} \) \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior absolute risk around \(a\) & \(\displaystyle |x-a|,\quad a\in\mathbb R\) & \(\displaystyle \begin{aligned} & 2\sigma\phi\!\left(\frac{a-y}{\sigma}\right) + (a-y) \left\{ 2\Phi\!\left(\frac{a-y}{\sigma}\right)-1 \right\} + \sigma^2 \left\{ 1-2\Phi\!\left(\frac{a-y}{\sigma}\right) \right\} \frac{\partial_y f_{Y\mid\Sigma}(y\mid\sigma^2)} {f_{Y\mid\Sigma}(y\mid\sigma^2)} \\ &\quad+ \frac{2}{f_{Y\mid\Sigma}(y\mid\sigma^2)} \lim_{n\to\infty} \sum_{\ell=2}^{\infty} \frac{(-1)^\ell}{\ell!} \left( \frac{\sigma^2}{\sqrt{\sigma^2+n^{-2}}} \right)^\ell \sqrt{\sigma^2+n^{-2}} \Phi^{(\ell-1)} \!\left( \frac{a+n^{-1/2}-y}{\sqrt{\sigma^2+n^{-2}}} \right) \partial_y^\ell f_{Y\mid\Sigma}(y\mid\sigma^2) \end{aligned} \) \\ \multicolumn{3}{@l}{Panel B. Tweedie representations for transformations}\\ Posterior variance & \(\displaystyle \text{--}\) & \makebox[\linewidth][c]{\(\displaystyle \sigma^2+\sigma^4 \left\{ \frac{\partial_y^2 f_{Y\mid\Sigma}(y\mid\sigma^2)} {f_{Y\mid\Sigma}(y\mid\sigma^2)} - \left( \frac{\partial_y f_{Y\mid\Sigma}(y\mid\sigma^2)} {f_{Y\mid\Sigma}(y\mid\sigma^2)} \right)^2 \right\} \)} \\ \addlinespace[0.75em] \addlinespace[0.75em] Posterior \(k\)-th centered moment & \(\displaystyle \text{--}\) & \makebox[\linewidth][c]{\(\displaystyle \sum_{r=0}^{k} \binom{k}{r}\sigma^{2r} \left[ \sum_{j=0}^{\lfloor (k-r)/2\rfloor} \frac{(k-r)!}{2^j j!(k-r-2j)!} \sigma^{2j} \left( -\sigma^2 \frac{\partial_y f_{Y\mid\Sigma}(y\mid\sigma^2)} {f_{Y\mid\Sigma}(y\mid\sigma^2)} \right)^{k-r-2j} \right] \frac{\partial_y^r f_{Y\mid\Sigma}(y\mid\sigma^2)} {f_{Y\mid\Sigma}(y\mid\sigma^2)} \)} \\ \bottomrule \end{tabular} \begin{tablenotes} • Notes: Throughout, \(\sigma>0\), \(f_{Y\mid\Sigma}(y\mid\sigma^2)>0\), and all derivatives are with respect to \(y\). Also, \(\Phi\) denotes the standard normal distribution function and \(\Phi^{(\ell)}\) its \(\ell\)-th derivative. Here \((z)_+=\max\{z,0\}\) \end{tablenotes} \end{threeparttable} \end{table}

Conclusion

Under Gaussian noise, Tweedie’s formula expresses the posterior mean as a functional of the observed marginal density, thereby avoiding nonparametric deconvolution. This paper develops a general framework for analogous identities governing conditional expectations in additive-noise models. The main contribution is twofold. First, I characterize when posterior functionals admit a representation through a linear map acting on the observed density---a structure I call the Tweedie functional---and establish existence, uniqueness, and continuity under general conditions. Second, I provide a constructive approach based on Fourier analysis for deriving such representations, showing that the relevant operator can be obtained by extending the inverse Fourier transform of an explicitly characterized tempered distribution to the space of admissible observed densities.

This perspective turns the search for Tweedie-type formulas into a tractable problem in the calculus of tempered distributions. It recovers the classical Gaussian identity, yields new posterior-mean representations for a broad family of noise laws, and, in some cases, delivers expressions that can be unbiasedly estimated from the data. The applications also cover posterior functionals under the Laplace mechanism used in differential privacy. The framework therefore provides a unified route to evaluating posterior functionals without explicitly recovering the latent distribution, and without being tied to a single noise structure.

The heteroskedastic application shows that the method is not confined to the standard additive model. By conditioning on the realized covariance matrix and applying the appropriate change of variables, the model regains the structure needed for the theory to apply. This yields Tweedie representations for posterior means and other functionals directly in terms of the observable conditional density, without imposing additional structure on the model.