EconBase
← Back to paper

Estimation of distribution functions, their jumps and interval probabilities under measurement error

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

84,630 characters

\begin{center}
\large \textsc{Estimation of distribution functions, their jumps and interval probabilities under measurement error}\normalsize\\[.2in]

\begin{tabular}{c}
\multicolumn{1}{c}{ \sc Kairat Mynbaev \rm} \\[.1in]
International School of Economics \\
Kazakh-British Technical University \\
Tolebi 59 \\
Almaty 050000, Kazakhstan \\
email: kairat\[email removed] \\
\end{tabular}\\[.1in]

\begin{tabular}{c}
\multicolumn{1}{c}{\sc Carlos Martins-Filho \rm} \\[.1in]
Department of Economics \\
University of Colorado  \\
Boulder, CO 80309-0256, USA  \\
email: [email removed]  \\
\end{tabular}\\[.1in]

and  \\[.1in]

\begin{tabular}{c}
\multicolumn{1}{c}{\sc Chad Brown \rm} \\[.1in]
Department of Economics \\
University of Manchester  \\
Manchester,  M13 9PL,  UK\\
email: [email removed]  \\
\end{tabular}\\[.1in]

August 2026\\[.1in]
\end{center}


\noindent \bf Abstract. \rm
We consider the classical additive measurement-error model $X=Y+Z$,
where the latent random variable $Y$ has unknown distribution $F_Y$ and the
error $Z$ has a known distribution.
We develop direct estimators for three functionals of $F_Y$:
(i) $F_Y(x)$ at continuity points;
(ii) interval probabilities $F_Y(y)-F_Y(x)$ when $x<y$ are continuity points;
and
(iii) the size of a jump at a prespecified discontinuity.
We derive non-asymptotic bias and variance bounds, and establish asymptotic unbiasedness and consistency.
Unlike previous work, we do not require $F_Y$ to admit a density, have a mixture representation, or satisfy global Sobolev smoothness assumptions.
The framework accommodates arbitrary latent distributions, including those with both discrete and continuous components, and distributions with multiple jumps.
These results rely on a link between Fourier inversion theorems and the algebraic structure of a class of estimators proposed in \cite{Mynbaev2022}.
A simulation study evaluates feasible tuning procedures and, where available, compares the finite-sample performance of the proposed estimators with existing methods.
\\[.1in]

\noindent \bf Keywords and phrases. \rm Error-in-measurement model; nonparametric deconvolution estimation; distribution function jump estimation; Fourier inversion theorems.\\

\noindent \textbf{MSC 2020}: Primary 60E10,  62G05; secondary 62G20.\\[.25in]

\clearpage
\setlength{\baselineskip}{24pt}
\pagestyle{plain}
\setcounter{page}{1}
\setcounter{footnote}{0}




\section{Introduction}

Distribution-function, interval-probability, and point-mass estimation are fundamental statistical problems, but observations are often contaminated by measurement error.
We consider the classical additive measurement-error model $X:=Y+Z$ where $Y$ is a latent random variable with unknown distribution $F_Y$ and the measurement error $Z$ is independent of $Y$ and has a known distribution.
Based on a random sample from $X$, we develop direct deconvolution estimators for three functionals of $F_Y$:
$F_Y(x)$ when $x$ is a continuity point;
interval probabilities $F_Y(x,y) := F_Y(y)-F_Y(x)$ when $x<y$ and both are continuity points;
and
the size of a jump $p_x:= F_Y(x)-\underset{\epsilon \downarrow 0}{\lim} \,F_Y(x-\epsilon)$ when $x$ is a known point of discontinuity.


The deconvolution literature has primarily focused on estimating the \textit{density} of $Y$, when it exists; see \citet{Meister2009} and the references therein. Early estimators of the distribution function $F_Y$ were consequently obtained by integrating suitable density deconvolution estimators, as in \citet{Zhang1990}, \citet{Fan1991}, and \citet{Hall2008}. More recently, \citet{Dattner2011} and \citet{Dattner2013} used Fourier inversion representations of distribution functions (see \citealp{Gil-Pelaez1951}) to construct pointwise estimators of $F_Y$ without first estimating its density.
These methods are more closely related to ours, but their bias and convergence-rate analyses assume that $F_Y$ possesses a density satisfying global Sobolev smoothness conditions. Thus, even direct distribution-function estimators have generally been studied under global regularity restrictions on the latent distribution.
The literature on interval probabilities remains even more limited; although they can be estimated by differencing existing pointwise estimators, direct estimators and their theoretical properties have received little attention.


Estimation of jump sizes is important whenever the latent distribution possesses atoms, as occurs naturally for mixed discrete-continuous distributions, censored variables, heaped survey responses, and economic variables with point masses. While jump estimation is rather straightforward in the absence of measurement error, additive noise makes the recovery of jumps a substantially more difficult deconvolution problem.


Estimation of jumps in a latent distribution function under additive measurement error
has previously been considered by \cite{vanEs2008} and \cite{gugushvili2011}, who assume a single atom at the origin together with an absolutely continuous component, and by \cite{Lee2013}, who allow finitely many atoms and continuous components.
These approaches rely on a specified discrete-continuous mixture structure for $F_Y$ and regularity conditions on the continuous component.
In contrast, we consider an arbitrary jump at any prespecified location of an arbitrary latent distribution. Apart from the target jump, no structural restrictions are placed on  $F_Y$: it may have countably many additional atoms and need not admit a discrete-continuous mixture representation with a globally smooth density.






The three estimators we propose are constructed within a kernel-regularization framework motivated by the generalized Fourier inversion theorem established in \citet{Mynbaev2022}.
For each estimator, we establish non-asymptotic bias and variance bounds, asymptotic unbiasedness, and consistency.
These results impose no restrictions on $F_Y$ beyond continuity at the evaluation points for $F_Y(x)$ and $F_Y(x,y)$, or a prespecified jump location for $p_x$.
Our non-asymptotic mean squared error (MSE) bounds also yield explicit convergence rates when $F_Y$ satisfies mild local regularity and tail conditions.  These conditions do not require a density or mixture representation and impose no global Sobolev smoothness.
Thus, while much of the previous literature has focused on minimax rates over global smoothness classes, we provide finite-sample and asymptotic guarantees under substantially weaker assumptions on the latent distribution.






A simulation study evaluates the finite-sample performance of all three estimators using feasible tuning procedures.
For the $F_Y(x)$ and $F_Y(x,y)$ estimators, we select tuning parameters using asymptotic rules and reconvolution cross-validation. For the $p_x$ estimator, we consider an asymptotic rule and a Goldenshluger-Lepski procedure for bandwidth selection.
We compare our estimator of $F_Y(x)$ with those of \citet{Hall2008} and \citet{Dattner2013}; for interval probabilities $F_Y(x,y)$, we also include the implied estimators obtained by differencing their pointwise estimators.




The rest of the paper is organized as follows. Section~\ref{sec:motiv} provides the theoretical motivation for our estimators.
Section~\ref{sec:DirDecEst} develops the main results for the pointwise distribution-function, interval-probability, and jump-size estimators.
Section~\ref{sec:sims} considers implementation and presents the simulation study.
Section~\ref{sec:conclusion} concludes and gives directions for future research.
Appendix~\ref{app:proofs} contains proofs for the theoretical results in the main text,
Appendix~\ref{app:noncompH} provides extensions to non-compact regularization kernels,
and Appendix~\ref{app:sims} provides details on the simulation study and implementation of the proposed estimators.









\noindent \bf Notation. \rm
We adopt the following notation throughout the paper.
$F$ denotes a (proper) distribution function, and when necessary we write $F_X$ to denote the distribution function associated with the random variable $X$.
$\chi_S$ is the indicator function for the set $S$.
$C(f)$ denotes the points of continuity of the function $f$, and $J(f)$ denotes the points where $f$ has a jump discontinuity.
$C_b$ denotes the space of uniformly bounded continuous functions on $\mathds{R}$ with norm $\|f\|_{C_b}:=\underset{x \in \mathds{R}}{\sup}|f(x)|$;
$L(\mathds{R})$ denotes the space of measurable functions $f:(\mathds{R}, \mathcal{B}) \to (\mathds{R}, \mathcal{B})$ with norm $\|f\|_{L}=\int_{\mathds{R}}|f(t)|dt < \infty$ where $\mathcal{B}$ are the Borel sets of $\mathds{R}$.
$\mathcal{F}$ and $\mathcal{F}^{-1}$ denote, respectively, the Fourier and inverse Fourier transforms, where
for a function $f \in L(\mathds{R})$, $\left(\mathcal{F}f \right)(t):=\int_{\mathds{R}}e^{ist}f (s)ds$ and $\left(\mathcal{F}^{-1}f \right)(t):=\frac{1}{2\pi }\int_{\mathds{R}}e^{-ist}f (s)ds$,
and
for a distribution function $F$, $\left(\mathcal{F}F \right)(t):=\int_{\mathds{R}}e^{ist}dF(s)$ and $\left(\mathcal{F}^{-1}F \right)(t)=\frac{1}{2\pi }\int_{\mathds{R}}e^{-ist}dF(s)$.
$f \star g$ denotes the convolution of $f$ and $g$, where $f$ and $g$ may be functions or measures.
For a function $f:\mathds{R} \to \mathds{R}$, $\supp \, f$ denotes its support.
For a set function $G$ defined on intervals $[a,b]\subset \mathds{R}$, we write
$\|G\|_\infty := \underset{a,b\in\mathds{R}, a<b}{\sup}|G([a,b])|$,
and
\begin{equation}\label{phiGdef}
    \phi _{G}(N) :=\max \left\{ \underset{b<-N}{\sup}|G([a,b])|,\ \underset{a>N}{\sup}|G([a,b])|,\
\underset{a<-N,\ b>N}{\sup}|G([a,b])-1|\right\}\quad \mbox{for}\;\; N>0.
\end{equation}
For $f:\mathds{R}\to\mathds{R}$ and $h>0$ we write $\omega_f(x,h):=\underset{0<u\le h}{\sup} \left| f(x+u)-f(x-u) \right|.$


\section{Motivation for the new deconvolution estimators}\label{sec:motiv}

The estimators proposed in this paper are motivated by a general class of nonparametric estimators introduced in \cite{Mynbaev2022}, together with a generalized Fourier inversion theorem.  \cite{Mynbaev2022} considered the estimation of $F(x,y):=F(y)-F(x)$ for $x<y$ and $x,y \in C(F)$ based on a random sample $\{X_j\}_{j=1}^n$, $n \in \mathds{N}$, where $F$ is the distribution function associated with the random variable $X$.  In a setting where there is no measurement error they suggested the  nonparametric estimator
\begin{equation}
\tilde{F}(x,y)=\frac{1}{n}\sum_{i=1}^{n}G\left( \left[ \frac{X_{i}-y}{h},\frac{X_{i}-x}{h}\right] \right),  \label{1}
\end{equation}
where $h>0$ is a bandwidth and $G$ is a function on intervals $[a,b]$, where $a,b \in \mathds{R}$.  They showed that if $G$ satisfies
$\left\Vert G\right\Vert _{\infty} < \infty$
and
$\lim_{N\to\infty}\phi _{G}(N)=0$,
then for any $h>0$ and $\delta \in (0,1)$,
\begin{equation}
\left\vert E \left(\tilde{F}(x,y)\right)-F(x,y)\right\vert \leq \phi _{G}(h^{\delta
-1})+\left( 1+\left\Vert G\right\Vert _{\infty}\right) \left[ \omega_F (x,h^{\delta
})+\omega_F (y,h^{\delta })\right]\!.  \label{4}
\end{equation}
It follows immediately from \eqref{4} that, for any distribution function $F$ and any $x,y\in C(F)$ with $x<y$,
\begin{equation}\label{2}
E\left( \tilde{F}(x,y) \right)\rightarrow F(x,y),\quad\text{as }h\rightarrow 0
.
\end{equation}
Although the estimator \eqref{1} was developed independently of Fourier methods, its expectation admits a representation that coincides with a generalized Fourier inversion formula. This observation provides the key motivation for the deconvolution estimators proposed in this paper.  Specifically, \cite{Mynbaev2022} identified a link between $E(\tilde{F}(x,y))$ and classical Fourier inversion theorems.  Recall that L\'evy's inversion theorem (see \citealp{Lukacs1970}) states that
\begin{equation}
\frac{1}{2\pi }\int_{-1/h}^{1/h}\frac{e^{-ixt}-e^{-iyt}}{it}(\mathcal{F}F)
(t)dt\rightarrow F(x,y),\mbox{ as } h\rightarrow 0,  \label{7}
\end{equation}
and Borovkov's inversion theorem (see \citealp{Borovkov2013}) gives
\begin{equation}
\frac{1}{2\pi }\int_{\mathds{R}}\frac{e^{-ixt}-e^{-iyt}}{it}e^{-h^{2}t^{2}}(\mathcal{F}F)
(t)dt\rightarrow F(x,y),\ h\rightarrow 0.  \label{8}
\end{equation}
The left-hand sides of \eqref{7} and \eqref{8} are special cases of
$
\frac{1}{2\pi }\int_{\mathds{R}}\frac{e^{-ixt}-e^{-iyt}}{it}H(ht)(\mathcal{F}F) (t)dt
$ for $H\in L(\mathds{R})$.  The function $H(h \cdot)$ plays the role of a regularization kernel. Different choices of $H$ recover classical Fourier inversion formulas and lead to different deconvolution estimators.  \cite{Mynbaev2022} showed that if $G([a,b])=\frac{1}{2\pi}\int_a^b(\mathcal{F}H)(u)du$, then
\begin{equation}
E\left(\tilde{F}(x,y) \right)=\int_{\mathds{R}}G\left( \left[ \frac{t-y}{h},\frac{t-x}{h}\right]
\right) dF(t)=\frac{1}{2\pi }\int_{\mathds{R}}\frac{e^{-ixt}-e^{-iyt}}{it}H(ht)(\mathcal{F}F)
(t)dt.  \label{11}
\end{equation}
The asymptotic unbiasedness of $\tilde{F}(x,y)$ in equation \eqref{2} and the link revealed in equation \eqref{11} provides the motivation for the estimators we propose in subsequent sections.


\section{Direct deconvolution estimators}\label{sec:DirDecEst}

This section defines and provides the main results for the three direct deconvolution estimators. Throughout,
suppose $X,\,Y,$ and $Z$ are random variables defined on the probability space $(\Omega, \mathcal{F},P)$ with
\begin{equation}\label{sum}
X=Y+Z
\end{equation}
and let $F_Y$ be the distribution function of $Y$. Estimation is to be conducted based on a random sample $\{X_j\}_{j=1}^n$ of observations on $X$, which has distribution function $F_X$.  $Z$ is an unobserved measurement error, independent of $Y$, with known distribution function denoted by $F_Z$.

\subsection{Estimation of $F_Y(x)$ for $x \in C(F_Y)$}\label{sec:FYx}

We consider the estimation of $F_Y(x)$ where $x \in C(F_Y)$.
Independence of $Y$ and $Z$ and equation \eqref{sum} imply that $F_X=F_Y\star F_Z$.  Furthermore, if $\Phi_X:=\mathcal{F}F_X$ denotes the Fourier transform of $F_X$, we have $\Phi _{X}=\Phi _{Y}\Phi_{Z}$.  If $\Phi_{Z}(t) \neq 0$ for all $t \in \mathds{R}$ we can write $\Phi _{Y}(t)=\Phi_X(t)/\Phi_{Z}(t)$.  If $\Phi_X/\Phi_{Z}$ is integrable we have $F_Y(x) =\left( \mathcal{F}^{-1}\frac{\Phi _{X}}{\Phi _{Z}}\right)(x)$.  In the general case where $\Phi_X/\Phi_{Z}$ is not integrable, we regularize it by multiplying by $H(h\cdot)$ with $h>0$, $H \in L(\mathds{R})$ chosen in such a way as to have
\begin{equation}\label{moteq}
F_Y(x)=\underset{h\to 0}{\lim}\,\mathcal{F}^{-1}\left\{\frac{\Phi _{X}}{\Phi _{Z}}H(h\cdot)\right\}(x).
\end{equation}
Equations \eqref{11} and \eqref{moteq} motivate our estimator for $F_Y(x)$.
For $h,\,\lambda>0$ define
\begin{equation}\label{decon}
\widehat{F}_{Y}(x):=\frac{1}{n}\sum_{j=1}^{n}\! K_{h,\lambda }\!\left(
X_{j}-x\right)\!,
\;\; \mbox{where} \;\;\;
K_{h,\lambda }(t):=\frac{1}{2\pi }\int_{\mathds{R}}e^{its}\alpha _{h,\lambda}\!\left( s\right) ds
,
\quad
\alpha _{h,\lambda }(s):=\frac{e^{is/\lambda }-1}{is}\frac{H(hs)}{\Phi _{Z}(s)}.
\end{equation}
We make the following assumption to ensure that $\widehat{F}_{Y}(x)$ is well defined.
    \begin{condition}\label{A1}
    1. (Identifiability) $\Phi _{Z}(s)\neq 0$ for each $s\in \mathds{R}$; 2. (Regularity) $\alpha _{h,\lambda } \in L(\mathds{R})$ for each $h,\,\lambda >0$; 3. (Symmetry) $H$ is even and real valued.
    \end{condition}

Given Assumptions \ref{A1}.1 and \ref{A1}.2, the function $K_{h,\lambda }$ is continuous and bounded for each $h,\lambda >0$.
Given Assumption \ref{A1}.3,
\begin{equation*}
\overline{\alpha _{h,\lambda }(s)}
=\frac{e^{-is/\lambda }-1}{-is}\frac{H(-hs)}{\Phi_{Z}(-s)}
=\alpha _{h,\lambda }(-s),
\quad \text{ and } \quad
\overline{K_{h,\lambda }(t)}
=\frac{1}{2\pi }\int_{\mathds{R}}e^{-its}\alpha_{h,\lambda }\left( -s\right) ds
=K_{h,\lambda }(t),
\end{equation*}
so $K_{h,\lambda }(t)$ is real-valued.
Hence, under Assumption \ref{A1},  $\widehat{F}_{Y}(x)$ is continuous, almost surely bounded,  real-valued, and  $E\left(\widehat{F}_{Y}(x)\right)<\infty$.



The algebraic structure of $\widehat{F}_{Y}(x)$ is explained by the following lemma.
This
makes precise the connection between
$\widehat{F}_{Y}(x)$ and the generalized Fourier inversion representation
in Section~\ref{sec:motiv}.
For $H\in L(\mathds{R})$, define
$$
G_H([a,b])
    :=\frac{1}{2\pi}\int_a^b(\mathcal FH)(v)\,dv,
    \qquad \mbox{for $a,b\in \mathds{R}$,\, $a<b$.}
$$
We note that $H \in L(\mathds{R})$ implies $\mathcal{F}H\in C_b(\mathds{R})$ so $G_H([a,b])$ is well defined.

\begin{lemma}\label{lem:algebra_FYx}\label{lem1}
Let  $H\in L(\mathds{R})$.
For any distribution function $F$, $h,\lambda>0$, and $x\in\mathds{R}$,
\[
\frac{1}{2\pi}\int_{\mathds{R}}
    \frac{e^{-it(x-1/\lambda)}-e^{-itx}}{it}
    H(ht)(\mathcal FF)(t)\,dt
=
\int_{\mathds{R}}
G_H\!\left(
\left[
\frac{u-x}{h},
\frac{u-(x-1/\lambda)}{h}
\right]\right)dF(u).
\]
Under the measurement-error model and Assumption~\ref{A1}, taking $F=F_Y$
gives
\[
E\left(\widehat{F}_{Y}(x)\right)
=
\int_{\mathds{R}}
G_H\!\left(
\left[
\frac{u-x}{h},
\frac{u-(x-1/\lambda)}{h}
\right]\right)dF_Y(u).
\]
\end{lemma}

Building on the approach of \cite{Mynbaev2022} discussed in Section \ref{sec:motiv}, the next assumption is used to obtain a finite-sample bias bound and asymptotic unbiasedness.

\begin{condition}\label{AG}
(Inversion kernel)
$H\in L(\mathds{R})$ and the interval function $G_H$ satisfies
$\left\Vert G_H\right\Vert _{\infty} < \infty$
and
$\lim_{N\to\infty}\phi _{G_H}(N)=0$ where
$\phi _{G}$ is defined in \eqref{phiGdef}.
\end{condition}

\begin{remark}\label{rem:inversion_kernel}
For generality, the conditions in Assumption~\ref{AG} are stated in terms of $G_H$.
A convenient sufficient condition stated directly in terms of $H$ is
\[
    H\in L(\mathds{R}),\qquad
    \mathcal FH\in L(\mathds{R}),\qquad
    H(0)=1,
\]
with $H$ continuous at zero. These conditions are not necessary. For
example,
$H=\chi_{[-1,1]}$ satisfies
Assumption~\ref{AG}, although
$
    (\mathcal FH)(u)
    =
    {2\sin (u)}/{u}
    \notin
    L(\mathds{R}).
$
Examples of kernels that satisfy Assumption \ref{AG}
include the Gaussian, Laplace, and Bartlett kernels,
as well as $H(t)=(1-t^2)^4\chi_{[-1,1]}(t)$
which will be used for the simulation study in Section~\ref{sec:sims}.
The assumption that $G_H$ is a bounded function of intervals $[a,b]$ is added since it does not follow from $\mathcal{F}H \in C_b(\mathds{R})$.
If, in addition, $\mathcal{F}H \ge 0$ then $\frac{1}{2\pi}(\mathcal{F}H)$ is a density function associated with the probability measure $G_H$.

\end{remark}


\begin{theorem}\label{thm7}
Suppose Assumptions \ref{A1} and \ref{AG} hold.
Then, for any $F_{Y},$ and $x\in C(F_{Y})$
\begin{equation} \label{unqua2}
\left|E\left(\widehat{F}_{Y}(x)\right)-F_{Y}(x)\right|\rightarrow 0,
\quad \mbox{ as $h,\,\lambda \rightarrow 0$.
}
\end{equation}
Furthermore, for any $\delta \in (0,1)$, $h \in (0,1]$ and $\lambda >0$ such that $\lambda h^\delta \le 1/2$, we have the following non-asymptotic bound
\begin{equation}\label{gen_bound2}
\left|E\left(\widehat{F}_{Y}(x)\right)-F_{Y}(x)\right| \leq \phi _{G_H}(h^{\delta -1})+(1+\left\Vert
G_H\right\Vert _{\infty})\omega_{F_Y} (x,h^{\delta })+(2+\left\Vert G_H\right\Vert _{\infty})F_{Y}\left(x-\frac{1}{\lambda} +1\right).
\end{equation}
\end{theorem}

We note that the bound on the bias of $\widehat{F}_{Y}(x)$ does not rely on any assumption on the distribution $F_Y$. Unlike \cite{Dattner2011} and \cite{Dattner2013}, whose bias analysis requires Sobolev smoothness of $F_Y$, inequality \eqref{gen_bound2} holds without assuming that $F_Y$ possesses a density.

\begin{remark}[Generalized Fourier inversion]\label{r2lem2}
The proof of Theorem~\ref{thm7}, applied to an arbitrary
distribution function $F$, also gives
\[
F(x)
=
\lim_{h,\lambda\to0}
\frac{1}{2\pi}\int_{\mathds{R}}
    \frac{e^{-it(x-1/\lambda)}-e^{-itx}}{it}
    H(ht)(\mathcal FF)(t)\,dt,
    \qquad x\in C(F).
\]
and the non-asymptotic bias bound in Theorem~\ref{thm7} provides a rate of convergence.
\cite{Adell2003} also obtained a rate but under smoothness conditions on $F$, viz., second-order moduli of smoothness for continuous functions or functions of bounded variation.  Our bound is expressed in terms of the local oscillation of the distribution function and applies when no density exists.
\end{remark}




\begin{remark}[Necessity]\label{remk1}
Suppose Assumption \ref{A1} holds, $H\in L(\mathds{R})$, and
$\mathcal FH\geq0$. Then the following statements are equivalent:
\begin{enumerate}
    \item Assumption~\ref{AG} holds;
    \item
    $
    |E(\widehat{F}_{Y}(x))-F_Y(x)|\rightarrow 0
    $
    as $h,\,\lambda \rightarrow 0$,
    for every distribution function $F_Y$ and every
    $x\in C(F_Y)$;
    \item $H$ is such that
    \begin{equation}\label{73.1}
        \mathcal{F}H\in L(\mathds{R}),
        \quad \mbox{ and }  \quad
        \frac{1}{2\pi }\int_{\mathds{R}}(\mathcal{F}H)(u)du=1.
    \end{equation}
\end{enumerate}

The implication $1\Rightarrow2$ follows from Theorem~\ref{thm7}, while
$3\Rightarrow1$ follows from the nonnegativity of $\mathcal FH$.
To prove $2\Rightarrow3$, take $F_Y$ to be the distribution function
of a point mass at zero and let $x=1$.
Then,
\[
\int_{\mathds{R}} G_H\left( \left[ \frac{t-1}{h},\frac{t-1+1/\lambda }{h}\right] \right)
dF_Y(t)=G_H\left( \left[ -\frac{1}{h},\frac{-1+1/\lambda}{h}\right]
\right).
\]
By Lemma~\ref{lem1}, choosing
$h_n=\lambda_n=1/n$ gives
$
G_H([-n,n^2-n])
=
E[\widehat{F}_{Y}(1)]
\rightarrow 1
$
as $n\to \infty$.
Because
$\lim_{n\to\infty}[-n,n^2-n] = \mathds{R}$
and $\mathcal FH\geq0$,
the monotone convergence theorem yields
$
\frac{1}{2\pi}\int_{\mathds{R}}(\mathcal FH)(u)\,du=1,
$
and hence $\mathcal FH\in L(\mathds{R})$.
\end{remark}

The next theorem provides a bound for the variance of $\widehat{F}_{Y}(x)$ and gives a sufficient condition for the convergence of $\widehat{F}_{Y}(x)$ in quadratic mean.  This, in turn, implies consistency of $\widehat{F}_{Y}(x)$.
\begin{theorem}\label{thm8}
Suppose Assumption \ref{A1} holds. Then,
\begin{equation}
V\left(\widehat{F}_{Y}(x)\right) \le \frac{1}{n}\frac{v_{h,\lambda}^2}{4 \pi^2}, \quad \mbox{ where } \quad v_{h,\lambda }:=\int_{\mathds{R}}|\alpha _{h,\lambda }(t)|dt.
\end{equation}
If, in addition, Assumption \ref{AG} holds, and
$h:=h_{n} \to 0$, $\lambda :=\lambda _{n} \to 0$ are chosen such that $v_{h,\lambda }/\sqrt{n}\rightarrow 0$ as $n\rightarrow \infty $,
then,
for any distribution function $F_Y$ and any $x\in C(F_{Y})$, we have
$\widehat{F}_{Y}(x)\overset{p}{\rightarrow }F_{Y}(x)$.
\end{theorem}
The usefulness of Theorem \ref{thm8} depends on obtaining conditions under which $v_{h,\lambda}/\sqrt{n} \to 0$ as $n \to \infty$.  The following lemmas give conditions to obtain the order of $v_{h,\lambda}$.
 Lemma \ref{lem6a} provides results when $H$ has compact support, and Lemma \ref{lem6a_NoncompH} in Appendix \ref{app:noncompH} provides an analogous result when $H$ does not have compact support.
Lemmas \ref{lem6a} and \ref{lem6a_NoncompH} both obtain orders that rely solely  on conditions imposed on $H$.

\begin{lemma}\label{lem6a}
Suppose Assumption \ref{A1} holds, $\supp H \subset [-\gamma,\gamma]$ for some $\gamma >0$, and
$C_H:=\underset{t \in [-\gamma,\gamma]}{\sup}|H(t)|<\infty$.
Then, there exist constants $C>0$ and $\beta\in(0,\pi/2)$, depending only on $C_H$ and $\Phi_Z$, such that for all $\lambda \in (0,\beta/\pi)$ and $h\in(0,\gamma/\beta)$,
$$
 v_{h,\lambda}
 \leq
 C\left[ \log \lambda^{-1}+  \underset{ \beta < t\le\gamma/h}{\int} \frac{1}{t|\Phi_Z(t)|}dt\right].
 $$
\end{lemma}
The order of $\underset{ \beta < t\le\gamma/h}{\int} \frac{1}{t|\Phi_Z(t)|}dt$
depends on the behavior of $|\Phi_Z(t)|$ as $|t| \to \infty $.  Following \cite{Fan1991}, \cite{Dattner2011} and \cite{Dattner2013} we consider two cases:
    \begin{subequations}\label{eq:error_smoothness}
    \begin{align}
    \text{(super-smooth errors)} &\qquad  |\Phi_Z(t)| \asymp \exp(-\tau |t|^\rho), \qquad \text{as } |t|\to\infty, \label{eq:error_smoothness_super} \\
    \text{(ordinary-smooth errors)} & \qquad |\Phi_Z(t)| \asymp |t|^{-\rho}, \qquad \text{as } |t|\to\infty, \label{eq:error_smoothness_ordinary}
    \end{align}
    \end{subequations}
where $\tau>0$ and $\rho>0$ in \eqref{eq:error_smoothness_super}, and $\rho>0$ in \eqref{eq:error_smoothness_ordinary}.
Theorem \ref{thm10a} gives conditions on $\lambda$ and $h$ that ensure $\frac{v_{h,\lambda}}{\sqrt{n}}=o(1)$ when $H$ has compact support under ordinary or super-smooth errors.
Theorem \ref{thm11a} in Appendix \ref{app:noncompH} considers the case where $H$ is not compactly supported and errors are super-smooth.
Combined with Theorem \ref{thm8}, these provide sufficient conditions to ensure consistency.





\begin{theorem}\label{thm10a}
Suppose Assumption \ref{A1}.1 holds,
$\supp H \subset [-\gamma,\gamma]$ for some $\gamma >0$, and $\underset{t \in [-\gamma,\gamma]}{\sup}|H(t)|<\infty$.
Let
$h:=h_{n} \to 0$, $\lambda :=\lambda _{n} \to 0$
and
 suppose $\log(\lambda_n^{-1})=o(\sqrt{n})$.
\begin{enumerate}
\item[(a)]
If
 there exist $\tau, \rho>0$ such that
 \eqref{eq:error_smoothness_super} holds,
and
$h_n^{\rho}\log n\ge2\tau\gamma^{\rho} $
for all $n$ sufficiently large,
then $\frac{v_{h,\lambda}}{\sqrt{n}}=o(1)$.
\item[(b)]
If there exists $\rho>0$ such that
\eqref{eq:error_smoothness_ordinary} holds,
and
$\displaystyle \lim_{n\to\infty} n h_n^{2 \rho} = \infty$,
then $\frac{v_{h,\lambda}}{\sqrt{n}}=o(1)$.
\end{enumerate}
\end{theorem}


\begin{remark}The condition
$\log(\lambda_n^{-1})=o(\sqrt{n})$
in Theorems \ref{thm10a} and \ref{thm11a} is mild and met if $\lambda \asymp n^{-a}$ for any $a>0$.  The conditions on $h$ are more demanding.  For super-smooth errors and compactly supported $H$, a slowly vanishing sequence such as $h_n=(\log n)^{-a}$ with $0< a <\rho^{-1}$ suffices.  In the case of ordinary-smooth errors a polynomial rate of decay, such as $h_n=n^{-a}$ for $0< a<\frac{1}{2\rho}$ suffices.
The condition on $h$ in Theorem \ref{thm11a} for non-compact $H$ is strictly stronger than that in part (a) of Theorem \ref{thm10a}.
\end{remark}

It follows directly from Theorems \ref{thm7} and \ref{thm8} that the mean squared error (MSE) for $\widehat{F}_{Y}(x)$ satisfies
\begin{align*}
MSE(\widehat{F}_{Y}(x)) &\le  \Big[\phi _{G_H}(h^{\delta -1})+(1+\left\Vert G_H\right\Vert _{\infty})\omega_{F_Y} (x,h^{\delta })+(2+\left\Vert G_H\right\Vert _{\infty})F_{Y}(x+1-1/\lambda ) \Big]^2
+
\frac{1}{n}\frac{v_{h,\lambda}^{2}}{4\pi^{2}}.
\end{align*}
The following theorem provides bounds on the
MSE
of $\widehat{F}_{Y}(x)$.
\begin{theorem}\label{thm:FYx_MSEbnd}
Suppose Assumptions \ref{A1} and \ref{AG} hold,
$\supp H \subset [-\gamma,\gamma]$ for some $\gamma >0$, and  $\underset{t \in [-\gamma,\gamma]}{\sup}|H(t)|<\infty$.
Then, there exists $C>0$ and $\beta\in \left(0,\min\{\gamma,\pi/2\}\right)$
such that, for every $\delta \in (0,1)$, $h\in (0,1]$, and $\lambda\in(0,\beta/\pi)$,
the following bound holds for any distribution function $F_Y$ and any $x\in C(F_{Y})$
\begin{align*}
MSE(\widehat{F}_{Y}(x))
& \le
\Big[\phi _{G_H}(h^{\delta -1})+(1+\left\Vert
G_H\right\Vert _{C_b})\omega_{F_Y} (x,h^{\delta }) +(2+\left\Vert G_H\right\Vert _{\infty})F_{Y}(x+1-1/\lambda ) \Big]^2
\\&
\quad +
\frac{1}{n} \frac{C}{ 4\pi^{2}}
\left( \log \lambda^{-1}+
\int_\beta^{\gamma/h}\!\frac{1}{t|\Phi_Z(t)|}dt\right)^2
\end{align*}
\end{theorem}

The next result illustrates the rates implied by Theorem \ref{thm:FYx_MSEbnd} under additional
local regularity and tail assumptions on $F_Y$. These assumptions are not
needed for consistency, but they make the dependence of the bound on $n$
explicit.

\begin{corollary}\label{coro1}
Suppose the conditions of Theorem \ref{thm:FYx_MSEbnd} a) hold and that, for some
$k,s,q>0$ and $C<\infty$,
\[
\phi_G(N)\le CN^{-k}, \qquad
\omega_{F_Y}(x,u)\le Cu^s,\qquad
F_Y(x+1-u_1)\le C u_1^{-q},
\]
for all sufficiently large $N,\,u_1>0$ and small $u>0$.
Define $r=\frac{ks}{k+s}$.

\begin{enumerate}
\item[(a)]
If there exist $\tau, \rho>0$ such that
    \eqref{eq:error_smoothness_super} holds
then, for any constant
$
A>\gamma (2\tau)^{1/\rho},
$
with
$
h_n
=
A(\log n)^{-1/\rho},
\,
\lambda_n=h_n^{\,r/q},
$
we have
\[
\operatorname{MSE}(\widehat F_Y(x))
=
O\!\left(
(\log n)^{-2r/\rho}
\right).
\]


\item[(b)]
If there exists $\rho>0$ such that
    \eqref{eq:error_smoothness_ordinary} holds,
and $
h_n\asymp n^{-1/\{2(\rho+r)\}},
\,\lambda_n=h_n^{\,r/q},
$ then
\[
\operatorname{MSE}(\widehat F_Y(x))
=
O\!\left(
n^{-r/(\rho+r)}
\right).
\]
\end{enumerate}
\end{corollary}

\cite{Dattner2011} and \cite{ Dattner2013} derive minimax optimal convergence rates under global Sobolev smoothness assumptions on $F_Y$. In contrast, Corollary \ref{coro1} establishes explicit rates under substantially weaker assumptions consisting only of a local modulus of continuity for the distribution function together with mild local tail conditions. Although the resulting rates need not match the minimax Sobolev rates when both sets of assumptions hold, they apply to a considerably broader class of distributions, including settings where no global smoothness of the distribution is available.


\subsection{Estimation of $F_{Y}(x,y)$ for $x,y \in C(F_Y)$}\label{sec:FYxy}
Although the interval probabilities $F_Y(x,y):=F_Y(y)-F_Y(x)$ for $x<y$ and $x,y \in C(F_Y)$ can be estimated by differencing pointwise estimators of $F_Y$, a direct estimator has several advantages. It avoids differencing two noisy estimators, yields sharper non-asymptotic bounds, and provides a natural analogue of the generalized inversion estimator developed in Section \ref{sec:FYx}.

The estimator we propose is obtained by replacing the interval $(x-1/\lambda,x)$ used in Section \ref{sec:FYx} by the fixed interval $(x,y)$.  Let $\{X_j\}_{j=1}^n$ be a random sample from $F_X$. For $h>0$ define
\begin{equation*}
\widehat{F}_{Y}(x,y):=\frac{1}{n}\sum_{j=1}^{n}\!K_{x,y,h}\!\left(X_j-x\right)\!,
\;\, \mbox{where} \;\,
K_{x,y,h}(t)=\frac{1}{2\pi }\!\int_{\mathds{R}\!}\!e^{its}\beta _{x,y,h}(s)ds
,
\;\;
\beta _{x,y,h}(s)=\frac{1-e^{-i(y-x)s}}{is}\frac{H(hs)}{\Phi_{Z}(s)}\!.
\end{equation*}
Assumptions \ref{A1}.1 and \ref{A1}.2 imply $\beta_{x,y,h} \in L(\mathds{R})$ for each $h>0$ and $x,y \in \mathds{R}$,
so $K_{x,y,h}$ is bounded and continuous for each $h>0$.
Furthermore, under Assumption~\ref{A1}.3
\begin{equation*}
\overline{\beta_{x,y,h}(s)}=\frac{1-e^{i(y-x)s}}{-is}\frac{H(-hs)}{\Phi _{Z}(-s)}=\beta _{x,y,h}(-s), \;\;\; \mbox{ and } \;\;\; \overline{K_{x,y,h}(t)}=\frac{1}{2\pi }\int_{\mathds{R}}e^{-its}\beta_{x,y,h}(-s)ds=K_{x,y,h}(t),
\end{equation*}
so $K_{x,y,h}$ is real-valued.
Hence, under Assumption \ref{A1},  $\widehat{F}_{Y}(x,y)$ is continuous, almost surely bounded,  real-valued, and  $E\left( \widehat{F}_{Y}(x,y) \right)<\infty$.

The following result gives a finite-sample bias bound and asymptotic unbiasedness under similar conditions as Theorem \ref{thm7}. Notably, this result places no assumptions on the distribution $F_Y$.
\begin{theorem}\label{thm13n}
Suppose Assumptions \ref{A1} and \ref{AG} hold.
Then, for any $F_{Y},$ and $x,y\in C(F_{Y})$
\begin{equation} \label{unqual}
\left|E\left(\widehat{F}_{Y}(x,y)\right)-F_{Y}(x,y)\right|
\rightarrow 0,
\quad \mbox{ as $h,\,\lambda \rightarrow 0$.  }
\end{equation}
Moreover, for $\delta \in (0,1)$, $h \in (0,1]$ and $h^\delta \le \frac{y-x}{2}$
\begin{eqnarray}
\left|E\left(\widehat{F}_{Y}(x,y)\right)-F_{Y}(x,y)\right| &\leq &\phi _{G}(h^{\delta
-1})+(1+\left\Vert G\right\Vert _{\infty})[\omega_{F_Y}(x,h^{\delta })+\omega_{F_Y}
(y,h^{\delta })].  \label{qual1}
\end{eqnarray}
\end{theorem}

Following the same reasoning as in Remark \ref{remk1}, we have that if $\mathcal{F}H \geq 0$, then  $\widehat{F}_{Y}(x,y)$ is asymptotically unbiased if, and only if, Assumption \ref{AG} holds, which is, in turn, equivalent to the conditions in equation \eqref{73.1}.



The next theorem provides a bound for the variance of $\widehat{F}_{Y}(x,y)$ and gives a sufficient condition for the convergence of $\widehat{F}_{Y}(x,y)$ in quadratic mean.  This, in turn, implies consistency of $\widehat{F}_{Y}(x,y)$.
\begin{theorem}\label{thm8a}
Suppose Assumption \ref{A1} holds. Then,
\begin{equation}
V\left(\widehat{F}_{Y}(x,y)\right) \le \frac{1}{n}\frac{v_h(x,y)^2}{4 \pi^2},
\quad \mbox{ where } \quad
v_h(x,y):=\int_{\mathds{R}}|\beta _{x,y,h}(t)|dt
.
\end{equation}
If, in addition, Assumption \ref{AG} holds, and
$h:=h_{n} \to 0$ is chosen so that $v_{h}(x,y)/\sqrt{n}\rightarrow 0$ as $n\rightarrow \infty $,
then, for any distribution function $F_Y$ and any $x\in C(F_{Y})$, we have
$\widehat{F}_{Y}(x,y)\overset{p}{\rightarrow }F_{Y}(x,y)$.
\end{theorem}
The following lemma is similar to Lemma \ref{lem6a} and gives the order of $v_h(x,y)$ under conditions on $H$. Lemma~\ref{lem6ab_NoncompH} in Appendix~\ref{app:noncompH} gives a similar result for $H$ with non-compact support.
\begin{lemma}\label{lem6ab}
Suppose Assumption \ref{A1} holds, $\supp H \subset [-\gamma,\gamma]$ for some $\gamma >0$, and
$C_H:=\underset{t \in [-\gamma,\gamma]}{\sup}|H(t)|<\infty$.
Then, there exist constants $C,\beta>0$, depending only on $C_H$ and $\Phi_Z$, such that
for all $h\in(0,\gamma/\beta)$ and $x,y\in\mathds{R}$ with $x<y$,
$$
v_{h}(x,y) \leq C\left[ \log(1+y-x) + \underset{ \beta < t\le\gamma/h}{\int} \frac{1}{t|\Phi_Z(t)|}dt\right]\!.
$$
\end{lemma}
As in the case of $v_{h,\lambda}$, the order of $v_h(x,y)$
depends on the behavior of $|\Phi_Z(t)|$ as $|t| \to \infty $.
The next theorem gives conditions on $h$ to ensure that $\frac{v_{h}(x,y)}{\sqrt{n}}=o(1)$,
which can be used with Theorem \ref{thm8a} to obtain consistency.
Theorem \ref{thm11ab} gives a similar result when $H$ has non-compact support.

\begin{theorem}\label{thm10ab}
Suppose $\supp H \subset [-\gamma,\gamma]$ for some $\gamma >0$ and $\underset{t \in [-\gamma,\gamma]}{\sup}|H(t)|<\infty$.
\begin{enumerate}
\item[(a)]
If there exist $\tau, \rho>0$ such that
\eqref{eq:error_smoothness_super} holds,
and if $h \to 0,\,\displaystyle\liminf_{n\to\infty}h_n^{\rho}\log n>2\tau\gamma^{\rho}$, then $\frac{v_{h}(x,y)}{\sqrt{n}}=o(1)$.
\item[(b)]
If there exists $\rho>0$ such that
\eqref{eq:error_smoothness_ordinary} holds,
and if  $n h^{2 \rho}  \to \infty$, then $\frac{v_{h}(x,y)}{\sqrt{n}}=o(1)$.
\end{enumerate}
\end{theorem}


It follows directly from Theorems \ref{thm13n} and \ref{thm8a} that the MSE for $\widehat{F}_{Y}(x,y)$ at $x,y \in C(F_Y)$ satisfies
$$
MSE(\widehat{F}_{Y}(x,y)) \le  \Big[ \phi _{G}(h^{\delta
-1})+(1+\left\Vert G\right\Vert _{\infty})(\omega_{F_Y}(x,h^{\delta })+\omega_{F_Y}
(y,h^{\delta })) \Big]^2
+
\frac{1}{n}\frac{v_h(x,y)^2}{4 \pi^2}.
$$
With this, the following theorem follows immediately from the bound on $v_h(x,y)$ obtained in Lemma \ref{lem6ab}.
\begin{theorem}\label{thm:FYxy_MSEbnd}
Suppose Assumptions \ref{A1} and \ref{AG} hold,
$\supp H \subset [-\gamma,\gamma]$ for some $\gamma >0$, and  $\underset{t \in [-\gamma,\gamma]}{\sup}|H(t)|<\infty$.
Then, there exist constants $C>0$ and $\beta \in (0,\gamma)$ such that, for any distribution function $F_Y$, and any $\delta \in (0,1)$, any $x,y\in C(F_{Y})$ with $x<y$ and any $h \in (0,1]$ satisfying $h^\delta\leq (y-x)/2$,
the following bound holds
\begin{align*}
MSE(\widehat{F}_{Y}(x,y))
&\le
\Big[ \phi _{G_H}(h^{\delta
-1})+(1+\left\Vert G_H\right\Vert _{\infty})(\omega_{F_Y}(x,h^{\delta })+\omega_{F_Y}
(y,h^{\delta })) \Big]^2
\\&
\quad +
\frac{1}{n} \frac{C}{ 4\pi^{2}}
\left( \log(1+y-x) + \underset{ \beta < t\le\gamma/h}{\int} \frac{1}{t|\Phi_Z(t)|}dt\right)^2\!
.
\end{align*}
\end{theorem}


\subsection{Estimation of a jump $p_{x}$ of $F_{Y}$}\label{sec:px}
Unlike our estimators for $F_Y(x)$ and $F_Y(x,y)$, estimation of jump sizes does not follow from a straightforward application of the generalized inversion formula.
Instead, it requires a different approximation argument based on localized kernels.
Since \(p_x\) can be written as \(\mathbb P(Y=x)\), the idea is to approximate the singleton indicator \(u\mapsto\chi_{{x}}(u)\) by a localized kernel \(W((u-x)/h)\) for $h>0$.


To motivate our estimator,
consider a distribution function $F$ with a jump of size $p_x=F(x)-\underset{\varepsilon \downarrow 0}{\lim}\,F(x-\varepsilon)$ at $x$.
Suppose $W: \mathds{R} \to \mathds{R}$ such that
\begin{equation}\label{A3}
W\text{ is continuous on }\mathbb R,\qquad
W(0)=1,\qquad
\lim_{|u|\to\infty}W(u)=0.
\end{equation}
Equation \eqref{A3} implies
$W$ is bounded and
$\underset{h \to 0}{\lim}W\left(\frac{u-x}{h}\right)= \chi_{{x}}(u)$,
so bounded convergence gives
\begin{equation}\label{jumpmot}
    \lim_{h\to 0}
    \int_{\mathds{R}}W\left(\frac{u-x}{h}\right)dF(u)
    =
    p_x
    .
\end{equation}
To quantify the approximation, define
\begin{equation*}
\omega _{W}(\varepsilon )
:=
\sup_{\left\vert v\right\vert \leq \varepsilon}\left\vert W(v)-1\right\vert
,\qquad
\phi _{W}(N)
:=
\sup_{\left\vert v\right\vert\geq N}\left\vert W(v)\right\vert
,\quad\mbox{and}\quad
\delta _{F}(x,\varepsilon):=\int_{\left\vert t-x\right\vert <\varepsilon }dF(t)-p_{x}\geq 0.
\end{equation*}
When equation \eqref{A3} holds, Lemma 3 in \cite{Mynbaev2022} showed that, for any $h, \,\varepsilon _{1}\in (0,1)$ and $\varepsilon_{2}>\varepsilon_{1}$
\begin{equation}\label{25n}
\left\vert \int_{\mathds{R}\!}\!W\!\left(\! \frac{u-x}{h}\right)\!dF(u)-p_{x}\right\vert
\leq
\omega _{W}(h^{\varepsilon_{2}-\varepsilon _{1}})\!\left[ p_{x}+\delta _{F}(x,h^{1-\varepsilon _{1}})\right]
+
\left( 1+\left\Vert W\right\Vert _{C_b}\right) \delta _{F}(x,h^{1-\varepsilon_{1}})
+
\phi _{W}(h^{-\varepsilon _{1}}),
\end{equation}
and we also have
$\underset{\varepsilon \rightarrow 0}{\lim}\omega _{W}(\varepsilon )= 0,$
$\underset{N \rightarrow \infty}{\lim}\phi _{W}(N)= 0,$
and
$\underset{\varepsilon \rightarrow 0}{\lim}\delta_{F}(x,\varepsilon )=0$
by continuity of probability measures.
Thus, unlike the bounds in Sections \ref{sec:FYx} and \ref{sec:FYxy}, the approximation error is governed by the local probability mass around the discontinuity rather than by a modulus of continuity.

We now connect this to a Fourier representation.
Let $H \in L(\mathds{R})$ be real-valued and even, and set $W:=\mathcal{F}H$.
If $H$ is a kernel, i.e., $\int_{\mathds{R}}H(t)dt=1$, then $\mathcal{F}H$ satisfies equation \eqref{A3} since $(\mathcal FH)(0)=1$ and $\mathcal FH$ is continuous and vanishes at infinity by the Riemann-Lebesgue lemma.
In addition, Fubini's Theorem gives the following relationship
\begin{equation}\label{e77n}
\int_{\mathds{R}}W\left(\frac{u-x}{h}\right)dF(u)
=
\int_{\mathds{R}}(\mathcal{F}H)\left(\frac{u-x}{h}\right)dF(u)
=
\int_{\mathds{R}}e^{-isx}(\mathcal{F}F)(s)hH(hs)ds.
\end{equation}






Equation \eqref{e77n} motivates the following estimator for $p_x:=F_Y(x)-\underset{\varepsilon \downarrow 0}{\lim}\,F_Y(x-\varepsilon)$.
Let $\{X_j\}_{j=1}^n$ be a random sample from $F_X$, and for $h>0$ define
\begin{equation*}
\hat{p}_{x}=\frac{1}{n}\sum_{j=1}^nK_{h}\!\left(X_{j}-x\right)\!,
\quad \mbox{where} \quad
K_{h}(t):=\int_{\mathds{R}}e^{ist}\gamma_{h}(s)ds
,
\quad
\gamma _{h}(s)=\frac{hH(hs)}{\Phi _{Z}(s)}.
\end{equation*}
To ensure $\hat{p}_{x}$ is well defined, we use Assumptions \ref{A1}.1 and \ref{A1}.3 again, but we replace Assumption \ref{A1}.2 with the following assumption.
\begin{condition}\label{A2}
(Regularity)
$\gamma _{h} \in L(\mathds{R})$ for each $h>0$.
\end{condition}
Then, under Assumptions \ref{A1}.1, \ref{A1}.3  and \ref{A2}, $K_{h}$ is bounded, continuous and real valued with
$
\overline{K_{h}(t)}=\int_{\mathds{R}}e^{-ist}\frac{hH(-hs)}{\Phi _{Z}(-s)}ds=K_{h}(t).
$
Consequently, $\hat{p}_{x}$ is also bounded, continuous and real valued.


The next theorem characterizes the asymptotic unbiasedness of $\hat{p}_x$.
Remarkably, asymptotic unbiasedness of $\hat{p}_x$ is equivalent to $\int_{\mathds{R}}H(t)dt=1$.
In addition, we obtain a non-asymptotic bound on the bias, since (see the proof of Theorem \ref{thm11})
$$
E(\hat{p}_{x})=\int_{\mathds{R}}(\mathcal{F}H)\left( \frac{u-x}{h}\right)dF_Y(u).
$$

\begin{theorem}\label{thm11}
Suppose $H\in L(\mathds{R})$, and Assumptions \ref{A1}.1, \ref{A1}.3, and \ref{A2} hold. Then the following statements are equivalent:
\begin{enumerate}
\item $\int_{\mathds{R}}H(t)dt=1$;

\item $E(\hat{p}_{x}) \rightarrow p_{x}$ as $h \to 0$ for any $F_{Y}$ and $x\in J(F_{Y})$;

\item
$W=\mathcal{F}H$ satisfies equation \eqref{A3}.
\end{enumerate}
Moreover, if these equivalent conditions hold, then, for every distribution function $F_{Y}$ and $x\in J(F_{Y})$, the following bias bound holds whenever
$h,\,\varepsilon _{1}\in (0,1)$ and $\varepsilon_{2}>\varepsilon _{1}$,
\begin{eqnarray*}
\left\vert E(\hat{p}_{x}) - p_{x} \right\vert
&\leq &
\omega _{\mathcal{F}H}(h^{\varepsilon_{2}-\varepsilon _{1}})\left[ p_{x}+\delta_{F_Y}(x,h^{1-\varepsilon _{1}})\right]
+
\left( 1+\left\Vert \mathcal{F}H\right\Vert _{C_b}\right) \delta_{F_Y}(x,h^{1-\varepsilon_{1}})
+
\phi _{\mathcal{F}H}(h^{-\varepsilon _{1}}).
\end{eqnarray*}
\end{theorem}


Theorem~\ref{thm11} establishes that $\hat p_x$ is asymptotically unbiased as $h\to 0$ under conditions only on the error distribution and localization kernel $H$. These place no additional restrictions on $F_Y$ and do not require the existence of a density,
a mixture representation, or global smoothness assumptions on the latent distribution.

\begin{remark}\label{rem5}
The normalizations for $H$ differ across the three estimators.
For $\widehat{F}_{Y}(x)$ and $\widehat{F}_{Y}(x,y)$ the sufficient conditions given in Remark \ref{rem:inversion_kernel}
include $H(0)=1$.
By contrast, Theorem~\ref{thm11} shows $\int_{\mathds{R}}H(t)dt=1$, i.e. $(\mathcal{F} H)(0)=1$, is necessary and sufficient for asymptotic unbiasedness of $\hat p_x$ for any distribution function $F_Y$ at every $x\in J(F_Y)$.
Note, e.g., that the Gaussian kernel
$H(t)=e^{-t^{2}/2}$ satisfies the first normalization but not the second.  Relatedly, $K_{h,\lambda}$ and
$K_{x,y,h}$ carry the factor $1/(2\pi)$ while $K_h$ does not.
\end{remark}

The next theorem provides a bound for the variance $V(\hat{p}_x)$. This, together with Theorem \ref{thm11}, gives sufficient conditions for convergence in quadratic mean of $\hat{p}(x)$, which implies consistency.
\begin{theorem}\label{thm12}
Given Assumptions \ref{A1}.1, \ref{A1}.3 and \ref{A2}
\begin{equation}
V(\hat{p}_x)\le \frac{1}{n}u_h^2, \mbox{ where $u_{h}=\int_{\mathds{R}}|\gamma _{h}(t)|dt.$}
\end{equation}
Suppose $h=h_{n}\to 0$ is chosen so that $u_{h}/\sqrt{n}\rightarrow 0$ as $n\rightarrow \infty$.  Then, under the conditions of Theorem \ref{thm11}, $\hat{p}_{x} \overset{p}{\rightarrow} p_x$.
\end{theorem}
Similar to results in the previous two sections, the next theorem gives conditions on $h$ to ensure that $\frac{u_{h}}{\sqrt{n}}=o(1)$ under ordinary and super smoothness, but in this case, only for compactly supported $H$.

\begin{theorem}\label{thm10j}
Suppose Assumption \ref{A1}.1 holds, $\supp H \subset [-\gamma,\gamma]$ for some $\gamma >0$, and $\underset{t \in [-\gamma,\gamma]}{\sup}|H(t)|<\infty$. Let $h:=h_{n} \to 0$.
\begin{enumerate}
    \item[(a)]
    If there exist $\tau, \rho>0$ such that
    \eqref{eq:error_smoothness_super} holds,
    and $h_n^{\rho}\log n\ge2\tau\gamma^{\rho}$ for all $n$ sufficiently large,
    then $\frac{u_{h}}{\sqrt{n}}=o(1)$.
    \item[(b)]
     If there exists $\rho>0$ such that
    \eqref{eq:error_smoothness_ordinary} holds,
     and $\displaystyle \lim_{n\to\infty} n h_n^{2 \rho}  = \infty$, then $\frac{u_{h}}{\sqrt{n}}=o(1)$.
\end{enumerate}
\end{theorem}
It follows directly from Theorems \ref{thm11} and \ref{thm12} that the mean squared error for $\hat{p}_x$ at $x \in J(F_Y)$ satisfies
$$
MSE(\hat{p}_x)
\le
\Big[ \omega _{\mathcal{F}H}(h^{\varepsilon_{2}-\varepsilon _{1}})\left( p_{x}+\delta _{F_Y}(x,h^{1-\varepsilon _{1}})\right)  +\left( 1+\left\Vert \mathcal{F}H\right\Vert _{C_b}\right) \delta _{F_Y}(x,h^{1-\varepsilon_{1}}) \notag\\
+\phi _{\mathcal{F}H}(h^{-\varepsilon _{1}}) \Big]^2
+
\frac{1}{n}u_h^2
.
$$
Given the bound on $u_h$ and equation \eqref{eq10j} from the proof of Theorem \ref{thm10j} we have the following theorem, which is stated without proof.
\begin{theorem}\label{thm14}
Suppose the conditions of Theorem \ref{thm11} hold, $\supp H \subset [-\gamma,\gamma]$ for some $\gamma >0$, and
$\underset{t \in [-\gamma,\gamma]}{\sup}|H(t)|<\infty$.
Then,
there exist constants $C, \, \beta>0$ such that,
for every distribution function $F_{Y}$ and $x\in J(F_{Y})$, the following holds whenever
$h,\,\varepsilon _{1}\in (0,1)$ and $\varepsilon_{2}>\varepsilon _{1}$,
\begin{align*}
MSE(\hat{p}_x)
&\le
\Big[
    \omega _{\mathcal{F}H}(h^{\varepsilon_{2}-\varepsilon _{1}})\left( p_{x}+\delta _{F_Y}(x,h^{1-\varepsilon _{1}})\right)
    +
    \left( 1+\left\Vert \mathcal{F}H\right\Vert _{C_b}\right) \delta _{F_Y}(x,h^{1-\varepsilon_{1}})
    +\phi _{\mathcal{F}H}(h^{-\varepsilon _{1}})
\Big]^2
\\&
\quad
+
C \frac{h^2}{n} \left(  \underset{ \beta < t\le\gamma/h}{\int} \frac{1}{|\Phi_Z(t)|}dt \right)^2
.
\end{align*}
\end{theorem}
\noindent Under additional assumptions on $\mathcal{F}H$ and $\delta_{F_Y}$, the following corollary obtains explicit rates of decay for $MSE(\hat{p}_x)$.

\begin{corollary}\label{coro2}
Suppose the conditions of Theorem \ref{thm11} hold, $\supp H \subset [-\gamma,\gamma]$ for some $\gamma >0$, and
$\underset{t \in [-\gamma,\gamma]}{\sup}|H(t)|<\infty$.
Fix a distribution function $F_Y$ and $x\in J(F_Y)$. Assume that, for some $k,m>0$ and $C<\infty$
$$
\varphi_{\mathcal FH}(N)\leq CN^{-k},
\qquad
\omega_{\mathcal FH}(\varepsilon)\leq C\varepsilon^m,
\quad \text{and}\quad
\delta_{F_Y}(x,\eta)\leq\psi(\eta),
$$
for all sufficiently large $N$ and all sufficiently small $\varepsilon,\eta>0$, where $\psi$ is nondecreasing with $\psi(0^+)=0$.
Fix $\varepsilon_1\in(0,1)$ and $\varepsilon_2>\varepsilon_1$, and define
$c:=\min\!\left\{m(\varepsilon_2-\varepsilon_1),\,k\varepsilon_1\right\}.$
\begin{enumerate}
    \item[(a)]
    Suppose there exist $\tau, \rho>0$ such that \eqref{eq:error_smoothness_super} holds,
    and take $h_n=A(\log n)^{-1/\rho}$ for some $A^{\rho}>4\tau\gamma^{\rho}$. Then,
\[
 \mathrm{MSE}(\hat p_x)
 =O\!\Big((\log n)^{-2c/\rho}
       +\psi\big(A^{1-\varepsilon_1}(\log n)^{-(1-\varepsilon_1)/\rho}\big)^{2}\Big).
\]
In particular, if $\psi(\eta)= C\eta^{s}$ for some $s>0$, take $\varepsilon_1=\tfrac{s}{k+s}$,
$\varepsilon_2=1$. If $m\ge s$, then $c=r:=\tfrac{ks}{k+s}$, and
\[
 \mathrm{MSE}(\hat p_x)=O\!\big((\log n)^{-2r/\rho}\big).
\]
\item[(b)]
Suppose there exists $\rho>0$ such that \eqref{eq:error_smoothness_ordinary} holds and that $\psi(\eta)= C\eta^{s}$ for some $s>0$.
Take $\varepsilon_1=\tfrac{s}{k+s}$, $\varepsilon_2=1$, and $r:=\tfrac{ks}{k+s}$.
If $m\geq s$ and $h_n\asymp n^{-1/\{2(r+\rho)\}}$ then
\[
 \mathrm{MSE}(\hat p_x)=O\!\big(n^{-r/(r+\rho)}\big).
\]
\end{enumerate}

\end{corollary}





The estimator proposed in this section extends the scope
of existing work on the estimation of an arbitrary jump under contaminated sampling.
Whereas prior work (e.g., \citealp{vanEs2008,gugushvili2011,Lee2013})
has largely relied on a specified discrete-continuous mixture structure for $F_Y$ and global smoothness conditions, the present estimator applies to an arbitrary jump of a completely general
distribution function.
We do not pursue pursue minimax optimality; in exchange, our methodology applies to a much broader class of latent distributions.



\section{Implementation and simulation study}\label{sec:sims}




In this section we conduct simulation studies for the estimators proposed in this paper: $\widehat{F}_Y(x)$, $\widehat{F}_Y(x,y)$, and  $\hat{p}_x$.
For each estimator, we consider multiple feasible methods for tuning-parameter selection and compare performance across different sample sizes and distributions for $Y$ and $Z$.


\subsection{Estimation of ${F}_Y(x)$}\label{sec:FYxSims}
This subsection studies the finite-sample performance of the proposed estimator of $F_Y(x)$ for $x \in C(F_Y)$.





We consider five distributions for $Y$: two smooth distributions and three non-smooth distributions with kink points. The smooth distributions are the standard normal distribution $Y\sim N(0,1)$ and the Gamma distribution $Y\sim\Gamma(3, 1/\sqrt{3})$, as considered by  \citet{Dattner2013} and \citet{Hall2008}.
The non-smooth distributions are the one-kink design $Y\sim K_1$, the gap-kink design $Y\sim K_2$, and the asymmetric piecewise-uniform design $Y\sim K_3$, shown in Figure \ref{fig:kink-dgps}.
The precise definitions of these distributions are given in Appendix~\ref{app:FYkink_def}.



\begin{figure}[h]
\centering
\makebox[\textwidth][c]{
\begin{tikzpicture}


\pgfmathsetmacro{\PhiMinusOne}{normcdf(-1)}
\pgfmathsetmacro{\Cgap}{2*\PhiMinusOne}

\begin{groupplot}[
    group style={group size=3 by 1, horizontal sep=1.35cm},
    width=0.35\textwidth,
    ymin=0, ymax=1,
    samples=200,
    axis lines=box,
    grid=major,
    grid style      =   {gray!30, line width=0.15pt},
    tick style      =   {gray!30, line width=0.20pt},
    axis line style =   {gray!75, line width=0.20pt},
    every axis plot/.append style={
        black,
        line width=0.55pt,
        mark=none,
        smooth
    },
    xlabel={$x$},
    ylabel={$F_Y(x)$},
    title style={font=\small},
    label style={font=\tiny},
    ticklabel style={font=\tiny},
    ytick={0,0.2,0.4,0.6,0.8,1},
    scaled ticks=false,
    enlargelimits=false,
    clip=true
]

\nextgroupplot[
    title={$Y\sim K_1$},
    xmin=-3, xmax=3,
    xtick={-3,-1.5,0,1.5,3}
]

\addplot[domain=-3:0] {3*normcdf(x)/2};

\addplot[domain=0:3] {(1+normcdf(x))/2};

\nextgroupplot[
    title={$Y\sim K_{\mathrm{2}}$},
    xmin=-3, xmax=3,
    xtick={-3,-1,0,1,3}
]

\addplot[domain=-3:-1] {normcdf(x)/\Cgap};

\addplot[domain=-1:1] {\PhiMinusOne/\Cgap};

\addplot[domain=1:3] {
    (\PhiMinusOne + normcdf(x) - normcdf(1))/\Cgap
};

\nextgroupplot[
    title={$Y\sim K_{3}$},
    xmin=-2.25, xmax=2.25,
    xtick={-2,-0.5,0.25,2},
    xticklabels={$-2$,$-0.5$,$0.25$,$2$}
]

\addplot[domain=-2.25:-2] {0};

\addplot[domain=-2:-0.5] {0.20*(x+2)/1.5};

\addplot[domain=-0.5:0.25] {0.20 + 0.60*(x+0.5)/0.75};

\addplot[domain=0.25:2] {0.80 + 0.20*(x-0.25)/1.75};

\addplot[domain=2:2.25] {1};

\end{groupplot}
\end{tikzpicture}

}
\caption{Kink distribution functions used in the simulation study.}
\label{fig:kink-dgps}
\end{figure}


We consider three measurement-error distributions for $Z$, using rescaled
versions of error families considered in the simulation studies of
\citet{Dattner2013} and \citet{Hall2008}. The first is the normal distribution
$Z\sim N(0,1/5)$, where $1/5$ denotes the standard deviation. The second is
the Gamma distribution
$Z\sim\Gamma\left(2,1/(5\sqrt{2})\right)$, with shape $2$ and scale
$1/(5\sqrt{2})$.
The third is the symmetrized Gamma distribution from
\citet{Hall2008}.
We denote this as $Z\sim \text{S}\Gamma\left(2,1/(5\sqrt{2})\right)$ when $Z$ is distributed as $U_1 - U_2$ where $U_1$, $U_2$ are independent
$\Gamma\left(1,1/(5\sqrt{2})\right)$ random variables.




These distributions allow us to vary error asymmetry and smoothness while
holding the error variance fixed. The normal and symmetrized-Gamma errors are
symmetric, whereas the Gamma errors are asymmetric; the normal errors are
supersmooth, whereas the two Gamma-based errors are ordinary-smooth.
All three error distributions have standard deviation $\sigma_Z=1/5$.
Consequently, for a fixed $Y$ distribution, the contamination ratio $\sigma_Z/\sigma_Y$ is the same across the three $Z$ distributions.
Across the $Y$ distributions, this ratio is
$0.20$ for $Y\sim N(0,1)$ and
$Y\sim\Gamma(3,1/\sqrt{3})$, $0.22$ for $Y\sim K_1$, $0.13$ for
$Y\sim K_2$, and $0.24$ for $Y\sim K_3$.






The five distributions of $Y$ and three distributions of $Z$ yield 15
measurement-error designs. We consider each design at sample sizes
$n\in\{200,1000,5000\}$, giving 45 simulation settings, and generate 500
independent samples for each setting. For every sample, we estimate $F_Y(x)$
at a collection of evaluation points. For the smooth distributions of $Y$, we
use the five quantiles satisfying
$
F_Y(x)\in\{0.1,0.3,0.5,0.7,0.9\}.
$
For the non-smooth distributions, we focus on their kink points: $x=0$ for
$Y\sim K_1$; $x\in\{-1,1\}$ for $Y\sim K_2$; and
$x\in\{-1/2,1/4\}$ for $Y\sim K_3$.
The non-smooth designs therefore assess estimator performance at points where
$F_Y$ is not differentiable, whereas the smooth designs assess performance at
points where $F_Y$ is continuously differentiable.


In all of the above designs we report two feasible implementations of $\widehat{F}_{Y}(x)$, which choose $(h,\lambda)$ by different procedures.
Both implementations use the Fourier kernel
    \begin{equation}\label{eq:Hsimsdef}
        H(t) \;=\;
        \begin{cases}
        (1-t^2)^{4}, & |t| \leq 1, \\[1.2em]
        0, & |t| > 1.
        \end{cases}
    \end{equation}

The first implementation, denoted by $\widehat{F}_{Y}^{AR}(x)$, chooses $(h,\lambda)$ using a simple asymptotic rate rule motivated by Corollary \ref{coro1}.
Applying the
corollary with $k=4$ and $s=q=1$ gives
\[
(h_n,\lambda_n)=
\begin{cases}
\left(\dfrac{1}{4}n^{-5/28},\,h_n^{4/5}\right),
    &
    \text{if }
    Z\sim\Gamma\left(2,\frac{1}{5\sqrt{2}}\right)
    \text{ or }
    Z\sim\mathrm{S}\Gamma\left(2,\frac{1}{5\sqrt{2}}\right)\!,
    \\[0.8em]
\left(\dfrac{1.01}{5}(\log n)^{-1/2},\,h_n^{4/5}\right),
    &
    \text{if } Z\sim N\left(0,\frac{1}{5}\right)\!
    ,
\end{cases}
\]
where the factor $1.01$ ensures that the super-smooth-error rule satisfies
$
h_n
>\gamma (2\tau/\log(n))^{1/\rho}
.
$

The second implementation, denoted $\widehat{F}_{Y}^{CV}(x)$, chooses $(h,\lambda)$ using a
reconvolution cross-validation procedure. Since $Y$ is latent, each candidate $(h,\lambda)$ is evaluated by reconvolving $\widehat{F}_{Y}$ with the known distribution of $Z$ and comparing the resulting estimate of $F_X$ with the empirical distribution of the observed $X$-sample. The criterion includes a
variance penalty term based on Theorem \ref{thm8}.
We fix the variance-penalty coefficient for $\widehat{F}_{Y}^{CV}(x)$ at $\kappa=0.0025$ across all DGPs and sample sizes.
A detailed description of $\widehat{F}_{Y}^{CV}(x)$ is given in Appendix~\ref{app:FYx_CV}.
Table~\ref{tab:FYx_CV_sensitivity} in Appendix~\ref{app:Full_sim_results} reports results for $\kappa\in\{0.001,0.0025,0.005\}$, which indicate the overall performance is relatively stable across alternative $\kappa$ choices.




For comparison, we include two deconvolution estimators from the existing
literature. The first, $\widehat{F}_{DR}(x)$, is the adaptive estimator of
\citet[Section~2.2]{Dattner2013}.
The second, $\widehat{F}_{HL}(x)$, is the estimator of
\citet{Hall2008}, with bandwidth selected using the normal-reference procedure
described in their Section~4.2.
Implementation details for
$\widehat{F}_{DR}(x)$ and $\widehat{F}_{HL}(x)$ are provided in
Appendix~\ref{app:SimsEstImplDet}.












\subsubsection{Numerical results}
{
\setlength{\tabcolsep}{9pt}
\begin{table}[h!t]
\caption{
    Monte Carlo MSE$\times 100$ and variance$\times 100$ (in parentheses) for estimators of $F_Y(x)$.
    }\label{tab:FYx_main}
\adjustbox{max width=\textwidth, center = \textwidth}{
\small
\begin{tabular}{lcccc}
\toprule
&
$\widehat{F}_{Y}^{CV}(x)$ &
$\widehat{F}_{Y}^{AR}(x)$ &
$\widehat{F}_{DR}(x)$ &
$\widehat{F}_{HL}(x)$ \\
\midrule
\multicolumn{1}{l}{\rule[-.10ex]{0pt}{3.25ex}\!Smooth $F_Y$, $Z\!\sim\! N\!\left(0,\frac{1}{5}\right)\!$} &  &  &  &  \\
\quad $n=200$ & 0.095 (0.082) & 0.080 (0.075) & \textbf{0.079 (0.074)} & 0.091 (0.090) \\
\quad $n=1000$ & 0.018 (0.017) & 0.017 (0.015) & \textbf{0.015 (0.014)} & 0.018 (0.018) \\
\quad $n=5000$ & 0.005 (0.004) & 0.005 (0.003) & \textbf{0.004 (0.003)} & 0.004 (0.004) \\
\multicolumn{1}{l}{\rule[-.10ex]{0pt}{3.25ex}\!Smooth $F_Y$, $Z\!\sim\! \Gamma\!\left(2,\frac{1}{5\sqrt{2}}\right)\!$} &  &  &  &  \\
\quad $n=200$ & 0.087 (0.074) & 0.071 (0.063) & \textbf{0.069 (0.064)} & 0.079 (0.078) \\
\quad $n=1000$ & 0.018 (0.017) & 0.017 (0.015) & \textbf{0.015 (0.014)} & 0.017 (0.017) \\
\quad $n=5000$ & 0.004 (0.004) & 0.004 (0.003) & \textbf{0.004 (0.003)} & 0.004 (0.004) \\
\multicolumn{1}{l}{\rule[-.10ex]{0pt}{3.25ex}\!Smooth $F_Y$, $Z\!\sim\! \mathrm{S}\Gamma\!\left(2,\frac{1}{5\sqrt{2}}\right)\!$} &  &  &  &  \\
\quad $n=200$ & 0.093 (0.081) & 0.077 (0.070) & \textbf{0.076 (0.071)} & 0.092 (0.092) \\
\quad $n=1000$ & 0.018 (0.017) & 0.017 (0.015) & \textbf{0.015 (0.014)} & 0.018 (0.018) \\
\quad $n=5000$ & 0.004 (0.004) & 0.004 (0.003) & \textbf{0.003 (0.003)} & 0.004 (0.004) \\
\midrule
\multicolumn{1}{l}{\rule[-.10ex]{0pt}{3.25ex}\!Non-smooth $F_Y$, $Z\!\sim\! N\!\left(0,\frac{1}{5}\right)\!$} &  &  &  &  \\
\quad $n=200$ & 0.361 (0.144) & 0.443 (0.092) & 0.373 (0.104) & \textbf{0.267 (0.109)} \\
\quad $n=1000$ & \textbf{0.162 (0.037)} & 0.300 (0.019) & 0.234 (0.020) & 0.171 (0.022) \\
\quad $n=5000$ & \textbf{0.113 (0.008)} & 0.232 (0.004) & 0.183 (0.004) & 0.131 (0.005) \\
\multicolumn{1}{l}{\rule[-.10ex]{0pt}{3.25ex}\!Non-smooth $F_Y$, $Z\!\sim\! \Gamma\!\left(2,\frac{1}{5\sqrt{2}}\right)\!$} &  &  &  &  \\
\quad $n=200$ & 0.325 (0.148) & 0.516 (0.082) & 0.376 (0.098) & \textbf{0.247 (0.100)} \\
\quad $n=1000$ & \textbf{0.091 (0.030)} & 0.271 (0.018) & 0.230 (0.019) & 0.144 (0.020) \\
\quad $n=5000$ & \textbf{0.036 (0.007)} & 0.151 (0.004) & 0.185 (0.004) & 0.099 (0.005) \\
\multicolumn{1}{l}{\rule[-.10ex]{0pt}{3.25ex}\!Non-smooth $F_Y$, $Z\!\sim\! \mathrm{S}\Gamma\!\left(2,\frac{1}{5\sqrt{2}}\right)\!$} &  &  &  &  \\
\quad $n=200$ & 0.342 (0.165) & 0.540 (0.095) & 0.399 (0.112) & \textbf{0.227 (0.123)} \\
\quad $n=1000$ & \textbf{0.087 (0.030)} & 0.269 (0.018) & 0.230 (0.020) & 0.107 (0.022) \\
\quad $n=5000$ & \textbf{0.034 (0.007)} & 0.150 (0.004) & 0.184 (0.003) & 0.073 (0.004) \\
\bottomrule
\end{tabular}

}

    
    \smallskip
    \footnotesize
    \noindent
    \textit{Note:} Boldface identifies the estimator with
        the smallest unrounded MSE in each row.
        Smooth $F_Y$ includes
        $Y\sim N(0,1)$ and $Y\sim\Gamma(3,1/\sqrt{3})$; non-smooth $F_Y$ includes
        $Y\sim K_1$, $K_2,$ and $K_3$.
        Within each $Y$ design,
        empirical MSE and variance are first averaged over the
        evaluation points; these design-level quantities are then averaged
        across the $Y$ designs in each group and multiplied by 100 for readability.
        Full DGP-by-DGP results are reported in Tables \ref{tab:FYx_app_smth} and \ref{tab:FYx_app_nonsmth} in Appendix~\ref{app:Full_sim_results}.

\end{table}
}


Table \ref{tab:FYx_main} reports empirical MSE and variance for the four estimators of $F_Y(x)$ across the smooth and
non-smooth $Y$ distribution designs. MSE decreases with the sample size for every
estimator, but the relative rankings differ markedly between the smooth and
non-smooth designs.


For the smooth designs, $\widehat{F}_{DR}(x)$ has the smallest grouped MSE
for every measurement-error distribution and sample size. However, its
advantage is generally modest, with $\widehat{F}_{Y}^{AR}(x)$ also performing well.



When estimating non-smooth $F_Y$ at kink points, the ranking reverses.
Across all three error distributions,
$\widehat{F}_{HL}(x)$ has the smallest grouped MSE when $n=200$, and
$\widehat{F}_{Y}^{CV}(x)$ has the smallest grouped MSE when $n=1000$ and $n=5000$.
The advantage of $\widehat{F}_{Y}^{CV}(x)$ is most evident under
ordinary-smooth errors.
At $n=5000$, for example,
its MSE is approximately $14\%$, $64\%$, and $53\%$ smaller than that of
$\widehat{F}_{HL}(x)$
under normal, Gamma, and symmetrized-Gamma errors,
respectively.
Table~\ref{tab:FYx_app_nonsmth} in Appendix~\ref{app:Full_sim_results} reports results for each $Y$ distribution separately,
which shows that this pattern is not driven by a single non-smooth distribution: $\widehat{F}_{Y}^{CV}(x)$ has the smallest MSE in all nine non-smooth DGP--error combinations when $n=5000$, and in seven of the nine combinations when $n=1000$.



Because MSE and variance are reported on the same scale, their difference
equals squared bias.
Bias is relatively small in the smooth designs, and the lower-variance estimators tend to perform best.
At the kink points, squared bias accounts for a much larger share of MSE. In these cases, $\widehat{F}_{Y}^{CV}(x)$ trades higher variance for a reduction in squared bias, resulting in lower MSE once the sample
size is sufficiently large.
Importantly, the variance of $\widehat{F}_{Y}^{CV}(x)$ is not uniformly larger across the simulation designs, so its performance
at the kink points does not simply reflect a generally more variable estimator.



\subsection{Estimation of ${F}_Y(x,y)$}
This subsection studies the finite-sample performance of the proposed estimator of the interval probability $F_Y(x,y)$ for $x<y$ and $x,y \in C(F_Y)$.
We use the same distributions for $Y$ and $Z$, sample sizes, and number of Monte Carlo replications as in Section \ref{sec:FYxSims}.

For the two smooth $Y$ distributions, we estimate $F_Y(x,y)$ at the points
$
x=F_Y^{-1}\!\left(0.5-{p}/{2}\right),
$
and
$
y=F_Y^{-1}\!\left(0.5+{p}/{2}\right),
$
for $p\in\{0.05,0.20,0.80\}$,
so that $F_Y(x,y)=p$.
 For the non-smooth distributions, we use the following
interval pairs:
    $
    (x,y)\in
    \left\{
        (-1,0),\ (0,1),\
        \left(-\frac32,\frac32\right)
    \right\}
    $
    for $Y\sim K_1$;
    $
    (x,y)\in
    \left\{
        \left(-1,\frac32\right),\
        \left(-\frac32,1\right),\
        \left(-\frac32,\frac32\right)
    \right\}
    $
    for $Y\sim K_2$;
    and
    $
    (x,y)\in
    \left\{
        \left(-\frac54,-\frac12\right),\
        \left(-\frac12,\frac14\right),\
        \left(\frac14,\frac98\right)
    \right\}
    $
    for $Y\sim K_3$.
Seven of these nine intervals have at least one kink point as an endpoint;
the remaining two are wider intervals that span the non-smooth region. The
design therefore examines interval-probability estimation both when an
endpoint is a point of non-differentiability and when the interval contains
non-smooth features in its interior.


We consider two feasible implementations of $\widehat{F}_{Y}(x,y)$ which are analogous to those considered in Section \ref{sec:FYxSims} and use the same $H$ defined as in \eqref{eq:Hsimsdef}.
The first implementation, denoted
$\widehat{F}_{Y}^{AR}(x,y)$, uses the rate-based bandwidth
\[
h_n=
\begin{cases}
\dfrac{1}{5}n^{-5/28},
    & Z\sim\Gamma\left(2,\dfrac{1}{5\sqrt{2}}\right)
      \text{ or }
      Z\sim\mathrm{S}\Gamma\left(2,\dfrac{1}{5\sqrt{2}}\right),\\[1em]
\dfrac{0.8}{5}(\log n)^{-1/2},
    & Z\sim N\left(0,\dfrac{1}{5}\right).
\end{cases}
\]
The second, denoted $\widehat F_Y^{CV}(x,y)$, chooses $h$ by a similar reconvolution cross-validation procedure as before; each candidate $h$ is evaluated by reconvolving the corresponding estimate with the known distribution of $Z$ and comparing the resulting estimate of interval probabilities for $X$ with empirical interval probabilities from the observed $X$-sample.
The CV criterion includes a variance penalty term following from Theorem \ref{thm8a}, and
we fix the variance-penalty weight for $\widehat{F}_{Y}^{CV}(x,y)$ at $\kappa=0.01$ across all designs.
Table~\ref{tab:FYxy_CV_sensitivity} in Appendix~\ref{app:Full_sim_results} reports results for $\kappa\in\{0.005,0.010,0.0125\}$, which indicate that overall performance is relatively stable across alternative $\kappa$ choices.
A detailed description of $\widehat{F}_{Y}^{CV}(x,y)$ is provided in Appendix~\ref{app:FYxy_CV}.



Although the estimators of \citet{Dattner2013}  and \citet{Hall2008} are developed for the pointwise distribution function $F_Y(x)$, they imply natural interval-probability estimators.
For comparison, we therefore include
$$
\widehat{F}_{DR}(x,y) = \min\left\{1,\max\right\{0,\widehat{F}_{DR}(y)-\widehat{F}_{DR}(x)\left\}\right\}
, \quad \text{and} \quad
\widehat{F}_{HL}(x,y) = \min\left\{1,\max\right\{0,\widehat{F}_{HL}(y)-\widehat{F}_{HL}(x)\left\}\right\}
.
$$




\subsubsection{Numerical results}

{
\setlength{\tabcolsep}{9pt}
\begin{table}[h!t]
    \caption{
        Monte Carlo MSE$\times 100$ and variance$\times 100$ (in parentheses) for estimators of $F_Y(x,y)$.
        }\label{tab:FYxy_main}
    \adjustbox{max width=\textwidth, center = \textwidth}{
    \small
\begin{tabular}{lcccc}
\toprule
&
$\widehat{F}_{Y}^{CV}(x,y)$ &
$\widehat{F}_{Y}^{AR}(x,y)$ &
$\widehat{F}_{DR}(x,y)$ &
$\widehat{F}_{HL}(x,y)$ \\
\midrule
\multicolumn{1}{l}{\rule[-.10ex]{0pt}{3.25ex}\!Smooth $F_Y$, $Z\!\sim\! N\!\left(0,\frac{1}{5}\right)\!$} &  &  &  &  \\
\quad $n=200$ & 0.055 (0.048) & 0.048 (0.043) & \textbf{0.043 (0.037)} & 0.062 (0.061) \\
\quad $n=1000$ & 0.013 (0.011) & 0.013 (0.010) & \textbf{0.010 (0.007)} & 0.014 (0.013) \\
\quad $n=5000$ & 0.004 (0.004) & 0.005 (0.003) & \textbf{0.003 (0.002)} & 0.004 (0.003) \\
\multicolumn{1}{l}{\rule[-.10ex]{0pt}{3.25ex}\!Smooth $F_Y$, $Z\!\sim\! \Gamma\!\left(2,\frac{1}{5\sqrt{2}}\right)\!$} &  &  &  &  \\
\quad $n=200$ & 0.054 (0.048) & 0.046 (0.040) & \textbf{0.042 (0.035)} & 0.057 (0.056) \\
\quad $n=1000$ & 0.012 (0.011) & 0.012 (0.010) & \textbf{0.010 (0.007)} & 0.013 (0.012) \\
\quad $n=5000$ & 0.003 (0.003) & 0.003 (0.003) & \textbf{0.003 (0.002)} & 0.003 (0.003) \\
\multicolumn{1}{l}{\rule[-.10ex]{0pt}{3.25ex}\!Smooth $F_Y$, $Z\!\sim\! \mathrm{S}\Gamma\!\left(2,\frac{1}{5\sqrt{2}}\right)\!$} &  &  &  &  \\
\quad $n=200$ & 0.054 (0.048) & 0.045 (0.040) & \textbf{0.040 (0.034)} & 0.069 (0.068) \\
\quad $n=1000$ & 0.013 (0.011) & 0.013 (0.011) & \textbf{0.010 (0.007)} & 0.015 (0.014) \\
\quad $n=5000$ & 0.003 (0.003) & 0.003 (0.003) & \textbf{0.003 (0.002)} & 0.004 (0.003) \\
\midrule
\multicolumn{1}{l}{\rule[-.10ex]{0pt}{3.25ex}\!Non-smooth $F_Y$, $Z\!\sim\! N\!\left(0,\frac{1}{5}\right)\!$} &  &  &  &  \\
\quad $n=200$ & 0.319 (0.116) & 0.393 (0.074) & 0.390 (0.085) & \textbf{0.274 (0.095)} \\
\quad $n=1000$ & \textbf{0.180 (0.033)} & 0.270 (0.017) & 0.234 (0.018) & 0.188 (0.021) \\
\quad $n=5000$ & \textbf{0.138 (0.008)} & 0.210 (0.004) & 0.176 (0.004) & 0.144 (0.005) \\
\multicolumn{1}{l}{\rule[-.10ex]{0pt}{3.25ex}\!Non-smooth $F_Y$, $Z\!\sim\! \Gamma\!\left(2,\frac{1}{5\sqrt{2}}\right)\!$} &  &  &  &  \\
\quad $n=200$ & 0.283 (0.116) & 0.414 (0.071) & 0.402 (0.084) & \textbf{0.261 (0.091)} \\
\quad $n=1000$ & \textbf{0.106 (0.027)} & 0.210 (0.017) & 0.232 (0.017) & 0.152 (0.018) \\
\quad $n=5000$ & \textbf{0.052 (0.007)} & 0.115 (0.004) & 0.177 (0.004) & 0.103 (0.004) \\
\multicolumn{1}{l}{\rule[-.10ex]{0pt}{3.25ex}\!Non-smooth $F_Y$, $Z\!\sim\! \mathrm{S}\Gamma\!\left(2,\frac{1}{5\sqrt{2}}\right)\!$} &  &  &  &  \\
\quad $n=200$ & 0.289 (0.128) & 0.424 (0.077) & 0.414 (0.090) & \textbf{0.230 (0.111)} \\
\quad $n=1000$ & \textbf{0.101 (0.031)} & 0.210 (0.018) & 0.233 (0.018) & 0.116 (0.022) \\
\quad $n=5000$ & \textbf{0.045 (0.008)} & 0.115 (0.004) & 0.176 (0.003) & 0.077 (0.005) \\
\bottomrule
\end{tabular}

    }

    
    \smallskip
    \footnotesize
    \noindent
    \textit{Note:} Boldface identifies the estimator with
        the smallest unrounded empirical MSE in each row.
        Smooth $F_Y$ includes
        $Y\sim N(0,1)$ and $Y\sim\Gamma(3,1/\sqrt{3})$; non-smooth $F_Y$ includes
        $Y\sim K_1$, $K_2,$ and $K_3$.
        Within each $Y$ design,
        empirical MSE and variance are first averaged over the
        evaluation points; these design-level quantities are then averaged
        across the $Y$ designs in each group and multiplied by 100 for readability.
        Full DGP-by-DGP results are reported in Tables \ref{tab:FYxy_app_smth} and \ref{tab:FYxy_app_nonsmth} in Appendix~\ref{app:Full_sim_results}.

\end{table}
}


Table~\ref{tab:FYxy_main}
reports empirical MSE and variance for the four estimators of $F_Y(x,y)$.
The patterns are broadly similar to those
for pointwise estimation of $F_Y(x)$.

For smooth designs,
$\widehat{F}_{DR}(x,y)$ has the smallest grouped MSE throughout, and $\widehat{F}_{Y}^{AR}(x,y)$ performs competitively.
However, the DGP-level results from Table~\ref{tab:FYxy_app_smth} in Appendix~\ref{app:Full_sim_results} show that the advantage of $\widehat{F}_{DR}(x,y)$ is concentrated in the $Y\sim \Gamma\big(3,1/\sqrt{3}\big)$ designs, and $\widehat{F}_{Y}^{AR}(x,y)$ generally has a small advantage in the $Y\sim N(0,1)$ designs.


For non-smooth $F_Y$, $\widehat{F}_{HL}(x,y)$ has the smallest grouped MSE when $n=200$, whereas $\widehat{F}_{Y}^{CV}(x,y)$ performs best under all three error distributions when $n=1000$ and $n=5000$.
As with the $F_Y(x)$ results, the advantage of $\widehat{F}_{Y}^{CV}(x,y)$ is most prominent under ordinary-smooth errors.
For example, at $n=5000$, $\widehat{F}_{Y}^{CV}(x,y)$ reduces MSE  relative to $\widehat{F}_{HL}(x,y)$ by approximately $4\%$, $50\%$, and $42\%$ under normal, Gamma, and symmetrized-Gamma errors, respectively.
Table~\ref{tab:FYxy_app_nonsmth} in Appendix~\ref{app:Full_sim_results} shows that
$\widehat{F}_{Y}^{CV}(x,y)$ has the smallest MSE in all nine non-smooth combinations when
$n=5000$, and has the smallest MSE in four of the nine non-smooth designs when $n=1000$.

The corresponding bias--variance pattern is similar to the
$F_Y(x)$ estimation results.
In the non-smooth designs, the higher variance of $\widehat{F}_{Y}^{CV}(x,y)$ is offset by a larger reduction in squared
bias when sample size is large, whereas in smooth designs the variance of $\widehat{F}_{Y}^{CV}(x,y)$ is comparable to
that of the other estimators.




\subsection{Estimation of ${p}_x$}
This subsection studies the finite-sample performance of the proposed estimator of a jump
$p_x=F_Y(x)-\underset{\varepsilon \downarrow 0}{\lim}\,F_Y(x-\varepsilon)$
at the point $x\in J(F_Y)$.
We consider three types of jump distributions for $F_Y$.
Let $\Phi$ denote the standard normal distribution function.
The first two,
defined as in \citet{Mynbaev2022},
have a single jump of size $p$ at $x=0$
    $$
    F_Y^{(J_1)}(x;p)
    \;=\;
    \begin{cases}
            \Phi(x),
        &
            x <0,
        \\
            \frac{1}{2} + p,
        &
            x = 0,
        \\
            \int_{-\infty}^x
            \frac{1}{\sqrt{2\pi}}e^{-\frac{1}{2}(t-\mu_R)^2}
            \,dt
        &
            x >0,
    \end{cases}
    \qquad
    F_Y^{(J_2)}(x;p)
    \;=\;
    \begin{cases}
            \frac{1}{2}e^{x/8},
        &
            x <0,
        \\%[1.2em]
            \frac{1}{2} + p,
        &
            x = 0,
        \\%[1.2em]
            \int_{-\infty}^x
            \frac{1}{\sqrt{2\pi}}e^{-\frac{1}{2}(t-\mu_R)^2}
            \,dt
        &
            x >0,
    \end{cases}
    $$
where $\mu_R=-\Phi^{-1}(1/2+p)$.
The third distribution has three jumps, each of
size $p$, at $x\in\{-2,0,2\}$,
    $$
    F_Y^{(J_3)}(x;p) \;=\;
        p \cdot \chi_{\{x \geq -2\}}
        \;+\; p \cdot \chi_{\{x \geq 0\}}
        \;+\; p \cdot \chi_{\{x \geq 2\}}
        \;+\; (1-3p) \cdot
        \Phi\!\left(x\right)
        .
    $$
We write $Y\sim J_1$, $J_2$, or $J_3$ when $F_Y$ is given by $F_Y^{(J_1)}$, $F_Y^{(J_2)}$, or $F_Y^{(J_3)}$, respectively.
Figure~\ref{fig:jump-dgps} illustrates these three distribution functions for $p_x=0.20$.
\begin{figure}[h!]
\centering
\makebox[\textwidth][c]{
\begin{tikzpicture}

\begin{groupplot}[
    group style={group size=3 by 1, horizontal sep=1.35cm},
    width=0.35\textwidth,
    ymin=0, ymax=1,
    samples=200,
    axis lines=box,
    grid=major,
    grid style      =   {gray!30, line width=0.15pt},
    tick style      =   {gray!30, line width=0.20pt},
    axis line style =   {gray!75, line width=0.20pt},
    every axis plot/.append style={
        black,
        line width=0.55pt,
        mark=none,
        smooth
    },
    xlabel={$x$},
    ylabel={$F_Y(x)$},
    title style={font=\small},
    label style={font=\tiny},
    ticklabel style={font=\tiny},
    ytick={0,0.2,...,1},
    scaled ticks=false,
    enlargelimits=false,
    clip=true
]

\nextgroupplot[
    title={$Y\sim J_1$},
    xmin=-3, xmax=3,
    xtick={-3,-1.5,0,1.5,3},
]

\addplot[domain=-3:0] {normcdf(x)};
\addplot[domain=0:3] {normcdf(x-(-0.5244005))};

\addplot[sharp plot] coordinates {
    (0,0.5)
    (0,{0.5+0.20})
};

\nextgroupplot[
    title={$Y\sim J_2$},
    xmin=-30, xmax=5,
    xtick={-40,-30,-20,-10,0},
]

\addplot[domain=-40:0] {0.5*exp(x/8)};
\addplot[domain=0:5] {normcdf(x-(-0.5244005))};

\addplot[sharp plot] coordinates {
    (0,0.5)
    (0,{0.5+0.20})
};

\nextgroupplot[
    title={$Y\sim J_3$},
    xmin=-3, xmax=3,
    xtick={-4,-2,0,2,4},
]

\addplot[domain=-4:-2] {(1-3*0.20)*normcdf(x)};
\addplot[domain=-2:0] {0.20 + (1-3*0.20)*normcdf(x)};
\addplot[domain=0:2] {2*0.20 + (1-3*0.20)*normcdf(x)};
\addplot[domain=2:4] {3*0.20 + (1-3*0.20)*normcdf(x)};

\addplot[sharp plot] coordinates {
    (-2,{(1-3*0.20)*normcdf(-2)})
    (-2,{0.20+(1-3*0.20)*normcdf(-2)})
};

\addplot[sharp plot] coordinates {
    (0,{0.20+(1-3*0.20)*normcdf(0)})
    (0,{2*0.20+(1-3*0.20)*normcdf(0)})
};

\addplot[sharp plot] coordinates {
    (2,{2*0.20+(1-3*0.20)*normcdf(2)})
    (2,{3*0.20+(1-3*0.20)*normcdf(2)})
};

\end{groupplot}
\end{tikzpicture}
}
\caption{Jump distribution functions used in the simulation study, illustrated for $p_x=0.20$.}
\label{fig:jump-dgps}
\end{figure}


For each $p\in\{ 0.01, 0.10, 0.20\}$ we estimate the jump at $x=0$ for $F_Y^{(J_1)}$ and $F_Y^{(J_2)}$, and the jumps at $x\in\{-2,0,2\}$ for  $F_Y^{(J_3)}$.
We consider the same error distributions and sample sizes as before,
$
    n\in\{200,1000,5000\},
$
and generate 500 independent Monte Carlo samples for each design.



We report two feasible implementations of the proposed jump estimator. Both use an Epanechnikov kernel,
\[
    H(t)=\frac{3}{4}(1-t^2)\chi_{\{|t|\leq 1\}}.
\]

The first implementation, denoted $\hat{p}_{x}^{AR}$, chooses $h$ according to an asymptotic rule motivated by Corollary~\ref{coro2}. Applying the corollary with $k=m=2$ and $s=1$ gives
\[
h_n=
\begin{cases}
Cn^{-3/16},
    &
    \text{if }
    Z\sim\Gamma\left(2,\frac{1}{5\sqrt{2}}\right)
    \text{ or }
    Z\sim\mathrm{S}\Gamma\left(2,\frac{1}{5\sqrt{2}}\right)\!,
    \\[0.8em]
1.01\frac{\sqrt{2}}{5}
(\log n)^{-1/2},
    &
    \text{if } Z\sim N\left(0,\frac{1}{5}\right)\!
    ,
\end{cases}
\]
where
$C$
is calibrated so that
$h_{1000} \approx 0.015$ for ordinary-smooth errors,
and
the factor
$1.01$ provides a margin above the lower bound on the
super-smooth constant in Corollary~\ref{coro2}.



The second implementation, denoted $\hat{p}_{x}^{GL}$, chooses $h$ using a feasible Goldenshluger--Lepski style bandwidth selector, motivated by the procedure developed in \citet{goldenshluger_universal_2008}.
For each candidate $h$ in a finite grid $\mathcal H_{GL}$, we compute $\widehat p_{x,h}$ and the variance bound
$
    V_h
    =
    u_h^2/n
$
defined as in Theorem \ref{thm12}.
With this, we construct a proxy MSE criterion
\[
    Q_{\kappa_A,\kappa_V}(h)
    =
    A_{\kappa_A}(h)+\kappa_V V_h
,\quad \text{where}\quad
    A_{\kappa_A}(h)
    =
    \max_{h'\in\mathcal H_{GL}: h'\leq h}
    \left[
        \left(
            \widehat p_{x,h}-\widehat p_{x,h'}
        \right)^2
        -
        \kappa_A\left(V_h+V_{h'}\right)
    \right]_+,
\]
and $[a]_+=\max\{a,0\}$.
Then, the corresponding estimator is
\[
    \hat{p}_{x}^{GL}
=
    \widehat p_{x,h_{GL}(x)}
,\quad \text{where} \quad
    h_{GL}(x)
=
    \arg\min_{h\in\mathcal H_{GL}}
    Q_{\kappa_A,\kappa_V}(h)
    .
\]
The idea is that
$A_{\kappa_A}(h)$ acts as a feasible proxy for squared smoothing bias: pairwise differences between $\widehat p_{x,h}$ and finer-bandwidth estimators are interpreted as evidence of bias, not just random noise, only when they exceed the threshold $\kappa_A\left(V_h+V_{h'}\right)$.
In our simulations, the candidate grid $\mathcal H_{GL}$ is constructed around the asymptotic-rate
bandwidth $h_{n}$.
For the tuning parameters we use $\kappa_A=0.05,$ and $\kappa_V=0.25.$



\subsubsection{Numerical results}

{
\setlength{\tabcolsep}{5pt}
\begin{table}[h!t]
    \caption{
        Monte Carlo MSE$\times 100$ and variance$\times 100$ (in parentheses) for estimators of $p_x$.
        }\label{tab:jump_main}
    \adjustbox{max width=\textwidth, center = \textwidth}{
    \small
\begin{tabular}{lcccccc}
\toprule
& \multicolumn{2}{c}{$p_x = 0.01$} & \multicolumn{2}{c}{$p_x = 0.10$} & \multicolumn{2}{c}{$p_x = 0.20$} \\
\cmidrule(lr){2-3}
\cmidrule(lr){4-5}
\cmidrule(lr){6-7}
& $\hat{p}_{x}^{GL}$ & $\hat{p}_{x}^{AR}$ & $\hat{p}_{x}^{GL}$ & $\hat{p}_{x}^{AR}$ & $\hat{p}_{x}^{GL}$ & $\hat{p}_{x}^{AR}$ \\
\midrule
\multicolumn{1}{l}{\rule[-1.6ex]{0pt}{6.0ex}\!$Z\!\sim\! N\!\left(0,\frac{1}{5}\right)\!$} &  &  &  &  &  &  \\
\quad $n=200$ & \textbf{2.626 (0.127)} & 3.040 (0.088) & \textbf{2.358 (0.155)} & 2.613 (0.113) & \textbf{1.944 (0.189)} & 2.086 (0.136) \\
\quad $n=1000$ & \textbf{1.737 (0.038)} & 2.299 (0.018) & \textbf{1.525 (0.051)} & 1.956 (0.025) & \textbf{1.250 (0.059)} & 1.556 (0.030) \\
\quad $n=5000$ & \textbf{1.235 (0.019)} & 1.828 (0.005) & \textbf{1.089 (0.022)} & 1.571 (0.006) & \textbf{0.896 (0.023)} & 1.251 (0.007) \\
\multicolumn{1}{l}{\rule[-1.6ex]{0pt}{6.0ex}\!$Z\!\sim\! \Gamma\!\left(2,\frac{1}{5\sqrt{2}}\right)\!$} &  &  &  &  &  &  \\
\quad $n=200$ & 0.922 (0.199) & \textbf{0.823 (0.544)} & \textbf{0.871 (0.300)} & 1.349 (1.207) & \textbf{0.852 (0.440)} & 1.854 (1.775) \\
\quad $n=1000$ & 0.479 (0.068) & \textbf{0.411 (0.270)} & \textbf{0.425 (0.095)} & 0.614 (0.567) & \textbf{0.371 (0.131)} & 0.816 (0.785) \\
\quad $n=5000$ & 0.274 (0.032) & \textbf{0.211 (0.142)} & \textbf{0.239 (0.041)} & 0.316 (0.289) & \textbf{0.202 (0.047)} & 0.371 (0.355) \\
\multicolumn{1}{l}{\rule[-1.6ex]{0pt}{6.0ex}\!$Z\!\sim\! \mathrm{S}\Gamma\!\left(2,\frac{1}{5\sqrt{2}}\right)\!$} &  &  &  &  &  &  \\
\quad $n=200$ & \textbf{0.830 (0.225)} & 1.034 (0.697) & \textbf{0.967 (0.484)} & 1.950 (1.785) & \textbf{1.132 (0.793)} & 3.163 (3.029) \\
\quad $n=1000$ & \textbf{0.482 (0.101)} & 0.592 (0.390) & \textbf{0.472 (0.190)} & 1.003 (0.919) & \textbf{0.544 (0.353)} & 1.623 (1.575) \\
\quad $n=5000$ & \textbf{0.264 (0.037)} & 0.268 (0.183) & \textbf{0.261 (0.089)} & 0.611 (0.566) & \textbf{0.279 (0.174)} & 0.982 (0.968) \\
\bottomrule
\end{tabular}

    }

    
    \smallskip
    \footnotesize
    \noindent
    \textit{Note:} Boldface identifies the estimator with
        the smallest unrounded empirical MSE in each row.
    For each measurement-error distribution,
    jump size, and sample size, entries are averaged over the three jump
    designs $J_1,J_2,J_3$, with the $J_3$ entries first averaged over the jump
    locations $x=-2,0,2$.
    Full DGP-by-DGP results
    are reported in Table \ref{tab:jump_appendix} in Appendix~\ref{app:Full_sim_results}.
    

\end{table}
}



Table~\ref{tab:jump_main} reports the simulation results for estimating the jump
size $p_x$. MSE decreases with $n$ for both estimators in every grouped design.
Performance nevertheless depends strongly on the measurement-error
distribution. The normal-error designs produce the largest MSEs despite having
the smallest empirical variances, indicating that estimation under super-smooth
errors is primarily limited by regularization bias. MSE is substantially
smaller under the two ordinary-smooth error distributions.

Under normal measurement error, $\hat{p}_{x}^{GL}$ has lower MSE than
$\hat{p}_{x}^{AR}$ for every jump size and sample size, although $\hat{p}_{x}^{AR}$ has lower
variance. Thus, the improvement from GL selection does not arise from choosing
a uniformly less variable estimator; instead, its higher variance is more than
offset by a reduction in squared bias. The relative MSE advantage of
$\hat{p}_{x}^{GL}$ also generally increases with the sample size.

Under Gamma measurement error, the ranking depends on the jump size. At each
sample size, $\hat{p}_{x}^{AR}$ has lower MSE when $p_x=0.01$, whereas $\hat{p}_{x}^{GL}$
performs better when $p_x=0.10$ or $p_x=0.20$. Under symmetrized-Gamma
errors, $\hat{p}_{x}^{GL}$ has lower MSE throughout. In the ordinary-smooth designs,
the asymptotic-rate estimator often has substantially greater variance,
particularly for the larger jumps.

The DGP-specific results in Table~\ref{tab:jump_appendix} in Appendix~\ref{app:Full_sim_results} reveal some
heterogeneity behind these grouped comparisons, particularly at the smallest
sample size. The grouped rankings should therefore not be interpreted as
uniform across every latent distribution. Because the baseline GL estimator
is bias-dominated in several designs, Appendix
Table~\ref{tab:jump_GL_sensitivity} also reports sensitivity to reducing either
of the GL tuning coefficients $\kappa_A$ and $\kappa_V$. These less
conservative specifications tend to improve performance under normal errors
and for smaller jumps, but can perform worse for larger jumps under
symmetrized-Gamma errors. Overall, the results are similar across the three GL
specifications, and no alternative uniformly dominates.











\section{Conclusion}\label{sec:conclusion}
In this paper we proposed three nonparametric deconvolution estimators under the classical additive error-in-measurement model: an estimator of the distribution function at its continuity points, an estimator of interval probabilities, and a new estimator of the sizes of jump discontinuities.
These estimators extend the scope of existing work by accommodating much broader classes of latent distributions.
In particular, whereas existing jump-size estimators under additive measurement error impose discrete-continuous mixture structures and smoothness conditions, to the best of our knowledge the estimator proposed here is the first to accommodate an otherwise unrestricted latent distribution.


The proposed estimators are motivated by a generalized Fourier inversion theorem established in \cite{Mynbaev2022}, which reveals a direct connection between generalized Fourier inversion formulas and a broad class of kernel-based estimators. This connection provides a unified framework for constructing deconvolution estimators of several distributional functionals. Unlike existing deconvolution estimators based on Sobolev smoothness assumptions, the non-asymptotic bias bounds obtained here require neither the existence of a density nor global smoothness of the latent distribution. Instead, they depend primarily on local properties of the distribution function together with regularity assumptions on the measurement-error distribution and the regularization kernel.

For each estimator we derived explicit non-asymptotic bounds on the bias, variance, and mean squared error, and established asymptotic unbiasedness and consistency under both ordinary-smooth and super-smooth measurement-error distributions. Under additional local regularity assumptions we also obtained explicit convergence rates for the mean squared error. In particular, the pointwise and interval estimators remain valid for distributions that need not possess densities, while the jump estimator provides a direct method for recovering point masses obscured by measurement error.

The simulation study demonstrates that the proposed procedures perform well across a wide range of latent distributions and measurement-error models. The cross-validation implementation of the pointwise and interval estimators provides stable performance across smooth and nonsmooth designs, while the Goldenshluger–Lepski bandwidth selection procedure substantially improves the finite-sample performance of the jump estimator in many settings.

Several directions for future research remain. An important extension would be to establish minimax optimality of the proposed estimators over suitable classes of distributions. Another natural direction is the development of adaptive bandwidth selection procedures supported by asymptotic optimality theory, particularly for the jump estimator. It would also be of interest to extend the methodology to settings with unknown measurement-error distributions, repeated measurements, Berkson-type measurement errors, or multivariate latent variables. Finally, the development of confidence intervals and confidence bands for the proposed estimators remains an open problem.