EconBase
← Back to paper

Inference based on Kotlarski's Identity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

72,892 characters · 16 sections · 53 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference based on Kotlarski's Identity

abstractKotlarski's identity has been widely used in applied economic research. However, how to conduct inference based on this popular identification approach has been an open question for two decades. This paper addresses this open problem by constructing a novel confidence band for the density function of a latent variable in repeated measurement error model. The confidence band builds on our finding that we can rewrite Kotlarski's identity as a system of linear moment restrictions. The confidence band controls the asymptotic size uniformly over a class of data generating processes, and it is consistent against all fixed alternatives. Simulation studies support our theoretical results. \begin{description} • {\bf Keywords:} deconvolution, Kotlarski's identity, measurement error, uniform confidence band. \end{description}

Introduction

Empirical researchers are often interested in recovering features of unobserved variables in economic models. Kotlarski's identity \citep*{kotlarski1967} -- see also rao:1992 -- is one of the most popular tools used to identify probability density functions of unobserved latent variables. Since its first introduction to econometrics by \citet*{li/vuong:1998}, Kotlarski's identity has been widely used in economics when data admit repeated measurements. Examples of research topics that use Kotlarski's identity include, but are not limited to, empirical auctions LiPerrigneVuong2000,krasnokutskaya:2011, income dynamics bonhomme/robin:2010, and labor economics cunha/heckman/navarro:2005,cunha/heckman/schennach:2010,bonhomme/sauder:2011,kennan/walker:2011. In these applications, researchers are interested in identifying the probability density function $f_X$ of a latent variable $X$ among others. The variable $X$ of interest is not observed in data, but two measurements $(Y_1,Y_2)$ are available in data with classical errors, $U_1 = Y_1-X$ and $U_2 = Y_2-X$. Kotlarski's identity is a nonparametric identifying restriction for the probability density function $f_X$ of $X$ implied by this setup.

The existing econometric literature on Kotlarski's identity focuses on identification and consistent estimation of $f_X$ and related objects li/vuong:1998,Li2002,schennach:2004,Schennach2004et,Schennach2008,bonhomme/robin:2010,Evdokimov2010,ZindeWalsh2014,SongSchennachWhite2015,FirpoGalvaoSong2017 -- also see surveys on this literature by Chen/Hong/Nekipelov:2011 and Schennach:2016. On the other hand, satisfactory inference methods for $f_X$ are missing in this literature -- in fact, even the sharp rate of convergence is unknown for the estimators based on Kotlarski's identity under unrestrictive assumptions, and hence a limit distribution result is unavailable under such assumptions. Indeed some empirical papers implement nonparametric bootstrap without a theoretical guarantee. In light of the demands by the empirical researchers for an inference method, and given the current unavailability of theoretically supported methods of inference, we propose a method of inference based on Kotlarski's identity in this paper.

This paper proposes an inference method based on Kotlarski's identity by developing a confidence band for $f_X$. Our construction of confidence bands works as follows. First, we derive linear complex-valued moment restrictions based on Kotlarski's identity. Second, we let the Hermite polynomial sieve chen:2007 approximate unknown probability density functions. Third, for a given sieve dimension and for a given class of probability density functions, we compute a bias bound for the linear complex-valued moment restrictions, and slack the linear complex-valued moment restrictions by this bias bound. Fourth, applying \citet*{chernozhukov/chetverikov/kato:2017}, we compute the uniform norm of the self-normalized process of the slacked linear complex-valued moment restrictions as the test statistic for each point in a set of sieve coefficients. Fifth, inverting this test statistic in the spirit of anderson/rubin:1949 yields a confidence set of sieve approximations to possible probability density functions. Sixth, for a given sieve dimension and for a given class for probability density functions, we compute a bias bound for sieve approximations of probability density functions, and the desired confidence band is obtained by uniformly enlarging the set of sieve approximations by this bias bound.

The process of identifying $f_X$ in additive measurement error models is called deconvolution -- for solving convolution integral equations. There are a number of existing papers on nonparametric inference in deconvolution. bissantz/lutz/holzmann/munk:2007, bissantz/holtzmann:2008, es/gugushvili:2008, lounici2011, and schmidt-hieber2013 develop uniform confidence bands for $f_X$ under the assumption of known error distributions.\footnote{These paper are based on the literature on deconvolution under known error distribution carroll/hall:1988,stefanski/carroll:1990,fan:1991b,carrasco2011spectral. fan:1991a develops a point-wise asymptotic inference result in this framework.} In most economic applications, however, it is not plausible to assume that the error distributions are known. More recently, kato/sasaki:2018 and adusumilli/otsu/whang:2017 develop uniform confidence bands for $f_X$ and the distribution function, respectively, without assuming that the error distributions are known, but they both assume that at least one error distribution is symmetric.\footnote{These paper are based on the literature on deconvolution under unknown error distribution with auxiliary data or symmetric error distributions diggle/hall:1993,horowitz/markatou:1996,neumann:1997,efromovich:1997,delaigle/hall/meister:2008,johannes:2009,comte/lacour:2011,delaigle/hall:2015.} Kotlarski's identity is a powerful device for new identification results which require neither the known error distribution assumption nor the symmetric error distribution assumption. This useful feature attracts many economic applications including those listed above, but no econometrician has developed a method of inference in this framework for twenty years ever since its first introduction by \citet*{li/vuong:1998} until our present paper.

It it not surprising that such an inference method has been missing for long in the literature, given the technical difficulties of the problem. Deconvolution is an ill-posed inverse problem, and inference under this problem is known to be challenging -- see bissantz/lutz/holzmann/munk:2007,bissantz/holtzmann:2008,lounici2011,horowitz/lee:2012,hall/horowitz:2013,schmidt-hieber2013,adusumilli/otsu/whang:2017,kato/sasaki:2017b,kato/sasaki:2018,babii:2018,chen/christensen:2018 for existing papers developing confidence bands in ill-posed inverse problems for example. We take a robust inference approach \`a la anderson/rubin:1949, and directly work with the moment restrictions based on Kotlarski's identity. A positive side product of taking this approach is that we do not need to assume the non-vanishing characteristic functions (i.e., we do not need the completeness), which is commonly assumed for nonparametric identification or inversion.

It is also worth mentioning that we chose to use the Hermite polynomial sieve among other sieves in this paper. The Hermite polynomial sieve has been in fact already known in the literature to be useful to approximate “smooth density with unbounded support” chen:2007 -- also see her discussion of gallant/nychka:1987 therein. In addition to this known advantage, we also find this sieve particularly useful for the deconvolution problem. Note that the deconvolution problem involves applications of the Fourier transform operation and the inverse Fourier transform operation. To our convenience, the Hermite functions are eigen-functions of the Fourier transform operator. While we deal with simultaneous restrictions in terms of density and characteristic functions, we can use the Hermite polynomial sieve to approximate both the density and characteristic functions without having to apply the Fourier transform or the Fourier inverse because of the eigen-function property. This convenient property saves computational time and resources as costly numerical integration within each iteration of a numerical optimization routine would be necessary if any other sieve were used. Furthermore, we find that a couple of properties of the Hermite functions (namely the Schr\"odinger equation for a harmonic oscillator and a pair of recursive equations) can be exploited to obtain informative bias bounds of the Hermite polynomial sieve, which in turn contributes to informative inference we establish in this paper.

The rest of the paper is organized as follows. Section (ref) derives linear complex-valued moment restrictions based on Kotlarski's identity. Section (ref) presents how to compute the confidence band. Section (ref) presents asymptotic properties of the confidence band. Section (ref) discusses practical considerations. Section (ref) illustrates simulation studies. The paper concludes in Section (ref). All mathematical derivations and details are delegated to the appendix.

Linear Complex-Valued Moment Restrictions

Consider the repeated measurement model

equation[equation omitted — 93 chars of source]

where $Y_1$ and $Y_2$ are observed, but none of $X$, $U_1$, or $U_2$ is observed. We are interested in making inference on the probability density function $f_X$ of $X$. We equip this model with the following assumption.

assumption${}$ \begin{enumerate} • $X, \ U_1$, and $U_2$ are continuous random variables with finite first moments, and $U_1$ has mean zero. • $X, \ U_1$, and $U_2$ are mutually independent. \end{enumerate}

This assumption is standard in the literature on identification and estimation based on Kotlarski's identity li/vuong:1998. In fact, the existing literature imposes an additional assumption, namely the identification condition (non-vanishing characteristic function or the completeness) -- see Lemma (ref) ahead for a specific condition. We do not invoke such an identification assumption for the purpose of identification-roust inference -- see Remark (ref) ahead for further details.

We now fix basic notations. In what follows, $\mathbb{E}_P$ and $\mathbb{V}_P$ denote the expectation and variance operators, respectively, with respect to a joint distribution $P$ of $(Y_1,Y_2)$. Analogously, $\mathbb{E}_n$ and $\mathbb{V}_n$ denote the expectation and variance operators, respectively, with respect the empirical distribution of $n$ independent copies of $(Y_1,Y_2)$. We let $i=\sqrt{-1}$ denote the imaginary unit. For the set of absolutely integrable functions, $\mathcal{L}^1$, we define the Fourier transform $\mathcal{F}$ on $\mathcal{L}^{1}$ by $[\mathcal{F}f](t)=\int_{-\infty}^{\infty} e^{itx}f(x)dx$, and its inverse transform is $[\mathcal{F}^{-1}\phi](x)=\frac{1}{2\pi}\int_{-\infty}^{\infty} e^{-itx}\phi(t)dt$ -- see folland2007. In light of Assumption (ref) (i), we let $f_X$, $f_{U_1}$, and $f_{U_2}$ denote the density functions of $X$, $U_1$, and $U_2$, respectively. Further, we denote the characteristic functions of them by $\phi_X=\mathcal{F}f_X$, $\phi_{U_1}=\mathcal{F}f_{U_1}$, and $\phi_{U_2}=\mathcal{F}f_{U_2}$. We first review the existing result of the identification.

lemma[Kotlarski's Identity] For every joint distribution $P$ of $(Y_1,Y_2)$ satisfying Assumption (ref) for ((ref)) and $\mathbb{E}_{P}\left[e^{itY_2}\right] \neq 0$ for all $t \in \mathbb{R}$, \begin{align} \phi_X(t) = \exp\left( \int_0^t \frac{i \mathbb{E}_P\left[Y_1 e^{i\tau Y_2}\right]}{\mathbb{E}_P\left[e^{i\tau Y_2}\right]} d\tau \right). \end{align}

This lemma presents Kotlarski's identity due to kotlarski1967 -- see also rao:1992. Since it is stated as a lemma, Kotlarski's identity is also known as Kotlarski's lemma or the lemma of Kotlarski in the econometrics literature. li/vuong:1998 first introduced it into econometrics and statistics, followed by a series of extensions Li2002,schennach:2004,bonhomme/robin:2010,Evdokimov2010. Some of these extensions relax the assumptions for identification and estimation in various ways. We do not need to rely on the prototypical assumptions for our purpose of inference, even though they are stated in Lemma (ref) for convenience of a concise review. Lemma (ref) shows that the characteristic function $\phi_X$ of $X$ is explicitly identified by the joint distribution of $(Y_1,Y_2)$. Under the additional assumption of absolutely integrable characteristic function $\phi_X$, the formula $f_X=\mathcal{F}^{-1}\phi_X$ in turn yields the identification of the probability density function $f_X$ of $X$.

Uniform convergence rates for the estimator of $f_X$ based on Kotlarski's identity are discovered in the existing literature li/vuong:1998,Li2002,schennach:2004,bonhomme/robin:2010,Evdokimov2010, but the sharp rates under unrestrictive assumptions are still unknown. In particular, limit distribution results under such assumptions are still unknown in the existing literature. This paper does not aim to derive a non-degenerate limit distribution for any estimator, but it aims to conduct an inference on $f_X$. With this said, our proposed inference does not rely on an explicit identifying formula. We argue that rewriting Kotlarski's lemma in terms of moment restrictions suffices and serves even more conveniently for the sake of conducting inference.

theorem[Linear Complex-Valued Moment Restrictions] For every joint distribution $P$ of $(Y_1,Y_2)$ satisfying Assumption (ref) for ((ref)), \begin{equation} \mathbb{E}_P\left[\left(iY_1\phi_X(t)-\phi_X^{(1)}(t)\right)\exp(itY_2)\right]=0 \end{equation} holds for every real $t$, where $\mathbb{E}_P$ is the expectation under $P$.

A proof is provided in Appendix (ref).

remarkTaking a few more steps beyond the claim in Theorem (ref) will lead us to the identification result of Lemma (ref) under the additional assumption of the invertibility or non-vanishing characteristic functions (also known as the completeness) -- see dhaultfoeuille:2011.\footnote{evdokimov2012some provides a relaxed assumption for the identification.} For the purpose of inference, however, it is not essential to solve the inverse problem, and thus we stop short of obtaining the explicit formula ((ref)), and only use the moment condition ((ref)). This idea is analogous to that of santos2011instrumental,santos2012inference, where robust inference for functional parameters is conducted without assuming the completeness.

Construction of the Confidence Band

Our objective is to construct a confidence band for the probability density function $f_X$ of $X$ on an interval $I \subset \mathbb{R}$. The construction procedure is based on the linear complex-valued moment restriction ((ref)). Throughout, we focus on the set of probability density functions given by

equation*[equation* omitted — 179 chars of source]

For this set of candidate probability density functions, we use $\mathcal{L}^1$ for applying the Fourier transform and the inverse, whereas $\mathcal{L}^2$ is used to approximate $\mathcal{L}$ by an orthonormal basis $\Psi=\{\psi_j \ : \ j=0,1,\ldots\}$ of $\mathcal{L}^2$ -- see Section (ref) for the example of the Hermite basis.

We use a $(q+1)$-dimensional sieve basis $\{\psi_0,\ldots,\psi_q\} \subset \Psi = \{\psi_j : j =0,1,\ldots\}$, with $\psi_j \in \mathcal{L}^1 \cap \mathcal{L}^2$ and $\mathcal{F} \psi_j \in \mathcal{L}^1$ for each $j \in \mathbb{N}$, to approximate the probability density function $f_X$. Let $\Theta^{q+1}\subset \mathbb{R}^{q+1}$ be a compact set, and write $\boldsymbol{\psi}=(\psi_0,\dots,\psi_q)^T$. With a uniform tolerance level $\eta > 0$, each $f \in \mathcal{L}$ is approximated by $ x \mapsto \boldsymbol{\psi}(x)^T \boldsymbol{\theta} $ for some $\boldsymbol{\theta}=(\theta_0,\ldots,\theta_q)^T \in \Theta^{q+1}$, i.e., $ \sup_{x \in I} \left| f(x) - \boldsymbol{\psi}(x)^T\boldsymbol{\theta} \right| \leq \eta. $ The set of values of the sieve coefficients $\boldsymbol{\theta} \in \Theta^{q+1}$ approximating a probability density function $f \in \mathcal{L}$ in this manner is denoted by

equation[equation omitted — 175 chars of source]

We next incorporate the linear complex-valued moment restrictions ((ref)) in this sieve framework. For every function $\psi \in \mathcal{L}$ and for every frequency $t \in \mathbb{R}$, define

align[align omitted — 381 chars of source]

where $\phi = \mathcal{F} \psi$. Note that $\mathrm{Re}(\cdot)$ (resp., $\mathrm{Im}(\cdot)$) denotes the real (resp., imaginary) part of a complex number. Further, stack these functions across $\psi \in \{\psi_0,\dots,\psi_q\}$ to define the random vector

align*[align* omitted — 183 chars of source]

With these notations, we now represent the linear complex-valued moment restrictions ((ref)) for the sieve approximation by

align[align omitted — 200 chars of source]

for all $t \in [-T,T]$ for $T \in (0,\infty)$, where $\delta(t)>0$ is the tolerance level of sieve approximation error for each $t \in [-T,T]$.

Our construction of the confidence band is based on a test statistic that quantifies the extent of deviation from the moment inequalities ((ref)). To construct a feasible test statistic, we use a grid $\{t_1,\dots,t_L\} \subset [-T,T]$ of $L$ frequencies. Define the test statistic by

equation*[equation* omitted — 390 chars of source]

for each $\boldsymbol{\theta}\in\Theta^{q+1}$. Let $\alpha\in(0,1/2)$ be fixed. We define the critical value $c(\alpha,\boldsymbol{\theta})$ of this statistic $T(\boldsymbol{\theta})$ by the conditional $(1-\alpha)$-th quantile of the multiplier bootstrap statistic $$ \sqrt{n}\max_{1\leq l\leq L}\max\left\{\frac{\left|\mathbb{E}_n[\boldsymbol{\epsilon}(\mathbf{R}_{t_l}-\mathbb{E}_n[\mathbf{R}_{t_l}])]^T\boldsymbol{\theta}\right|}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_n(\mathbf{R}_{t_l})\boldsymbol{\theta}}},\frac{\left|\mathbb{E}_n[\boldsymbol{\epsilon}(\mathbf{I}_{t_l}-\mathbb{E}_n[\mathbf{I}_{t_l}])]^T\boldsymbol{\theta}\right|}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_n(\mathbf{I}_{t_l})\boldsymbol{\theta}}} \right\} $$ given the data, where $\boldsymbol{\epsilon}_1,\ldots,\boldsymbol{\epsilon}_n$ are independent standard normal random variables independent of the data. As a more conservative yet simpler alternative following chernozhukov/chetverikov/kato:2017, we may define the critical value as

equation*[equation* omitted — 118 chars of source]

where $\Phi$ is the cumulative distribution function of the standard normal distribution. Our confidence band for the density function of $X$ is given by the $I$-restriction of

equation[equation omitted — 219 chars of source]

where $\mathbf{B}_{q+1,\eta}(f)$ is defined in ((ref)).

A practical procedure to obtain this confidence band is outlined as Algorithm (ref) below. While this algorithm prescribes a general procedure, we also elaborate on details of practical considerations in Section (ref), where we introduce the Hermite orthonormal sieve (Section (ref)), propose concrete choice rules for the tuning parameters (Section (ref)), and present a more concrete algorithm for these sieve and tuning parameters (Section (ref)).

algorithm[algorithm omitted — 1,086 chars of source]
remarkThe baseline idea of our confidence band construction is to discretize the continuum of moment conditions ((ref)) into a set of finite but many moment inequalities on the sieve coefficients $\boldsymbol{\theta}$, calibrate critical values by the multiplier bootstrap for the max statistic as in chernozhukov/chetverikov/kato:2017, and project a confidence set for $\boldsymbol{\theta}$ into a confidence band for $f_{X}$. A natural question would be whether we could use directly the continuum of moment conditions ((ref)) without discretization, similarly to, e.g., AndrewsShi2013,ANDREWS2017275. One difficulty, however, is that the moment functions corresponding to the moment condition ((ref)) at given $f_{X}$, i.e., $\{ (y_{1},y_{2}) \mapsto R_{f_{X},t}(y_{1},y_{2}), I_{f_{X},t}(y_{1},y_{2}) : t \in \mathbb{R} \}$, is not likely to be Donsker, in view of the fact that, e.g., the function class $\{ y_{2} \mapsto \cos (ty_{2}) : t \in \mathbb{R} \}$ is non-Donsker as soon as $Y_{2}$ has a density (which is the case under our assumption), so that the “manageability” condition in AndrewsShi2013,ANDREWS2017275 would not be satisfied in our case. Indeed, the preceding function class is not Glivenko-Cantelli from the Riemann-Lebesgue lemma and discreteness of the empirical distribution; see FeuervergerMureika1977 for details. Another potential approach would be to apply the method of continuum of moment conditions developed in CarrascoFlorens2000. Their analysis relies on point-identification of the parameter of interest and, more importantly, focused on a finite dimensional parameter of interest (so the convergence rate of their estimator is the parametric rate), so that their approach is not directly applicable to our problem.

Properties of the Confidence Band

In this section, we present theoretical properties of the confidence band ((ref)). Let $\mathcal{P}$ denote a given space to which the joint distribution of $(Y_1,Y_2)$ belongs. For every $P\in\mathcal{P}$, define the identified set

equation[equation omitted — 235 chars of source]

as the set of density functions for which the linear complex-valued moment restriction ((ref)) is satisfied. Furthermore, we define the sieve-approximation counterpart of $\mathcal{L}_0(P)$ by

equation[equation omitted — 326 chars of source]

where $\mathbf{B}_{q+1,\eta}(f)$ is defined in ((ref)). Here, the infimum over the empty set is understood to be the infinity.

The current section is structured as follows. First, we establish the size control for the confidence band ((ref)) to contain the approximation set $\mathcal{L}_0^\ast(P)$ in Section (ref). Second, we establish the containment of the identified set by the approximation set (i.e., $\mathcal{L}_0(P) \subset \mathcal{L}_0^\ast(P)$) in Section (ref). These two pieces of the results together show the validity of the confidence band ((ref)) to contain the identified set $\mathcal{L}_0(P)$. Lastly, Section (ref) presents power properties with local alternatives. Throughout, we assume to observe $n$ i.i.d. copies of $(Y_1,Y_2)$ drawn from $P \in \mathcal{P}$.

Size Control

We make the following assumption for a uniform size control.

assumption(i) $\mathbb{E}_P\left[Y_1^2\right]<\infty$. (ii) There are constants $0 < c_1<1/2$ and $C_1 > 0$ such that $$ \left(M_{L,q,3}^3(\boldsymbol{\theta},P) \vee M_{L,q,4}^2(\boldsymbol{\theta},P) \vee B_{L,q}(\boldsymbol{\theta},P) \right)^2 \log^{7/2}(4Ln)\leq C_1n^{1/2-c_1} $$ for all $P \in \mathcal{P}$ and $\theta \in \Theta^{q+1}$, where\\{ $ M_{L,q,k}(\boldsymbol{\theta},P) = \max_{1 \leq l \leq L} \max \left\{\mathbb{E}_P\left[\left|\frac{(\mathbf{R}_{t_l}-\mathbb{E}_P[\mathbf{R}_{t_l}])^T\boldsymbol{\theta}}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_P(\mathbf{R}_{t_l})\boldsymbol{\theta}}}\right|^k\right]^{1/k}, \mathbb{E}_P\left[\left|\frac{(\mathbf{I}_{t_l}-\mathbb{E}_P[\mathbf{I}_{t_l}])^T\boldsymbol{\theta}}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_P(\mathbf{I}_{t_l})\boldsymbol{\theta}}}\right|^k\right]^{1/k} \right\} $}\\ and { $ B_{L,q}(\boldsymbol{\theta},P) = \mathbb{E}_P\left[\max_{1 \leq l \leq L} \max \left\{\left|\frac{(\mathbf{R}_{t_l}-\mathbb{E}_P[\mathbf{R}_{t_l}])^T\boldsymbol{\theta}}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_P(\mathbf{R}_{t_l})\boldsymbol{\theta}}}\right|^4 , \left|\frac{(\mathbf{I}_{t_l}-\mathbb{E}_P[\mathbf{I}_{t_l}])^T\boldsymbol{\theta}}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_P(\mathbf{I}_{t_l})\boldsymbol{\theta}}}\right|^4 \right\} \right]^{1/4} $}.
remarkWhen $\mathbb{E}_P\left[Y_1^4\right]<\infty$ and $\{\psi_0,\ldots,\psi_q\}$ are the Hermite functions, Assumption (ref) (ii) holds if \begin{equation} (q+1)/ \min_{t\in[-T,T]}\min\left\{ \mathrm{eig}_{\min}(\mathbb{V}_P(\mathbf{R}_{t})), \mathrm{eig}_{\min}(\mathbb{V}_P(\mathbf{I}_{t})) \right\}=O(n^{1/6-c_1/3}\log^{-7/6}(4Ln)). \end{equation} A derivation of this condition is in Appendix (ref). Eq. ((ref)) restricts how fast $q$ and $T$ can increase to infinity.\footnote{Loosely speaking, the term $(q+1)/ \min_{t\in[-T,T]}\min\left\{ \mathrm{eig}_{\min}(\mathbb{V}_P(\mathbf{R}_{t})), \mathrm{eig}_{\min}(\mathbb{V}_P(\mathbf{I}_{t})) \right\}$ is increasing in $q$ and $T$. When $q$ increases, the numerator increases and the minimum eigenvalues of the $(q+1)$ square matrices, $\mathbb{V}_P(\mathbf{R}_{t})$ and $\mathbb{V}_P(\mathbf{I}_{t})$, can be smaller. Moreover, when $T$ increases, the denominator decreases.} On the other hand, Eq. ((ref)) does not restrict the choice of $L$ in the sense that, as long as $L$ grows at a polynomial rate of $n$, the term $\log^{-7/6}(4Ln)$ is negligible on the right hand side.
theorem[Size Control] Suppose that Assumption (ref) holds. Then, there exist positive constants $c$ and $C$ depending only on $c_{1}$ and $C_{1}$ such that $$ \inf_{P\in\mathcal{P}}\inf_{f\in\mathcal{L}_0^\ast(P)}\mathbb{P}_P(f\in \mathcal{C}_{n}(\alpha))\geq 1-\alpha-Cn^{-c}. $$

A proof is provided in Appenedix (ref). This theorem guarantees the size control for the event of the inclusion of the density function in the approximate identified set $\mathcal{L}_0^\ast(P)$ as opposed to the identified set $\mathcal{L}_0(P)$. The next section presents conditions under which the approximate identified set $\mathcal{L}_0^\ast(P)$ contains the identified set $\mathcal{L}_0(P)$ for every possible joint distribution $P \in \mathcal{P}$.

Bounds of Approximation Errors

Throughout, we equip $\mathcal{L}^2$ with the inner product $\langle{\cdot,\cdot}\rangle$ defined by $$ \langle{f_1,f_2}\rangle = \int_\mathbb{R} f_1(x) f_2(x) dx. $$

assumption${}$\\ (i) $\Psi = \{\psi_j \ : \ j=0,1,\ldots\}$ is an orthonormal basis of $\mathcal{L}^2$, i.e., orthonormal and complete in $\mathcal{L}^2$. \\ (ii) $\left( \langle{f,\psi_0}\rangle, \dots, \langle{f,\psi_q}\rangle\right) \in \Theta^{q+1}$ for all $f \in \mathcal{L}_0(P)$ for all $P \in \mathcal{P}$.

In Section (ref), we propose a concrete orthonormal basis $\Psi$ and the set $\Theta^{q+1}$ of coefficients to satisfy Assumption (ref). The following theorem provides a guide for choices of $\eta$ and $\delta(t_1),\dots,\delta(t_L)$ such that $\mathcal{L}_0^\ast(P)$ contains $\mathcal{L}_0(P)$ for every possible $P \in \mathcal{P}$.

theorem[Approximation] Suppose that Assumption (ref) is satisfied. If \begin{align} &\sup_{f\in\mathcal{L}}\sup_{t\in I}\left|\sum_{j=q+1}^{\infty}\langle{ f,\psi_j }\rangle \cdot \psi_j(t)\right| \leq \eta \qquadand \\ &\sup_{f\in\mathcal{L}}\left|\sum_{j=q+1}^{\infty} \langle{ f,\psi_j }\rangle \cdot \left( i\phi_j(t_l)\cdot\mathbb{E}_P\left[Y_1\exp(it_lY_2)\right] -\phi_j^{(1)}(t_l)\cdot\mathbb{E}_P\left[\exp(it_lY_2)\right] \right) \right| \leq \delta(t_l) \end{align} for every $l \in \{1,\ldots,L\}$, then $\mathcal{L}_0(P)\subset\mathcal{L}_0^\ast(P)$ holds for all $P \in \mathcal{P}$.

A proof is provided in Appendix (ref). In Section (ref), we provide concrete evaluations of the left-hand side of ((ref)) and ((ref)) under a concrete orthonormal basis $\Psi$, namely the Hermite orthonormal basis. Putting Theorems (ref) and (ref) together, we obtain the following result on the validity of the confidence band.

corollary[Validity of the Confidence Band] Suppose that the conditions of Theorems (ref) and (ref) are satisfied. Then, there exist positive constants $c$ and $C$ depending only on $c_{1}$ and $C_{1}$ such that $$ \inf_{P\in\mathcal{P}}\inf_{f\in\mathcal{L}_0(P)}\mathbb{P}_P(f\in \mathcal{C}_{n}(\alpha))\geq 1-\alpha-Cn^{-c}. $$

Power

We introduce the following short-hand notation for the random variable defined as the maximum deviation of the sample variance from the population variance: $$ B_{V}= \sup_{\boldsymbol{\theta}\in \mathbf{B}_{q+1,\eta}(f)}\max_{l=1,\ldots,L}\max\left\{\boldsymbol{\theta}^T(\mathbb{V}_n(\mathbf{R}_{t_l})-\mathbb{V}_P(\mathbf{R}_{t_l}))\boldsymbol{\theta},\boldsymbol{\theta}^T(\mathbb{V}_n(\mathbf{I}_{t_l})-\mathbb{V}_P(\mathbf{I}_{t_l}))\boldsymbol{\theta}\right\}. $$ The following theorem shows a power property of our proposed inference method.

theorem[Power] Suppose that Assumption (ref) (i) holds. For every $P\in\mathcal{P}$, every $f\in\mathcal{L}$, every $\nu>0$, and every $b\in(0,\infty)$, if there is $t_\ast \in \{t_1,\dots,t_L\}$ such that at least one of the following statements holds: \begin{align} \sqrt{n}\inf_{\boldsymbol{\theta}\in \mathbf{B}_{q+1,\eta}(f)}\frac{\mathbb{E}_P[\mathbf{R}_{t_\ast}]^T\boldsymbol{\theta}-\delta(t_\ast)}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_P(\mathbf{R}_{t_\ast})\boldsymbol{\theta}+\nu}} \geq(1+b)\cdot\mathbb{E}_P\left[\sup_{\boldsymbol{\theta}\in\mathbf{B}_{q+1,\eta}(f)}\frac{|\mathbb{G}_n[\mathbf{R}_{t_\ast}]^T\boldsymbol{\theta}|}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_n(\mathbf{R}_{t_\ast})\boldsymbol{\theta}}}\right]\notag\\ +\sqrt{2\log(4L)}+\sqrt{2\log(1/\alpha)} \end{align} \begin{align} \sqrt{n}\inf_{\boldsymbol{\theta}\in \mathbf{B}_{q+1,\eta}(f)}\frac{-\mathbb{E}_P[\mathbf{R}_{t_\ast}]^T\boldsymbol{\theta}-\delta(t_\ast)}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_P(\mathbf{R}_{t_\ast})\boldsymbol{\theta}+\nu}} \geq(1+b)\cdot\mathbb{E}_P\left[\sup_{\boldsymbol{\theta}\in\mathbf{B}_{q+1,\eta}(f)}\frac{|\mathbb{G}_n[\mathbf{R}_{t_\ast}]^T\boldsymbol{\theta}|}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_n(\mathbf{R}_{t_\ast})\boldsymbol{\theta}}}\right]\notag\\ +\sqrt{2\log(4L)}+\sqrt{2\log(1/\alpha)} \end{align} \begin{align} \sqrt{n}\inf_{\boldsymbol{\theta}\in \mathbf{B}_{q+1,\eta}(f)}\frac{\mathbb{E}_P[\mathbf{I}_{t_\ast}]^T\boldsymbol{\theta}-\delta(t_\ast)}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_P(\mathbf{I}_{t_\ast})\boldsymbol{\theta}+\nu}} \geq(1+b)\cdot\mathbb{E}_P\left[\sup_{\boldsymbol{\theta}\in\mathbf{B}_{q+1,\eta}(f)}\frac{|\mathbb{G}_n[\mathbf{I}_{t_\ast}]^T\boldsymbol{\theta}|}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_n(\mathbf{I}_{t_\ast})\boldsymbol{\theta}}}\right]\notag\\ +\sqrt{2\log(4L)}+\sqrt{2\log(1/\alpha)} \end{align} \begin{align} \sqrt{n}\inf_{\boldsymbol{\theta}\in \mathbf{B}_{q+1,\eta}(f)}\frac{-\mathbb{E}_P[\mathbf{I}_{t_\ast}]^T\boldsymbol{\theta}-\delta(t_\ast)}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_P(\mathbf{I}_{t_\ast})\boldsymbol{\theta}+\nu}} \geq(1+b)\cdot\mathbb{E}_P\left[\sup_{\boldsymbol{\theta}\in\mathbf{B}_{q+1,\eta}(f)}\frac{|\mathbb{G}_n[\mathbf{I}_{t_\ast}]^T\boldsymbol{\theta}|}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_n(\mathbf{I}_{t_\ast})\boldsymbol{\theta}}}\right]\notag\\ +\sqrt{2\log(4L)}+\sqrt{2\log(1/\alpha)}, \end{align} then $$ \mathbb{P}_P(f\notin \mathcal{C}_{n}(\alpha)) \geq \mathbb{P}_P\left(B_{V}\leq \nu\right)-\frac{1}{1+b}. $$

A proof is provided in Appendix (ref). According to this theorem, for any density function $f \in \mathcal{L}$ such that at least one of the moment inequalities violated at some frequency point $t_\ast$ in the grid $\{t_1,\dots,t_L\}$, then the probability that this density function does not belong to the confidence band is bounded below by $\mathbb{P}_P\left(B_{V}\leq \nu\right)-\frac{1}{1+b}$. Choosing sequences of $\nu$ and $b$ so that $\mathbb{P}_P\left(B_{V}\leq \nu\right)-\frac{1}{1+b} \rightarrow 1$ as $n \rightarrow \infty$, therefore, this theorem implies the consistency against all fixed alternatives.

Practical Considerations

The current section presents a guide to practice. Construction of the confidence band and its theoretical properties are presented in Sections (ref) and (ref) with abstract objects. These objects in particular include an orthonormal basis $\Psi = \{\psi_j \ : \ j=0,1,\ldots\}$, the sieve dimension $q$, the set $\Theta^{q+1}$ of sieve coefficients, and the tolerance levels, $\eta$, $\delta(t_1),\dots, \delta(t_L)$, of approximation errors. Section (ref) presents concrete choices of $\Psi = \{\psi_j \ : \ j=0,1,\ldots\}$ and $\Theta^{q+1}$. Section (ref) presents a concrete data-driven procedure for selecting the tolerance levels $\eta$, $\delta(t_1),\dots, \delta(t_L)$. Section (ref) presents a concrete implementation procedure for constructing the confidence band with these choices of the objects.

The Hermite Orthonormal Basis

There is a large extent of freedom of choice for an orthonormal basis $\Psi$ -- see chen:2007 for a list of options. We recommend the Hermite orthonormal basis in particular for its convenient properties and its nice compatibility with the deconvolution framework -- a Hermite function is an eigenfunction of the Fourier transform and the Fourier inverse.\footnote{The Hermite orthonormal sieve is not location invariant, and hence we recommend to location- and scale normalize the observed data using the empirical moments.} The Hermite functions take the form

equation[equation omitted — 121 chars of source]

$j=0,1,\ldots$, where $H_j$ is the Hermite polynomial defined by $$ H_j(x)=(-1)^{j} \cdot \exp(x^2) \cdot \frac{d^{j}}{dx^{j}}\exp(-x^2). $$ The Hermite functions are the eigenfunctions of the Fourier transform operator, and specifically, $\phi_j=\mathcal{F}\psi_j=i^j\sqrt{2\pi}\psi_j$ holds. For any $q \in \mathbb{N}$ to be chosen below, have the set of sieve coefficients satisfy

equation[equation omitted — 119 chars of source]

These concrete choices are made for the sake of satisfying Assumption (ref) so we can use Theorem (ref).

proposition[Sufficient Condition for Assumption (ref)] If $\Psi=\{\psi_j:j=0,1,\ldots\}$ is the sequence of the Hermite functions given in ((ref)), then it satisfies Assumption (ref) with the coefficient set given in ((ref)).

A proof is provided in Appendix (ref). We also present the following proposition which provides approximation bounds for the condition of Theorem (ref) to guarantee the containment of the identified set $\mathcal{L}_0(P)$ by the approximate identified set $\mathcal{L}_0^\ast(P)$ for every possible joint distribution $P \in \mathcal{P}$.

proposition[Approximation Bounds] Suppose that $\Psi=\{\psi_j:j=0,1,\ldots\}$ is the sequence of the Hermite functions given in ((ref)) and $\Theta^{q+1}$ satisfies ((ref)). If \begin{equation} \int \left( \frac{d^2}{d x^2} (f^{(2)}(x)+x^2f(x)) + x^2 \cdot (f^{(2)}(x)+x^2f(x)) \right)^2 dx \leq M, \end{equation} then, for each $q=1,2,\ldots$, \begin{align} &\sup_{f\in\mathcal{L}}\sup_{x\in I}\left|\sum_{j=q+1}^{\infty}\langle{ f,\psi_j }\rangle \cdot \psi_j(x)\right| \leq \frac{1.086435\pi^{-1/4}}{\sqrt{2q+3}}\sqrt{M\sum_{j=q+1}^{\infty}(2j+1)^{-3}}, \\ &\sup_{f\in\mathcal{L}}\sup_{t\in \mathbb{R}}\left|\sum_{j=q+1}^{\infty}\langle{ f,\psi_j }\rangle \cdot \phi_j(t)\right| \leq \frac{1.086435\pi^{-1/4}\sqrt{2\pi}}{\sqrt{2q+3}}\sqrt{M\sum_{j=q+1}^{\infty}(2j+1)^{-3}}, \qquadand \\ &\sup_{f\in\mathcal{L}}\left|\sum_{j=q+1}^{\infty} \langle{ f,\psi_j }\rangle \cdot \left(i\phi_j(t)\cdot\mathbb{E}_P\left[Y_1\exp(itY_2)\right]-\phi_j^{(1)}(t)\cdot\mathbb{E}_P\left[\exp(itY_2)\right]\right)\right| \notag\\ &\leq \frac{1.086435\pi^{-1/4}}{\sqrt{2\pi}}\cdot\sqrt{M\sum_{j=q+1}^{\infty}(2j+1)^{-3}}\left(\frac{\mathbb{E}_P\left[|Y_1|\right]}{\sqrt{2q+3}}+1\right) \qquadfor all t. \end{align}

A proof is provided in Appendix (ref). An admissible function class in terms of smoothness restriction can be specified by ((ref)). Note that this is analogous to the standard practice in the literature to work with Sovolev classes of functions. With the function class specified in this way, equations ((ref)) and ((ref)) prescribe possible choices of the tolerance levels $\eta,\delta(t_1),\dots,\delta(t_L)$ which admit $\mathcal{L}_0(P) \subset \mathcal{L}_0^\ast(P)$ for all $P \in \mathcal{P}$ per Theorem (ref). See Section (ref) for concrete choice procedures. Equation ((ref)) in addition suggests the worst approximation error for the characteristic function, which will be useful when we impose natural restrictions on the characteristic functions -- see Section (ref). Finally, we present the following proposition showing an alternative representation of the function class specification ((ref)).

proposition[Equivalent Smoothness Condition] Suppose that $x \mapsto \frac{d^2}{d x^2} (f''(x)+x^2f(x)) + x^2 \cdot (f''(x)+x^2f(x))$ is $\mathcal{L}^2$ and $\lim_{|x| \rightarrow \infty} x^2 f(x) = \lim_{|x| \rightarrow \infty} x^2 f^{(1)}(x) = \lim_{|x| \rightarrow \infty} f^{(2)}(x) = \lim_{|x| \rightarrow \infty} f^{(3)}(x) = 0$. Then ((ref)) is equivalent to \begin{equation} \int \left|\int\left(x^4 - (2t^2 + it)x^2 -(2 - 4it)x + (t^4 + 2)\right)e^{itx}f(x)dx\right|^2 dt \leq M. \end{equation}

A proof is provided in Appendix (ref). The components in the left-hand side of ((ref)) are estimable by li/vuong:1998 and its extensions, and thus Proposition (ref) provides a guideline for setting the smoothness bound parameter $M$ -- see Section (ref).

Choice of Tuning Parameters

In this section, we provide example procedures of choosing the smoothness bound $M$ and the tolerance levels $\eta, \delta(t_1),\dots,\delta(t_L) \in (0,\infty)$ in finite sample. Although we present a data driven choice of the smoothness bound $M$ below, we remark that the smoothness bound $M \in (0,\infty)$ as well as the sieve dimension $q \in \mathbb{N}$ could be imposed by a researcher in the spirit of honest inference armstrong/kolesar:2018.

{\bf Smoothness Bound $M$:} We choose $M$ to satisfy the condition ((ref)) of Proposition (ref). In light of Proposition (ref), we choose $M$ to satisfy ((ref)). By Minkowski's inequality, it is sufficient to have $M$ satisfy

align*[align* omitted — 387 chars of source]

Such a bound may be obtained by the plug-in Schennach2015. Specifically, one can choose

align[align omitted — 490 chars of source]

where $\widehat\varphi_X$, $\widehat\varphi_X^{(2)}$, and $\widehat\varphi_X^{(4)}$ can be computed based on li/vuong:1998 and its extensions -- see Appendix (ref) and Appendix (ref), respectively.

{\bf Tolerance Level $\eta$:} In light of Proposition (ref), we can choose the tolerance level $\eta$ in the following manner. By ((ref)), we can satisfy condition ((ref)) of Theorem (ref) if $\frac{1.086435\pi^{-1/4}}{\sqrt{2q+3}}\sqrt{M\sum_{j=q+1}^{\infty}(2j+1)^{-3}} \leq \eta$. Thus, having selected the smoothness bound $M$ and the sieve dimension $q$, one can set $\eta$ to

align[align omitted — 121 chars of source]

{\bf Tolerance Levels $\delta(t_1),\dots,\delta(t_L)$:} In light of Proposition (ref), we can choose the tolerance levels $\delta(t_1),\dots,\delta(t_L)$ of approximation errors in the following manner. By ((ref)), we can satisfy condition ((ref)) of Theorem (ref) if $\frac{1.086435\pi^{-1/4}}{\sqrt{2\pi}}\cdot\sqrt{M\sum_{j=q+1}^{\infty}(2j+1)^{-3}}\left(\frac{\mathbb{E}_P\left[|Y_1|\right]}{\sqrt{2q+3}}+1\right)\leq \delta(t_l)$ for all $l \in \{1,\dots,L\}$. Thus, having selected the smoothness bound $M$ and the sieve dimension $q$, one can set $\delta(t_l)$ to

align[align omitted — 157 chars of source]

for all $l \in \{1,\dots,L\}$, where we replaced $1.086435\pi^{-1/4}$ by one for slackness accounting for estimation of $\mathbb{E}_P\left[|Y_1|\right]$ by $\mathbb{E}_n\left[|Y_1|\right]$.

{\bf Sieve Dimension $q$}: The sieve dimension $q$ can be chosen by adapting a bandwidth selection method suggested in bissantz/lutz/holzmann/munk:2007 in density deconvolution with known error distribution; similar bandwidth selection rules are also used in kato/sasaki:2018 and adusumilli/otsu/whang:2017 in the deconvolution literature. Given $q$, choose tolerance levels $\eta, \delta (t_{1}),\dots,\delta(t_{L})$ depending on $q$, and then construct the upper and lower functions $f^{U}(x) = f_{q}^{U}(x)$ and $f^{L}(x) = f_{q}^{L}(x)$ according to Algorithm (ref). Then use the midpoint $\hat{f}_{q}(x) = \{ f_{q}^{U}(x) + f_{q}^{L}(x) \}/2$ as a surrogate of a point estimate of $f(x)$. Realizing that a sieve dimension corresponds to the reciprocal of a bandwidth, we suggest the following rule to choose $q$. Construct a candidate set for $q$ as $\{ q_{\min}, \dots, q_{\max} \}$, and compute the $L^{\infty}$-distance between the density estimates with adjacent sieve dimensions, $d_{q,q+1}^{(\infty)} = \sup_{x \in I} |\hat{f}_{q+1}(x) - \hat{f}_{q}(x)|$. Then we choose the smallest $q$ such that $d_{q,q+1}^{(\infty)}$ is larger than $\rho d_{q_{\min},q_{\min + 1}}^{(\infty)}$ for some $\rho > 1$ (or alternatively we can choose the largest $q$ such that $d_{q,q+1}^{(\infty)}$ is smaller than $\rho d_{q_{\min},q_{\min + 1}}^{(\infty)}$). In practice, it is recommended to make use of visual information on how $d_{q,q+1}^{(\infty)}$ behaves as $q$ decreases when determining the sieve dimension.

Implementation

When we implement the Anderson-Rubin-type inference, we would generally sweep across the parameter set $\Theta^{q+1}$ of sieve coefficients, and this operation may demand long computational time. However, we do not need to conduct the test at every point in $\Theta^{q+1}$, because properties of probability density functions and characteristic functions together with the sieve approximation property rule out substantially many elements of $\Theta^{q+1}$. From the property $f \geq 0$ of probability density functions and the approximation bound ((ref)), we can impose the restriction

equation[equation omitted — 215 chars of source]

Similarly, from the property $\left[ \mathcal{F}f \right](0) =1$ of characteristic functions and the approximation bound ((ref)), we can impose the restriction

equation[equation omitted — 243 chars of source]

Since the Hermite function is an eigenfunction of $\mathcal{F}$, ((ref)) can be simplified when the Hermite orthonormal basis $\boldsymbol{\Psi}$ (see Section (ref)) is used. Specifically, ((ref)) reduces to

equation[equation omitted — 275 chars of source]

Note that the left-hand side of ((ref)) does not require to compute an integral unlike that of ((ref)), which is a major advantage of using the Hermite orthonormal basis in the context of deconvolution. Use of the constraints ((ref)) and ((ref))/((ref)) is motivated by the definition of the confidence band ((ref)) consisting only of “density functions” $f \in \mathcal{L}$ which indexes the set $\mathbf{B}_{q+1,\eta}(f)$ of possible values of $\boldsymbol{\theta}$.

We also remark that we do not need to conduct a grid search for the purpose of drawing confidence bands. In fact, solving $\min_{\boldsymbol{\theta} \in \Theta^{q+1}} \boldsymbol{\psi}(x)^T \boldsymbol{\theta}$ subject to $T(\boldsymbol{\theta}) \leq c(\alpha)$ as well as ((ref)), ((ref)), and ((ref))/((ref)) yields the lower bound of the confidence band up to the approximation error $\eta$. Similarly, solving $\max_{\boldsymbol{\theta} \in \Theta^{q+1}} \boldsymbol{\psi}(x)^T \boldsymbol{\theta}$ subject to $T(\boldsymbol{\theta}) \leq c(\alpha)$ as well as ((ref)), ((ref)), and ((ref))/((ref)) yields the upper bound of the confidence band up to the approximation error $\eta$. Accounting for these points, we propose the following implementation algorithm.

algorithm[algorithm omitted — 923 chars of source]

We remark that, in computing the test statistic $T(\boldsymbol{\theta})$ in steps 2 and 3 of the algorithm above, the use of the Hermite orthonormal basis element $\psi = \psi_j$, $j=0,1,\dots$ simplifies ((ref)) and ((ref)) to

align*[align* omitted — 517 chars of source]

respectively. As such, one need not compute an integral to obtain the test statistic $T(\boldsymbol{\theta})$. This convenient property again follows from the fact that the Hermite function is an eigenfunction of $\mathcal{F}$.

Simulation Studies

In this section, we present and discuss finite-sample performance of the proposed method by simulation studies. Simulation outcomes that we present include the size under the null of the true distribution, the power under alternative distributions, and the lengths of confidence bands. The lengths will be further decomposed into the bias bound $\eta$ and the remaining lengths due to the stochastic part.

Simulation Setting

We employ three distribution families to generate the latent variable $X$ -- the normal distribution, the skew normal distribution, and the $t$ distribution. We employ the skew normal distribution and the $t$ distribution to see whether our method is effective for asymmetric distributions and super-Gaussian tails, respectively. Specifically, we generate a random sample of $(X,U_1,U_2)$ mutually independently according to the marginal laws:

align*[align* omitted — 377 chars of source]

Here, $N(\xi_1,\xi_2^2)$ denotes the normal distribution with mean $\xi_1$ and variance $\xi_2^2$, $SN(\xi_1,\xi_2,\xi_3)$ denotes the skew normal distribution with location $\xi_1$, scale $\xi_2$, and shape $\xi_3$, and $t_{\xi_4}$ denotes the $t$ distribution with $\xi_4$ degrees of freedom. The distribution parameters for the latent variable $X$ are set to $(\xi_1,\xi_2)=(0,1)$ for Model 1, $(\xi_1,\xi_2,\xi_3)=(0,1,1)$ for Model 2, and $\xi_4=5$ for Model 3. The choice of the normal error distribution, which is an instance of super-smooth distributions, imposes a difficult case in deconvolution -- see li/vuong:1998. The error variance parameters are set to $\sigma_{U_1} = \sigma_{U_2} = 0.5$ in each of the three models. We conduct experiments with three sample sizes $n=$ 250, 500, and 1,000, and run 2,500 Monte Carlo iterations for each set of simulations.

We follow the practical guideline provided in Appendix (ref) to construct confidence bands. As remarked previously, the smoothness bound $M \in (0,\infty)$ as well as the sieve dimension $q \in \mathbb{N}$ can be imposed by a researcher in the spirit of honest inference armstrong/kolesar:2018 -- see Section (ref). We experiment with the tuning parameters $q \in \{5,7,9\}$. The function classes are defined by ((ref)) with $M=15$, $25$, and $35$ for Models 1, 2, and 3, respectively. The frequency bound is set to $T=5$ and the number of frequency grid points is set to $L=50$. The interval on which the confidence band is formed is set to $I = \left[\mathbb{E}[X]-2\sqrt{\mathrm{Var}(X)},\mathbb{E}[X]+2\sqrt{\mathrm{Var}(X)}\right]$, where $\mathbb{E}[X]$ and $\mathrm{Var}(X)$ are the theoretical mean and the theoretical variance, respectively, of $X$ under the relevant model. To enjoy favorable speed of computation for numerous Monte Carlo iterations, we use the conservative critical value $c(\alpha)$. The level is set to $\alpha = 0.05$ throughout.

Simulation Results

Figure (ref) (A) shows the simulated frequencies that the confidence band formed under Model 1 covers alternative probability density functions for $N(\xi_1,\xi_2^2)$ indexed by location parameter values $\xi_1 \in [0.0,1.0]$ while the scale parameter is fixed at the true value $\xi_2=1.0$. The coverage frequency under $\xi_1=0.0$ indicate (the complement of) the size, whereas the coverage frequencies under $\xi_1 \in (0.0,1.0]$ indicate (the complement of) the power. Similarly, Figure (ref) (B) shows the simulated frequencies that the confidence band formed under Model 1 covers alternative probability density functions for $N(\xi_1,\xi_2^2)$ indexed by scale parameter values $\xi_2 \in [1.0,2.0]$ while the location parameter is fixed at the true value $\xi_1=0.0$. These results show the correct size and increasing power. The size entails over-coverage, which is still consistent with our theory on size control.

figure[figure omitted — 964 chars of source]
figure[figure omitted — 1,144 chars of source]

Figures (ref) (A) and (ref) (B) show analogous results to Figure (ref) (A) except that Model 2 and Model 3, respectively, are used instead of Model 1. For Model 2, the shape parameter is fixed at the true value $\xi_3=1.0$. These results evidence that the proposed method is similarly effective for the cases where the latent variable $X$ follows asymmetric distributions or distributions with super-Gaussian tails.

We next present average lengths of the confidence bands on $I$, and their decomposition into the bias bound $\eta$ and the remaining length due to the stochastic part. Table (ref) summarizes results on the lengths. There are a couple of features in these results that deserve discussions. First, observe that the confidence bands shrink as sample size increases when the sieve dimension $q$ is held fixed. On the other hand, the bias bound $\eta$ remains invariant across sample sizes while the sieve dimension $q$ is fixed. These results are natural features of our approach. Second, observe that the bias bound $\eta$ decreases as the sieve dimension $q$ increments.

table[table omitted — 1,840 chars of source]
figure[figure omitted — 1,587 chars of source]

Finally, we display instances of confidence bands in Figure (ref). The gray shades indicate the confidence bands including the bias bound and the stochastic parts together. The internal dark gray shades include only the stochastic parts. We also plot the true density functions and Li-Vuong estimates (with the choice of tuning parameter according to Appendix (ref)) as solid and dashed curves, respectively. While such instances of confidence bands will not tell us any evidence on the statistical properties, they at least inform how a confidence band may look in applications.

Simulations with Nonparametric Bootstrap

In this paper, we propose a novel inference method based on Kotlarski's identity. Under the absence of an inference method prior to our present paper, some existing applied papers that use Kotlarski's identity conduct the nonparametric bootstrap of the Li-Vuong estimator to draw confidence intervals though there is no theoretical guarantee for such a method to work. In the current subsection, we use simulation studies to assess the coverage performance for such a bootstrap approach. For the purpose of comparison, we use the same data generating processes as in the previous subsection, namely Model 1, Model 2, and Model 3.

Table (ref) summarizes the simulated frequencies that the na\"ive bootstrap confidence interval covers the true probability density function at locations $x \in \mathbb{E}[X]+\{-2,1,0,1,2\}\sqrt{\mathrm{Var}(X)}$. The tuning parameter is selected based on a commonly used approach in this literature -- see Appendix (ref). The coverage frequencies are close to the nominal probabilities only near the tails of the distributions, e.g., $x = \mathbb{E}[X] \pm 2\sqrt{\mathrm{Var}(X)}$, under Model 1 and Model 2. On the other hand, the na\"ive bootstrap suffers from under-coverage near the center of the distributions under Model 1 and Model 2. Furthermore, the na\"ive bootstrap suffers from even more severe under-coverage both near the center of the distribution and near the tails of the distribution under Model 3. These results suggest that one should substitute our proposed method for the traditional na\"ive bootstrap approach.

table[table omitted — 1,089 chars of source]

Conclusion

Since its introduction to econometrics by \citet*{li/vuong:1998}, Kotlarski's identity \citep*{kotlarski1967} -- see also rao:1992 -- has been widely used in empirical economics. Examples include applications to empirical auctions LiPerrigneVuong2000,krasnokutskaya:2011, income dynamics bonhomme/robin:2010, and labor economics cunha/heckman/navarro:2005,cunha/heckman/schennach:2010,bonhomme/sauder:2011,kennan/walker:2011. Despite its popular use in applications, a method of inference based on Kotlarski's identity has long been missing in the literature. After twenty years since \citet*{li/vuong:1998}, we now propose a method of inference based on Kotlarski's identity. Specifically, we develop confidence bands for the probability density function $f_X$ of $X$ in the repeated measurement model where two measurements $(Y_1,Y_2)$ of unobserved variable $X$ are available in data with additive independent errors, $U_1 = Y_1-X$ and $U_2 = Y_2-X$.

Our construction of confidence bands can be summarized as follows. First, we derive linear complex-valued moment restrictions based on Kotlarski's identity. Second, we let the Hermite polynomial sieve approximate unknown probability density functions. Third, for a given sieve dimension and for a given class for probability density functions, we compute a bias bound for the linear complex-valued moment restrictions, and slack the linear complex-valued moment restrictions by this bias bound. Fourth, we compute the uniform norm of the self-normalized process of the slacked linear complex-valued moment restrictions as the test statistics for each point in a set of sieve coefficients. Fifth, inverting this test statistic yields a confidence set of sieve approximations to possible probability density functions. Sixth, for a given sieve dimension and for a given class for probability density functions, we compute a bias bound for sieve approximations of probability density functions, and the desired confidence band is obtained by uniformly enlarging the set of sieve approximations by this bias bound.

We not only provide a method that works, but also care for its practicality. The Fourier transform and the inverse Fourier transform operations are known to be computationally costly in the deconvolution literature. By exploiting the property of the Hermite functions as eigen-functions of the Fourier transform operator, we propose to let the Hermite polynomial sieve approximate both the density and characteristic functions without having to implement numerical integrations within each iteration of a numerical optimization routine. This convenient feature of the proposed method saves computational resources. Furthermore, we also exploit a couple of other convenient properties of the Hermite functions (namely the Schr\"odinger equation for a harmonic oscillator and a pair of recursive equations), and consequently obtain informative bias bounds and thus informative inference. With these practical features of our method, simulation studies indeed conclude reasonably fast with informative inference results. The results evidence the efficacy of the proposed method. Since Kotlarski's identity is one of the most popular methods in a number of applied fields, including empirical auctions, income dynamics, and labor economics, we hope that our method will contribute to the practice of economic analyses in these and other topics.