Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
72,892 characters · 16 sections · 53 citation commands
Inference based on Kotlarski's Identity
Empirical researchers are often interested in recovering features of unobserved variables in economic models. Kotlarski's identity \citep*{kotlarski1967} -- see also rao:1992 -- is one of the most popular tools used to identify probability density functions of unobserved latent variables. Since its first introduction to econometrics by \citet*{li/vuong:1998}, Kotlarski's identity has been widely used in economics when data admit repeated measurements. Examples of research topics that use Kotlarski's identity include, but are not limited to, empirical auctions LiPerrigneVuong2000,krasnokutskaya:2011, income dynamics bonhomme/robin:2010, and labor economics cunha/heckman/navarro:2005,cunha/heckman/schennach:2010,bonhomme/sauder:2011,kennan/walker:2011. In these applications, researchers are interested in identifying the probability density function $f_X$ of a latent variable $X$ among others. The variable $X$ of interest is not observed in data, but two measurements $(Y_1,Y_2)$ are available in data with classical errors, $U_1 = Y_1-X$ and $U_2 = Y_2-X$. Kotlarski's identity is a nonparametric identifying restriction for the probability density function $f_X$ of $X$ implied by this setup.
The existing econometric literature on Kotlarski's identity focuses on identification and consistent estimation of $f_X$ and related objects li/vuong:1998,Li2002,schennach:2004,Schennach2004et,Schennach2008,bonhomme/robin:2010,Evdokimov2010,ZindeWalsh2014,SongSchennachWhite2015,FirpoGalvaoSong2017 -- also see surveys on this literature by Chen/Hong/Nekipelov:2011 and Schennach:2016. On the other hand, satisfactory inference methods for $f_X$ are missing in this literature -- in fact, even the sharp rate of convergence is unknown for the estimators based on Kotlarski's identity under unrestrictive assumptions, and hence a limit distribution result is unavailable under such assumptions. Indeed some empirical papers implement nonparametric bootstrap without a theoretical guarantee. In light of the demands by the empirical researchers for an inference method, and given the current unavailability of theoretically supported methods of inference, we propose a method of inference based on Kotlarski's identity in this paper.
This paper proposes an inference method based on Kotlarski's identity by developing a confidence band for $f_X$. Our construction of confidence bands works as follows. First, we derive linear complex-valued moment restrictions based on Kotlarski's identity. Second, we let the Hermite polynomial sieve chen:2007 approximate unknown probability density functions. Third, for a given sieve dimension and for a given class of probability density functions, we compute a bias bound for the linear complex-valued moment restrictions, and slack the linear complex-valued moment restrictions by this bias bound. Fourth, applying \citet*{chernozhukov/chetverikov/kato:2017}, we compute the uniform norm of the self-normalized process of the slacked linear complex-valued moment restrictions as the test statistic for each point in a set of sieve coefficients. Fifth, inverting this test statistic in the spirit of anderson/rubin:1949 yields a confidence set of sieve approximations to possible probability density functions. Sixth, for a given sieve dimension and for a given class for probability density functions, we compute a bias bound for sieve approximations of probability density functions, and the desired confidence band is obtained by uniformly enlarging the set of sieve approximations by this bias bound.
The process of identifying $f_X$ in additive measurement error models is called deconvolution -- for solving convolution integral equations. There are a number of existing papers on nonparametric inference in deconvolution. bissantz/lutz/holzmann/munk:2007, bissantz/holtzmann:2008, es/gugushvili:2008, lounici2011, and schmidt-hieber2013 develop uniform confidence bands for $f_X$ under the assumption of known error distributions.\footnote{These paper are based on the literature on deconvolution under known error distribution carroll/hall:1988,stefanski/carroll:1990,fan:1991b,carrasco2011spectral. fan:1991a develops a point-wise asymptotic inference result in this framework.} In most economic applications, however, it is not plausible to assume that the error distributions are known. More recently, kato/sasaki:2018 and adusumilli/otsu/whang:2017 develop uniform confidence bands for $f_X$ and the distribution function, respectively, without assuming that the error distributions are known, but they both assume that at least one error distribution is symmetric.\footnote{These paper are based on the literature on deconvolution under unknown error distribution with auxiliary data or symmetric error distributions diggle/hall:1993,horowitz/markatou:1996,neumann:1997,efromovich:1997,delaigle/hall/meister:2008,johannes:2009,comte/lacour:2011,delaigle/hall:2015.} Kotlarski's identity is a powerful device for new identification results which require neither the known error distribution assumption nor the symmetric error distribution assumption. This useful feature attracts many economic applications including those listed above, but no econometrician has developed a method of inference in this framework for twenty years ever since its first introduction by \citet*{li/vuong:1998} until our present paper.
It it not surprising that such an inference method has been missing for long in the literature, given the technical difficulties of the problem. Deconvolution is an ill-posed inverse problem, and inference under this problem is known to be challenging -- see bissantz/lutz/holzmann/munk:2007,bissantz/holtzmann:2008,lounici2011,horowitz/lee:2012,hall/horowitz:2013,schmidt-hieber2013,adusumilli/otsu/whang:2017,kato/sasaki:2017b,kato/sasaki:2018,babii:2018,chen/christensen:2018 for existing papers developing confidence bands in ill-posed inverse problems for example. We take a robust inference approach \`a la anderson/rubin:1949, and directly work with the moment restrictions based on Kotlarski's identity. A positive side product of taking this approach is that we do not need to assume the non-vanishing characteristic functions (i.e., we do not need the completeness), which is commonly assumed for nonparametric identification or inversion.
It is also worth mentioning that we chose to use the Hermite polynomial sieve among other sieves in this paper. The Hermite polynomial sieve has been in fact already known in the literature to be useful to approximate “smooth density with unbounded support” chen:2007 -- also see her discussion of gallant/nychka:1987 therein. In addition to this known advantage, we also find this sieve particularly useful for the deconvolution problem. Note that the deconvolution problem involves applications of the Fourier transform operation and the inverse Fourier transform operation. To our convenience, the Hermite functions are eigen-functions of the Fourier transform operator. While we deal with simultaneous restrictions in terms of density and characteristic functions, we can use the Hermite polynomial sieve to approximate both the density and characteristic functions without having to apply the Fourier transform or the Fourier inverse because of the eigen-function property. This convenient property saves computational time and resources as costly numerical integration within each iteration of a numerical optimization routine would be necessary if any other sieve were used. Furthermore, we find that a couple of properties of the Hermite functions (namely the Schr\"odinger equation for a harmonic oscillator and a pair of recursive equations) can be exploited to obtain informative bias bounds of the Hermite polynomial sieve, which in turn contributes to informative inference we establish in this paper.
The rest of the paper is organized as follows. Section (ref) derives linear complex-valued moment restrictions based on Kotlarski's identity. Section (ref) presents how to compute the confidence band. Section (ref) presents asymptotic properties of the confidence band. Section (ref) discusses practical considerations. Section (ref) illustrates simulation studies. The paper concludes in Section (ref). All mathematical derivations and details are delegated to the appendix.
Consider the repeated measurement model
where $Y_1$ and $Y_2$ are observed, but none of $X$, $U_1$, or $U_2$ is observed. We are interested in making inference on the probability density function $f_X$ of $X$. We equip this model with the following assumption.
This assumption is standard in the literature on identification and estimation based on Kotlarski's identity li/vuong:1998. In fact, the existing literature imposes an additional assumption, namely the identification condition (non-vanishing characteristic function or the completeness) -- see Lemma (ref) ahead for a specific condition. We do not invoke such an identification assumption for the purpose of identification-roust inference -- see Remark (ref) ahead for further details.
We now fix basic notations. In what follows, $\mathbb{E}_P$ and $\mathbb{V}_P$ denote the expectation and variance operators, respectively, with respect to a joint distribution $P$ of $(Y_1,Y_2)$. Analogously, $\mathbb{E}_n$ and $\mathbb{V}_n$ denote the expectation and variance operators, respectively, with respect the empirical distribution of $n$ independent copies of $(Y_1,Y_2)$. We let $i=\sqrt{-1}$ denote the imaginary unit. For the set of absolutely integrable functions, $\mathcal{L}^1$, we define the Fourier transform $\mathcal{F}$ on $\mathcal{L}^{1}$ by $[\mathcal{F}f](t)=\int_{-\infty}^{\infty} e^{itx}f(x)dx$, and its inverse transform is $[\mathcal{F}^{-1}\phi](x)=\frac{1}{2\pi}\int_{-\infty}^{\infty} e^{-itx}\phi(t)dt$ -- see folland2007. In light of Assumption (ref) (i), we let $f_X$, $f_{U_1}$, and $f_{U_2}$ denote the density functions of $X$, $U_1$, and $U_2$, respectively. Further, we denote the characteristic functions of them by $\phi_X=\mathcal{F}f_X$, $\phi_{U_1}=\mathcal{F}f_{U_1}$, and $\phi_{U_2}=\mathcal{F}f_{U_2}$. We first review the existing result of the identification.
This lemma presents Kotlarski's identity due to kotlarski1967 -- see also rao:1992. Since it is stated as a lemma, Kotlarski's identity is also known as Kotlarski's lemma or the lemma of Kotlarski in the econometrics literature. li/vuong:1998 first introduced it into econometrics and statistics, followed by a series of extensions Li2002,schennach:2004,bonhomme/robin:2010,Evdokimov2010. Some of these extensions relax the assumptions for identification and estimation in various ways. We do not need to rely on the prototypical assumptions for our purpose of inference, even though they are stated in Lemma (ref) for convenience of a concise review. Lemma (ref) shows that the characteristic function $\phi_X$ of $X$ is explicitly identified by the joint distribution of $(Y_1,Y_2)$. Under the additional assumption of absolutely integrable characteristic function $\phi_X$, the formula $f_X=\mathcal{F}^{-1}\phi_X$ in turn yields the identification of the probability density function $f_X$ of $X$.
Uniform convergence rates for the estimator of $f_X$ based on Kotlarski's identity are discovered in the existing literature li/vuong:1998,Li2002,schennach:2004,bonhomme/robin:2010,Evdokimov2010, but the sharp rates under unrestrictive assumptions are still unknown. In particular, limit distribution results under such assumptions are still unknown in the existing literature. This paper does not aim to derive a non-degenerate limit distribution for any estimator, but it aims to conduct an inference on $f_X$. With this said, our proposed inference does not rely on an explicit identifying formula. We argue that rewriting Kotlarski's lemma in terms of moment restrictions suffices and serves even more conveniently for the sake of conducting inference.
A proof is provided in Appendix (ref).
Our objective is to construct a confidence band for the probability density function $f_X$ of $X$ on an interval $I \subset \mathbb{R}$. The construction procedure is based on the linear complex-valued moment restriction ((ref)). Throughout, we focus on the set of probability density functions given by
For this set of candidate probability density functions, we use $\mathcal{L}^1$ for applying the Fourier transform and the inverse, whereas $\mathcal{L}^2$ is used to approximate $\mathcal{L}$ by an orthonormal basis $\Psi=\{\psi_j \ : \ j=0,1,\ldots\}$ of $\mathcal{L}^2$ -- see Section (ref) for the example of the Hermite basis.
We use a $(q+1)$-dimensional sieve basis $\{\psi_0,\ldots,\psi_q\} \subset \Psi = \{\psi_j : j =0,1,\ldots\}$, with $\psi_j \in \mathcal{L}^1 \cap \mathcal{L}^2$ and $\mathcal{F} \psi_j \in \mathcal{L}^1$ for each $j \in \mathbb{N}$, to approximate the probability density function $f_X$. Let $\Theta^{q+1}\subset \mathbb{R}^{q+1}$ be a compact set, and write $\boldsymbol{\psi}=(\psi_0,\dots,\psi_q)^T$. With a uniform tolerance level $\eta > 0$, each $f \in \mathcal{L}$ is approximated by $ x \mapsto \boldsymbol{\psi}(x)^T \boldsymbol{\theta} $ for some $\boldsymbol{\theta}=(\theta_0,\ldots,\theta_q)^T \in \Theta^{q+1}$, i.e., $ \sup_{x \in I} \left| f(x) - \boldsymbol{\psi}(x)^T\boldsymbol{\theta} \right| \leq \eta. $ The set of values of the sieve coefficients $\boldsymbol{\theta} \in \Theta^{q+1}$ approximating a probability density function $f \in \mathcal{L}$ in this manner is denoted by
We next incorporate the linear complex-valued moment restrictions ((ref)) in this sieve framework. For every function $\psi \in \mathcal{L}$ and for every frequency $t \in \mathbb{R}$, define
where $\phi = \mathcal{F} \psi$. Note that $\mathrm{Re}(\cdot)$ (resp., $\mathrm{Im}(\cdot)$) denotes the real (resp., imaginary) part of a complex number. Further, stack these functions across $\psi \in \{\psi_0,\dots,\psi_q\}$ to define the random vector
With these notations, we now represent the linear complex-valued moment restrictions ((ref)) for the sieve approximation by
for all $t \in [-T,T]$ for $T \in (0,\infty)$, where $\delta(t)>0$ is the tolerance level of sieve approximation error for each $t \in [-T,T]$.
Our construction of the confidence band is based on a test statistic that quantifies the extent of deviation from the moment inequalities ((ref)). To construct a feasible test statistic, we use a grid $\{t_1,\dots,t_L\} \subset [-T,T]$ of $L$ frequencies. Define the test statistic by
for each $\boldsymbol{\theta}\in\Theta^{q+1}$. Let $\alpha\in(0,1/2)$ be fixed. We define the critical value $c(\alpha,\boldsymbol{\theta})$ of this statistic $T(\boldsymbol{\theta})$ by the conditional $(1-\alpha)$-th quantile of the multiplier bootstrap statistic $$ \sqrt{n}\max_{1\leq l\leq L}\max\left\{\frac{\left|\mathbb{E}_n[\boldsymbol{\epsilon}(\mathbf{R}_{t_l}-\mathbb{E}_n[\mathbf{R}_{t_l}])]^T\boldsymbol{\theta}\right|}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_n(\mathbf{R}_{t_l})\boldsymbol{\theta}}},\frac{\left|\mathbb{E}_n[\boldsymbol{\epsilon}(\mathbf{I}_{t_l}-\mathbb{E}_n[\mathbf{I}_{t_l}])]^T\boldsymbol{\theta}\right|}{\sqrt{\boldsymbol{\theta}^T\mathbb{V}_n(\mathbf{I}_{t_l})\boldsymbol{\theta}}} \right\} $$ given the data, where $\boldsymbol{\epsilon}_1,\ldots,\boldsymbol{\epsilon}_n$ are independent standard normal random variables independent of the data. As a more conservative yet simpler alternative following chernozhukov/chetverikov/kato:2017, we may define the critical value as
where $\Phi$ is the cumulative distribution function of the standard normal distribution. Our confidence band for the density function of $X$ is given by the $I$-restriction of
where $\mathbf{B}_{q+1,\eta}(f)$ is defined in ((ref)).
A practical procedure to obtain this confidence band is outlined as Algorithm (ref) below. While this algorithm prescribes a general procedure, we also elaborate on details of practical considerations in Section (ref), where we introduce the Hermite orthonormal sieve (Section (ref)), propose concrete choice rules for the tuning parameters (Section (ref)), and present a more concrete algorithm for these sieve and tuning parameters (Section (ref)).
In this section, we present theoretical properties of the confidence band ((ref)). Let $\mathcal{P}$ denote a given space to which the joint distribution of $(Y_1,Y_2)$ belongs. For every $P\in\mathcal{P}$, define the identified set
as the set of density functions for which the linear complex-valued moment restriction ((ref)) is satisfied. Furthermore, we define the sieve-approximation counterpart of $\mathcal{L}_0(P)$ by
where $\mathbf{B}_{q+1,\eta}(f)$ is defined in ((ref)). Here, the infimum over the empty set is understood to be the infinity.
The current section is structured as follows. First, we establish the size control for the confidence band ((ref)) to contain the approximation set $\mathcal{L}_0^\ast(P)$ in Section (ref). Second, we establish the containment of the identified set by the approximation set (i.e., $\mathcal{L}_0(P) \subset \mathcal{L}_0^\ast(P)$) in Section (ref). These two pieces of the results together show the validity of the confidence band ((ref)) to contain the identified set $\mathcal{L}_0(P)$. Lastly, Section (ref) presents power properties with local alternatives. Throughout, we assume to observe $n$ i.i.d. copies of $(Y_1,Y_2)$ drawn from $P \in \mathcal{P}$.
We make the following assumption for a uniform size control.
A proof is provided in Appenedix (ref). This theorem guarantees the size control for the event of the inclusion of the density function in the approximate identified set $\mathcal{L}_0^\ast(P)$ as opposed to the identified set $\mathcal{L}_0(P)$. The next section presents conditions under which the approximate identified set $\mathcal{L}_0^\ast(P)$ contains the identified set $\mathcal{L}_0(P)$ for every possible joint distribution $P \in \mathcal{P}$.
Throughout, we equip $\mathcal{L}^2$ with the inner product $\langle{\cdot,\cdot}\rangle$ defined by $$ \langle{f_1,f_2}\rangle = \int_\mathbb{R} f_1(x) f_2(x) dx. $$
In Section (ref), we propose a concrete orthonormal basis $\Psi$ and the set $\Theta^{q+1}$ of coefficients to satisfy Assumption (ref). The following theorem provides a guide for choices of $\eta$ and $\delta(t_1),\dots,\delta(t_L)$ such that $\mathcal{L}_0^\ast(P)$ contains $\mathcal{L}_0(P)$ for every possible $P \in \mathcal{P}$.
A proof is provided in Appendix (ref). In Section (ref), we provide concrete evaluations of the left-hand side of ((ref)) and ((ref)) under a concrete orthonormal basis $\Psi$, namely the Hermite orthonormal basis. Putting Theorems (ref) and (ref) together, we obtain the following result on the validity of the confidence band.
We introduce the following short-hand notation for the random variable defined as the maximum deviation of the sample variance from the population variance: $$ B_{V}= \sup_{\boldsymbol{\theta}\in \mathbf{B}_{q+1,\eta}(f)}\max_{l=1,\ldots,L}\max\left\{\boldsymbol{\theta}^T(\mathbb{V}_n(\mathbf{R}_{t_l})-\mathbb{V}_P(\mathbf{R}_{t_l}))\boldsymbol{\theta},\boldsymbol{\theta}^T(\mathbb{V}_n(\mathbf{I}_{t_l})-\mathbb{V}_P(\mathbf{I}_{t_l}))\boldsymbol{\theta}\right\}. $$ The following theorem shows a power property of our proposed inference method.
A proof is provided in Appendix (ref). According to this theorem, for any density function $f \in \mathcal{L}$ such that at least one of the moment inequalities violated at some frequency point $t_\ast$ in the grid $\{t_1,\dots,t_L\}$, then the probability that this density function does not belong to the confidence band is bounded below by $\mathbb{P}_P\left(B_{V}\leq \nu\right)-\frac{1}{1+b}$. Choosing sequences of $\nu$ and $b$ so that $\mathbb{P}_P\left(B_{V}\leq \nu\right)-\frac{1}{1+b} \rightarrow 1$ as $n \rightarrow \infty$, therefore, this theorem implies the consistency against all fixed alternatives.
The current section presents a guide to practice. Construction of the confidence band and its theoretical properties are presented in Sections (ref) and (ref) with abstract objects. These objects in particular include an orthonormal basis $\Psi = \{\psi_j \ : \ j=0,1,\ldots\}$, the sieve dimension $q$, the set $\Theta^{q+1}$ of sieve coefficients, and the tolerance levels, $\eta$, $\delta(t_1),\dots, \delta(t_L)$, of approximation errors. Section (ref) presents concrete choices of $\Psi = \{\psi_j \ : \ j=0,1,\ldots\}$ and $\Theta^{q+1}$. Section (ref) presents a concrete data-driven procedure for selecting the tolerance levels $\eta$, $\delta(t_1),\dots, \delta(t_L)$. Section (ref) presents a concrete implementation procedure for constructing the confidence band with these choices of the objects.
There is a large extent of freedom of choice for an orthonormal basis $\Psi$ -- see chen:2007 for a list of options. We recommend the Hermite orthonormal basis in particular for its convenient properties and its nice compatibility with the deconvolution framework -- a Hermite function is an eigenfunction of the Fourier transform and the Fourier inverse.\footnote{The Hermite orthonormal sieve is not location invariant, and hence we recommend to location- and scale normalize the observed data using the empirical moments.} The Hermite functions take the form
$j=0,1,\ldots$, where $H_j$ is the Hermite polynomial defined by $$ H_j(x)=(-1)^{j} \cdot \exp(x^2) \cdot \frac{d^{j}}{dx^{j}}\exp(-x^2). $$ The Hermite functions are the eigenfunctions of the Fourier transform operator, and specifically, $\phi_j=\mathcal{F}\psi_j=i^j\sqrt{2\pi}\psi_j$ holds. For any $q \in \mathbb{N}$ to be chosen below, have the set of sieve coefficients satisfy
These concrete choices are made for the sake of satisfying Assumption (ref) so we can use Theorem (ref).
A proof is provided in Appendix (ref). We also present the following proposition which provides approximation bounds for the condition of Theorem (ref) to guarantee the containment of the identified set $\mathcal{L}_0(P)$ by the approximate identified set $\mathcal{L}_0^\ast(P)$ for every possible joint distribution $P \in \mathcal{P}$.
A proof is provided in Appendix (ref). An admissible function class in terms of smoothness restriction can be specified by ((ref)). Note that this is analogous to the standard practice in the literature to work with Sovolev classes of functions. With the function class specified in this way, equations ((ref)) and ((ref)) prescribe possible choices of the tolerance levels $\eta,\delta(t_1),\dots,\delta(t_L)$ which admit $\mathcal{L}_0(P) \subset \mathcal{L}_0^\ast(P)$ for all $P \in \mathcal{P}$ per Theorem (ref). See Section (ref) for concrete choice procedures. Equation ((ref)) in addition suggests the worst approximation error for the characteristic function, which will be useful when we impose natural restrictions on the characteristic functions -- see Section (ref). Finally, we present the following proposition showing an alternative representation of the function class specification ((ref)).
A proof is provided in Appendix (ref). The components in the left-hand side of ((ref)) are estimable by li/vuong:1998 and its extensions, and thus Proposition (ref) provides a guideline for setting the smoothness bound parameter $M$ -- see Section (ref).
In this section, we provide example procedures of choosing the smoothness bound $M$ and the tolerance levels $\eta, \delta(t_1),\dots,\delta(t_L) \in (0,\infty)$ in finite sample. Although we present a data driven choice of the smoothness bound $M$ below, we remark that the smoothness bound $M \in (0,\infty)$ as well as the sieve dimension $q \in \mathbb{N}$ could be imposed by a researcher in the spirit of honest inference armstrong/kolesar:2018.
{\bf Smoothness Bound $M$:} We choose $M$ to satisfy the condition ((ref)) of Proposition (ref). In light of Proposition (ref), we choose $M$ to satisfy ((ref)). By Minkowski's inequality, it is sufficient to have $M$ satisfy
Such a bound may be obtained by the plug-in Schennach2015. Specifically, one can choose
where $\widehat\varphi_X$, $\widehat\varphi_X^{(2)}$, and $\widehat\varphi_X^{(4)}$ can be computed based on li/vuong:1998 and its extensions -- see Appendix (ref) and Appendix (ref), respectively.
{\bf Tolerance Level $\eta$:} In light of Proposition (ref), we can choose the tolerance level $\eta$ in the following manner. By ((ref)), we can satisfy condition ((ref)) of Theorem (ref) if $\frac{1.086435\pi^{-1/4}}{\sqrt{2q+3}}\sqrt{M\sum_{j=q+1}^{\infty}(2j+1)^{-3}} \leq \eta$. Thus, having selected the smoothness bound $M$ and the sieve dimension $q$, one can set $\eta$ to
{\bf Tolerance Levels $\delta(t_1),\dots,\delta(t_L)$:} In light of Proposition (ref), we can choose the tolerance levels $\delta(t_1),\dots,\delta(t_L)$ of approximation errors in the following manner. By ((ref)), we can satisfy condition ((ref)) of Theorem (ref) if $\frac{1.086435\pi^{-1/4}}{\sqrt{2\pi}}\cdot\sqrt{M\sum_{j=q+1}^{\infty}(2j+1)^{-3}}\left(\frac{\mathbb{E}_P\left[|Y_1|\right]}{\sqrt{2q+3}}+1\right)\leq \delta(t_l)$ for all $l \in \{1,\dots,L\}$. Thus, having selected the smoothness bound $M$ and the sieve dimension $q$, one can set $\delta(t_l)$ to
for all $l \in \{1,\dots,L\}$, where we replaced $1.086435\pi^{-1/4}$ by one for slackness accounting for estimation of $\mathbb{E}_P\left[|Y_1|\right]$ by $\mathbb{E}_n\left[|Y_1|\right]$.
{\bf Sieve Dimension $q$}: The sieve dimension $q$ can be chosen by adapting a bandwidth selection method suggested in bissantz/lutz/holzmann/munk:2007 in density deconvolution with known error distribution; similar bandwidth selection rules are also used in kato/sasaki:2018 and adusumilli/otsu/whang:2017 in the deconvolution literature. Given $q$, choose tolerance levels $\eta, \delta (t_{1}),\dots,\delta(t_{L})$ depending on $q$, and then construct the upper and lower functions $f^{U}(x) = f_{q}^{U}(x)$ and $f^{L}(x) = f_{q}^{L}(x)$ according to Algorithm (ref). Then use the midpoint $\hat{f}_{q}(x) = \{ f_{q}^{U}(x) + f_{q}^{L}(x) \}/2$ as a surrogate of a point estimate of $f(x)$. Realizing that a sieve dimension corresponds to the reciprocal of a bandwidth, we suggest the following rule to choose $q$. Construct a candidate set for $q$ as $\{ q_{\min}, \dots, q_{\max} \}$, and compute the $L^{\infty}$-distance between the density estimates with adjacent sieve dimensions, $d_{q,q+1}^{(\infty)} = \sup_{x \in I} |\hat{f}_{q+1}(x) - \hat{f}_{q}(x)|$. Then we choose the smallest $q$ such that $d_{q,q+1}^{(\infty)}$ is larger than $\rho d_{q_{\min},q_{\min + 1}}^{(\infty)}$ for some $\rho > 1$ (or alternatively we can choose the largest $q$ such that $d_{q,q+1}^{(\infty)}$ is smaller than $\rho d_{q_{\min},q_{\min + 1}}^{(\infty)}$). In practice, it is recommended to make use of visual information on how $d_{q,q+1}^{(\infty)}$ behaves as $q$ decreases when determining the sieve dimension.
When we implement the Anderson-Rubin-type inference, we would generally sweep across the parameter set $\Theta^{q+1}$ of sieve coefficients, and this operation may demand long computational time. However, we do not need to conduct the test at every point in $\Theta^{q+1}$, because properties of probability density functions and characteristic functions together with the sieve approximation property rule out substantially many elements of $\Theta^{q+1}$. From the property $f \geq 0$ of probability density functions and the approximation bound ((ref)), we can impose the restriction
Similarly, from the property $\left[ \mathcal{F}f \right](0) =1$ of characteristic functions and the approximation bound ((ref)), we can impose the restriction
Since the Hermite function is an eigenfunction of $\mathcal{F}$, ((ref)) can be simplified when the Hermite orthonormal basis $\boldsymbol{\Psi}$ (see Section (ref)) is used. Specifically, ((ref)) reduces to
Note that the left-hand side of ((ref)) does not require to compute an integral unlike that of ((ref)), which is a major advantage of using the Hermite orthonormal basis in the context of deconvolution. Use of the constraints ((ref)) and ((ref))/((ref)) is motivated by the definition of the confidence band ((ref)) consisting only of “density functions” $f \in \mathcal{L}$ which indexes the set $\mathbf{B}_{q+1,\eta}(f)$ of possible values of $\boldsymbol{\theta}$.
We also remark that we do not need to conduct a grid search for the purpose of drawing confidence bands. In fact, solving $\min_{\boldsymbol{\theta} \in \Theta^{q+1}} \boldsymbol{\psi}(x)^T \boldsymbol{\theta}$ subject to $T(\boldsymbol{\theta}) \leq c(\alpha)$ as well as ((ref)), ((ref)), and ((ref))/((ref)) yields the lower bound of the confidence band up to the approximation error $\eta$. Similarly, solving $\max_{\boldsymbol{\theta} \in \Theta^{q+1}} \boldsymbol{\psi}(x)^T \boldsymbol{\theta}$ subject to $T(\boldsymbol{\theta}) \leq c(\alpha)$ as well as ((ref)), ((ref)), and ((ref))/((ref)) yields the upper bound of the confidence band up to the approximation error $\eta$. Accounting for these points, we propose the following implementation algorithm.
We remark that, in computing the test statistic $T(\boldsymbol{\theta})$ in steps 2 and 3 of the algorithm above, the use of the Hermite orthonormal basis element $\psi = \psi_j$, $j=0,1,\dots$ simplifies ((ref)) and ((ref)) to
respectively. As such, one need not compute an integral to obtain the test statistic $T(\boldsymbol{\theta})$. This convenient property again follows from the fact that the Hermite function is an eigenfunction of $\mathcal{F}$.
In this section, we present and discuss finite-sample performance of the proposed method by simulation studies. Simulation outcomes that we present include the size under the null of the true distribution, the power under alternative distributions, and the lengths of confidence bands. The lengths will be further decomposed into the bias bound $\eta$ and the remaining lengths due to the stochastic part.
We employ three distribution families to generate the latent variable $X$ -- the normal distribution, the skew normal distribution, and the $t$ distribution. We employ the skew normal distribution and the $t$ distribution to see whether our method is effective for asymmetric distributions and super-Gaussian tails, respectively. Specifically, we generate a random sample of $(X,U_1,U_2)$ mutually independently according to the marginal laws:
Here, $N(\xi_1,\xi_2^2)$ denotes the normal distribution with mean $\xi_1$ and variance $\xi_2^2$, $SN(\xi_1,\xi_2,\xi_3)$ denotes the skew normal distribution with location $\xi_1$, scale $\xi_2$, and shape $\xi_3$, and $t_{\xi_4}$ denotes the $t$ distribution with $\xi_4$ degrees of freedom. The distribution parameters for the latent variable $X$ are set to $(\xi_1,\xi_2)=(0,1)$ for Model 1, $(\xi_1,\xi_2,\xi_3)=(0,1,1)$ for Model 2, and $\xi_4=5$ for Model 3. The choice of the normal error distribution, which is an instance of super-smooth distributions, imposes a difficult case in deconvolution -- see li/vuong:1998. The error variance parameters are set to $\sigma_{U_1} = \sigma_{U_2} = 0.5$ in each of the three models. We conduct experiments with three sample sizes $n=$ 250, 500, and 1,000, and run 2,500 Monte Carlo iterations for each set of simulations.
We follow the practical guideline provided in Appendix (ref) to construct confidence bands. As remarked previously, the smoothness bound $M \in (0,\infty)$ as well as the sieve dimension $q \in \mathbb{N}$ can be imposed by a researcher in the spirit of honest inference armstrong/kolesar:2018 -- see Section (ref). We experiment with the tuning parameters $q \in \{5,7,9\}$. The function classes are defined by ((ref)) with $M=15$, $25$, and $35$ for Models 1, 2, and 3, respectively. The frequency bound is set to $T=5$ and the number of frequency grid points is set to $L=50$. The interval on which the confidence band is formed is set to $I = \left[\mathbb{E}[X]-2\sqrt{\mathrm{Var}(X)},\mathbb{E}[X]+2\sqrt{\mathrm{Var}(X)}\right]$, where $\mathbb{E}[X]$ and $\mathrm{Var}(X)$ are the theoretical mean and the theoretical variance, respectively, of $X$ under the relevant model. To enjoy favorable speed of computation for numerous Monte Carlo iterations, we use the conservative critical value $c(\alpha)$. The level is set to $\alpha = 0.05$ throughout.
Figure (ref) (A) shows the simulated frequencies that the confidence band formed under Model 1 covers alternative probability density functions for $N(\xi_1,\xi_2^2)$ indexed by location parameter values $\xi_1 \in [0.0,1.0]$ while the scale parameter is fixed at the true value $\xi_2=1.0$. The coverage frequency under $\xi_1=0.0$ indicate (the complement of) the size, whereas the coverage frequencies under $\xi_1 \in (0.0,1.0]$ indicate (the complement of) the power. Similarly, Figure (ref) (B) shows the simulated frequencies that the confidence band formed under Model 1 covers alternative probability density functions for $N(\xi_1,\xi_2^2)$ indexed by scale parameter values $\xi_2 \in [1.0,2.0]$ while the location parameter is fixed at the true value $\xi_1=0.0$. These results show the correct size and increasing power. The size entails over-coverage, which is still consistent with our theory on size control.
Figures (ref) (A) and (ref) (B) show analogous results to Figure (ref) (A) except that Model 2 and Model 3, respectively, are used instead of Model 1. For Model 2, the shape parameter is fixed at the true value $\xi_3=1.0$. These results evidence that the proposed method is similarly effective for the cases where the latent variable $X$ follows asymmetric distributions or distributions with super-Gaussian tails.
We next present average lengths of the confidence bands on $I$, and their decomposition into the bias bound $\eta$ and the remaining length due to the stochastic part. Table (ref) summarizes results on the lengths. There are a couple of features in these results that deserve discussions. First, observe that the confidence bands shrink as sample size increases when the sieve dimension $q$ is held fixed. On the other hand, the bias bound $\eta$ remains invariant across sample sizes while the sieve dimension $q$ is fixed. These results are natural features of our approach. Second, observe that the bias bound $\eta$ decreases as the sieve dimension $q$ increments.
Finally, we display instances of confidence bands in Figure (ref). The gray shades indicate the confidence bands including the bias bound and the stochastic parts together. The internal dark gray shades include only the stochastic parts. We also plot the true density functions and Li-Vuong estimates (with the choice of tuning parameter according to Appendix (ref)) as solid and dashed curves, respectively. While such instances of confidence bands will not tell us any evidence on the statistical properties, they at least inform how a confidence band may look in applications.
In this paper, we propose a novel inference method based on Kotlarski's identity. Under the absence of an inference method prior to our present paper, some existing applied papers that use Kotlarski's identity conduct the nonparametric bootstrap of the Li-Vuong estimator to draw confidence intervals though there is no theoretical guarantee for such a method to work. In the current subsection, we use simulation studies to assess the coverage performance for such a bootstrap approach. For the purpose of comparison, we use the same data generating processes as in the previous subsection, namely Model 1, Model 2, and Model 3.
Table (ref) summarizes the simulated frequencies that the na\"ive bootstrap confidence interval covers the true probability density function at locations $x \in \mathbb{E}[X]+\{-2,1,0,1,2\}\sqrt{\mathrm{Var}(X)}$. The tuning parameter is selected based on a commonly used approach in this literature -- see Appendix (ref). The coverage frequencies are close to the nominal probabilities only near the tails of the distributions, e.g., $x = \mathbb{E}[X] \pm 2\sqrt{\mathrm{Var}(X)}$, under Model 1 and Model 2. On the other hand, the na\"ive bootstrap suffers from under-coverage near the center of the distributions under Model 1 and Model 2. Furthermore, the na\"ive bootstrap suffers from even more severe under-coverage both near the center of the distribution and near the tails of the distribution under Model 3. These results suggest that one should substitute our proposed method for the traditional na\"ive bootstrap approach.
Since its introduction to econometrics by \citet*{li/vuong:1998}, Kotlarski's identity \citep*{kotlarski1967} -- see also rao:1992 -- has been widely used in empirical economics. Examples include applications to empirical auctions LiPerrigneVuong2000,krasnokutskaya:2011, income dynamics bonhomme/robin:2010, and labor economics cunha/heckman/navarro:2005,cunha/heckman/schennach:2010,bonhomme/sauder:2011,kennan/walker:2011. Despite its popular use in applications, a method of inference based on Kotlarski's identity has long been missing in the literature. After twenty years since \citet*{li/vuong:1998}, we now propose a method of inference based on Kotlarski's identity. Specifically, we develop confidence bands for the probability density function $f_X$ of $X$ in the repeated measurement model where two measurements $(Y_1,Y_2)$ of unobserved variable $X$ are available in data with additive independent errors, $U_1 = Y_1-X$ and $U_2 = Y_2-X$.
Our construction of confidence bands can be summarized as follows. First, we derive linear complex-valued moment restrictions based on Kotlarski's identity. Second, we let the Hermite polynomial sieve approximate unknown probability density functions. Third, for a given sieve dimension and for a given class for probability density functions, we compute a bias bound for the linear complex-valued moment restrictions, and slack the linear complex-valued moment restrictions by this bias bound. Fourth, we compute the uniform norm of the self-normalized process of the slacked linear complex-valued moment restrictions as the test statistics for each point in a set of sieve coefficients. Fifth, inverting this test statistic yields a confidence set of sieve approximations to possible probability density functions. Sixth, for a given sieve dimension and for a given class for probability density functions, we compute a bias bound for sieve approximations of probability density functions, and the desired confidence band is obtained by uniformly enlarging the set of sieve approximations by this bias bound.
We not only provide a method that works, but also care for its practicality. The Fourier transform and the inverse Fourier transform operations are known to be computationally costly in the deconvolution literature. By exploiting the property of the Hermite functions as eigen-functions of the Fourier transform operator, we propose to let the Hermite polynomial sieve approximate both the density and characteristic functions without having to implement numerical integrations within each iteration of a numerical optimization routine. This convenient feature of the proposed method saves computational resources. Furthermore, we also exploit a couple of other convenient properties of the Hermite functions (namely the Schr\"odinger equation for a harmonic oscillator and a pair of recursive equations), and consequently obtain informative bias bounds and thus informative inference. With these practical features of our method, simulation studies indeed conclude reasonably fast with informative inference results. The results evidence the efficacy of the proposed method. Since Kotlarski's identity is one of the most popular methods in a number of applied fields, including empirical auctions, income dynamics, and labor economics, we hope that our method will contribute to the practice of economic analyses in these and other topics.