EconBase
← Back to paper

A Unified Nonparametric Test of Transformations on Distribution Functions with Nuisance Parameters

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

84,535 characters · 12 sections · 101 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Unified Inference on Moment Restrictions with Nuisance Parameters

abstractThis paper proposes a simple unified inference approach on moment restrictions in the presence of nuisance parameters. The proposed test is constructed based on a new characterization that avoids the estimation of nuisance parameters and can be broadly applied across diverse settings. Under suitable conditions, the test is shown to be asymptotically size controlled and consistent for both independent and dependent samples. Monte Carlo simulations show that the test performs well in finite samples. Numerical results from the application to conditional moment restriction models with weak instruments demonstrate that the proposed method may improve upon existing approaches in the literature.

Keywords: Unified inference, moment restrictions, nuisance parameters, numerical delta method, weak instruments \thispagestyle{fancy}

\setlength\abovedisplayskip{2pt plus 0pt minus 2pt} \setlength\belowdisplayskip{2pt plus 0pt minus 2pt}

Introduction

The analysis of moment restriction models plays a central role in econometric theory and applications. Considerable efforts have been devoted to estimating unknown key parameters and to testing hypotheses related to these parameters in the moment restrictions; see, e.g., chamberlain1987asymptotic, NEWEY1993419, dominguez2004consistent, kitamura2004empirical, smith2007efficient, and lavergne2013smooth, among many others; see also kunitomo2011moment for an overview of the moment restriction-based econometric methods.

Valid statistical inference on these parameters relies crucially on the correct specification of the postulated moment restriction models. Assessing the suitability of the moment restrictions has therefore generated an extensive literature; see, e.g., bierens1982consistent, tauchen1985diagnostic, newey1985generalized, and donald2003empirical. In testing the moment restrictions, the unknown parameters may not be of primary interest under the null hypothesis and can be regarded as nuisance parameters. Handling nuisance parameters in the considered testing procedures is an important theoretical issue. Existing specification tests for moment restrictions typically employ procedures that first estimate the nuisance parameters and then test the moment restrictions using the estimators; see, e.g., tripathi2003testing, delgado2006consistent, and muandet2020kernel. As a result, classical approaches are generally model- or estimator-dependent, requiring different theories and implementation procedures for different cases. In addition, these approaches may encounter theoretical difficulties due to the estimation of nuisance parameters. For example, obtaining reliable estimates of nuisance parameters may be nontrivial in conditional moment restriction models when instruments are weak.

In this paper, we propose a unified testing framework for moment restrictions with nuisance parameters that is broadly applicable to various settings. The critical values of our test are constructed using the numerical delta methods developed by hong2018numerical and chen2019inference who provide novel methodologies for addressing nonstandard testing issues.\footnote{More discussions on this topic can be found in dumbgen1993nondifferentiable, andrews2000inconsistency, hirano2012impossibility, hansen2017regression, and fang2019inference. Other discussions and applications of related bootstrap methods can be found in beare2015nonparametric, Beare2016global, Seo2016tests, Beare2015improved, chen2019improved, hong2020numerical, Beare2017improved, and sun2018ivvalidity.} The proposed method in the paper effectively circumvents the estimation of nuisance parameters, thus providing a general and robust inferential tool for different settings where nuisance parameters are present. A comparison between the proposed test and existing approaches in conditional moment restriction models with weak instruments demonstrates that the test can achieve performance improvement.

We summarize the main features of the proposed test as follows: (i) It is case-independent; (ii) it is free of the estimation of nuisance parameters, and is particularly appealing in cases where desirable estimation is challenging; (iii) it is asymptotically size controlled and consistent against a broad class of alternatives to the null; (iv) it works for both independent and dependent samples; (v) the bootstrap test procedure is simple.

Now we introduce our testing framework. Let $d_{\theta}\in\mathbb{Z}_+$ and $d_z\in\mathbb{Z}_+$. Let $\Theta\subset\mathbb{R}^{d_\theta}$ be a parameter space. Let $\Psi=\{\psi_{x,\theta}:\mathbb{R}^{d_z}\to \mathbb{R}:x\in\mathbb{R}, \theta\in\Theta\}$ be a class of functions indexed by $(x,\theta)\in\mathbb{R}\times\Theta$ such that $\psi_{x,\theta}$ is measurable for all $(x,\theta)\in\mathbb{R}\times\Theta$. Throughout the paper, all random elements are defined on a probability space $\left( \Omega, \mathscr{F}, \mathbb{P} \right)$. Let $P$ be an unknown probability distribution on $(\mathbb{R}^{d_z}, \mathscr{B}(\mathbb{R}^{d_z}))$ and $Z\sim P$ be a random vector such that for every Borel set $B\subset \mathbb{R}^{d_z}$, $P(B)=\mathbb{P}(Z\in B)$. We are interested in the null hypothesis

align[align omitted — 163 chars of source]

This can be viewed as a specification test on a set of moment restrictions. The parameter $\theta$ in (ref) is the nuisance parameter we need to take into account. Let $\phi_P:\mathbb{R}\times\Theta\to\mathbb{R}$ be a function depending on $P$ such that $\phi_P(x,\theta)=P(\psi_{x,\theta})=\mathbb{E}_P[\psi_{x,\theta}(Z)]$ for every $(x,\theta)\in\mathbb{R}\times\Theta$. Clearly, the null hypothesis in (ref) is equivalent to

align[align omitted — 145 chars of source]

The above formulation can easily be extended to cases where $x\in\mathbb{R}^k$ for some $k>1$. To simplify exposition, we present the results for scalar $x$ in the main text.

The testing approach provided in the paper can be readily applied in a wide range of empirical studies. In the following, we present several important examples where the hypothesis of interest can be formulated into (ref).

Examples

exam[label=CMR](Conditional Moment Restrictions) Let $Z=(X,Y)$ be a $d_z$-dimensional random vector with scalar $X$ and $d_y$-dimensional vector $Y$, where $d_z=d_y+1\ge 2$. Let $g:\mathbb{R}^{d_y}\times \Theta\to\mathbb{R}$ be a known function. The null hypothesis of interest is \begin{align*} \mathrm{H}_0: For some \theta\in \Theta, \mathbb{E}_P[g(Y,\theta)|X]=0 almost surely. \end{align*} This null hypothesis is equivalent to \begin{align*} \mathrm{H}_0: For some \theta\in \Theta, \mathbb{E}_P [g(Y,\theta)\mathbbm{1}\{X\le x\}]=0 for all x\in \mathbb{R}. \end{align*} In this case, $\psi_{x,\theta}(z)=g(y,\theta)\mathbbm{1}\{w\le x\}$ for every $z=(w,y)\in\mathbb{R}\times\mathbb{R}^{d_y}$ and every $(x,\theta)\in\mathbb{R}\times\Theta$, and $\phi_P(x,\theta)=\mathbb{E}_P [g(Y,\theta)\mathbbm{1}\{X\le x\}]$ for every $(x,\theta)\in\mathbb{R}\times\Theta$. tripathi2003testing construct a smoothed empirical likelihood-based test for the conditional moment restrictions, escanciano2014specification use a projected empirical process to eliminate the estimation effect of nuisance parameters, dominguez2015simple introduce an omnibus test statistic as the minimized value of the objective function considered in dominguez2004consistent, and berger2022testing proposes a new empirical likelihood test for parameters of conditional moment restriction models. jun2009semiparametric propose semi-parametric tests of conditional moment restrictions with weak instruments. The null rejection probabilities of their tests are shown to be asymptotically no greater than the nominal significance level, suggesting possible conservativeness. Under suitable conditions, the test proposed in this paper has an exact asymptotic size, which allows for dependent data as well. The performance improvement of our method over existing approaches is illustrated through simulation studies in Section (ref), where the data generating processes (DGPs) are tailored to conditional moment restriction models with weak instruments.
exam[label=symmetry](Symmetry) Let $G$ be the cumulative distribution function of the random variable $Z$. The null hypothesis of symmetry about center $\theta$ is \begin{align*} \mathrm{H}_0: For some \theta\in \Theta, G(x)=1-G(2\theta-x) for all x\in \mathbb{R}. \end{align*} In this case, $\psi_{x,\theta}(z)=\mathbbm{1}\{z\le x\}+\mathbbm{1}\{z\le 2\theta-x\}-1$ for every $z\in\mathbb{R}$ and every $(x,\theta)\in\mathbb{R}\times\Theta$, and $\phi_P(x,\theta)=G(x)+G(2\theta-x)-1$ for every $(x,\theta)\in\mathbb{R}\times\Theta$. psaradakis2015quantile use a quantile-based measure of skewness to test symmetry about an unspecified center, psaradakis2016using considers the autoregressive sieve bootstrap to obtain critical values for tests of symmetry, and psaradakis2022using employ a U-statistic involving triples of observations to assess symmetry. psaradakis2019bootstrap provide an overview of symmetry tests.
exam[label=fit](Goodness of Fit) Let $G$ be the cumulative distribution function of the random variable $Z$. Suppose there is a given class of distribution functions $\{G_0 (\cdot, \theta):\theta\in\Theta\}$ so that $x\mapsto G_0(x,\theta)$ is a distribution function on $\mathbb{R}$ for every $\theta\in\Theta$. We assume the identifiability of $\theta$ in the sense that for all $\theta_1,\theta_2\in\Theta$ with $\theta_1\ne\theta_2$, there exists $x_0\in\mathbb{R}$ such that $G_0(x_0,\theta_1)\ne G_0(x_0,\theta_2)$. The null hypothesis of correct specification is \begin{align*} \mathrm{H}_0: For some \theta\in \Theta, G(x)=G_0(x,\theta) for all x\in \mathbb{R}. \end{align*} In this case, $\psi_{x,\theta}(z)=\mathbbm{1}\{z\le x\}-G_0(x,\theta)$ for every $z\in\mathbb{R}$ and every $(x,\theta)\in\mathbb{R}\times\Theta$, and $\phi_P(x,\theta)=G(x)-G_0(x,\theta)$. Goodness-of-fit tests based on parametric empirical processes have been extensively studied since durbin1973weak. For example, the martingale approach proposed by khmaladze1982martingale is applied to the problem of testing goodness of fit with estimated parameters, and genest2008validity consider goodness-of-fit tests using a parametric bootstrap approach. A more recent work is parker2013comparison, which recommends conducting sup-norm inference for tests based on durbin1985first’s approximations.
exam[label=LST](Location-scale Transformation) We wish to test the null hypothesis of equal distributions up to some location-scale transformation. This is a generalization of the classical two-sample problem. Let $Z=(X,Y)$ be a two-dimensional random vector and $H$ be the joint cumulative distribution function of $Z$ with marginal distribution functions $F$ (for $X$) and $G$ (for $Y$). The null hypothesis is \begin{align} \mathrm{H}_0: For some \theta=(\theta_1,\theta_2)\in \Theta, F(x)= G\left(\frac{x-\theta_1}{\theta_2}\right) for all x\in \mathbb{R}. \end{align} In this case, $\psi_{x,\theta}(z)=\mathbbm{1}\{z_1\le x\}-\mathbbm{1}\{z_2\le(x-\theta_1)/\theta_2\}$ for every $z=(z_1,z_2)\in\mathbb{R}^2$ and every $(x,\theta)\in\mathbb{R}\times\Theta$, and $\phi_P(x,\theta)=F(x)-G((x-\theta_1)/\theta_2)$. A substantial number of tests exist for comparing two or multiple distributions. See, for example, lehmann2005testing and chen2018modern for extensive reviews. hall2013new propose an extension of the Cram{\'e}r--von Mises type test based on empirical characteristic functions to examine whether the two samples come from the same location-scale family of distributions. henze2005checking and jimenez2017fast deal with the two-sample problem using similar test statistics. An important special case of Example (ref) is testing for heterogeneous treatment effects. We follow ding2016randomization and chung2021permutation and consider a randomized experiment model. Let $Y$ denote the observable outcome of interest, and $D$ denote the binary treatment variable. If an individual is randomly assigned to the treatment group and receives treatment, then $D=1$; otherwise, the individual is randomly assigned to the control group and does not receive treatment, with $D=0$. Suppose that $Y(1)$ is the potential outcome of an individual if treated, and $Y(0)$ is the potential outcome if not treated. The treatment effect is constant if $Y(1)-Y(0)=\theta$ almost surely for some fixed constant $\theta$; otherwise, the treatment effect is said to be heterogeneous. The null hypothesis of constant treatment effect is \begin{align} \mathrm{H}_0^{s}: For some \theta\in \Theta, Y(1)-Y(0)=\theta almost surely. \end{align} Hypothesis (ref) is a more restrictive sharp null and is usually not directly testable. A necessary and weaker condition of this sharp null hypothesis, which is considered by ding2016randomization and chung2021permutation, is \begin{align*} \mathrm{H}_0: For some \theta\in \Theta, F(x)=G(x-\theta) for all x\in\mathbb{R}, \end{align*} where $F$ and $G$ are the CDFs of $Y(1)$ and $Y(0)$, respectively. Clearly, this condition can be incorporated into (ref).

{Organization of the Paper}: Section (ref) provides the framework and develops theoretical results for testing general moment restrictions in the presence of nuisance parameters. Section (ref) extends the results to dependent data. Section (ref) provides Monte Carlo simulation evidence to show the performance of the test in finite samples. Section (ref) concludes the paper. Auxiliary lemmas, analyses and extensions of examples, all mathematical proofs, and additional simulation results are collected in the Online Supplementary Appendix.

{Notation:} We introduce some notation following the convention van1996weak,kosorok2008introduction. We use $M^\mathsf{T}$ to denote the transpose of a matrix $M$. For $a,b\in\mathbb{R}$, we define $a\wedge b=\min\{a,b\}$ and $a\vee b=\max\{a,b\}$. We use two forms of indicator functions: $\mathbbm{1}\{S\}=1$ if the statement $S$ is true, and $\mathbbm{1}\{S\}=0$ otherwise; $\mathbbm{1}_A(x)=1$ if $x\in A$, and $\mathbbm{1}_A(x)=0$ if $x\notin A$. For an arbitrary set $A$, let $\ell^\infty(A)$ be the set of bounded real-valued functions on $A$. Equip $\ell^\infty(A)$ with the supremum norm $\left\Vert \cdot \right\Vert_{\infty}$ such that $\left\Vert f \right\Vert_{\infty}=\sup_{x\in A} \left\vert f(x) \right\vert$ for every $f\in \ell^\infty(A)$. For a subset $B$ of a metric space, let $\mathcal{C}(B)$ be the set of continuous real-valued functions on $B$, and $\mathcal{C}_\mathrm{b}(B)$ be the set of bounded continuous functions on $B$, that is, $\mathcal{C}_\mathrm{b}(B)=\mathcal{C}(B)\cap \ell^\infty(B)$. Following the notation of van1996weak, for every normed space $\mathbb{B}$ with a norm $\left\Vert \cdot \right\Vert_\mathbb{B}$, we define

align*[align* omitted — 274 chars of source]

Let $\mathbb{F}$ be an arbitrary vector space equipped with a norm $\Vert\cdot\Vert_{\mathbb{F}}$. For every $C\subset \mathbb{F}$ and every $\varepsilon>0$, define the $\varepsilon$-neighborhood of $C$ to be

align*[align* omitted — 132 chars of source]

For every measure $\nu$ on $\left( \mathbb{R}, \mathscr{B} (\mathbb{R}) \right)$, let $L^p(\nu)$ be the set of functions such that

align*[align* omitted — 137 chars of source]

with $p\ge 1$. Equip $L^p(\nu)$ with the norm $\left\Vert \cdot \right\Vert_{L^p(\nu)}$ such that

align*[align* omitted — 129 chars of source]

for every $f\in L^p(\nu)$.

Let $\mu$ be the Lebesgue measure on $\left( \mathbb{R}, \mathscr{B}(\mathbb{R}) \right)$, where $\mathscr{B}(\mathbb{R})$ denotes the collection of Borel sets in $\mathbb{R}$. For an arbitrary space $\mathcal{F}$, we say $\mathbb{W}$ is a $P$-Brownian bridge in $\ell^\infty(\mathcal{F})$ if and only if $\mathbb{W}$ is a tight Borel measurable Gaussian process with $\mathbb{E}_P[\mathbb{W}(f_1)]=0$ and $\mathbb{E}_P[\mathbb{W}(f_1)\mathbb{W}(f_2)]=P(f_1 f_2)-P(f_1)P(f_2)$ for all $f_1,f_2\in\mathcal{F}$. Let $\leadsto$ denote the weak convergence defined in van1996weak. Let $\overset{\mathbb{P}}{\leadsto}$ and $\overset{\text{a.s.}}{\leadsto}$ denote the weak convergence in probability conditional on the sample and almost sure weak convergence conditional on the sample, respectively, as defined in kosorok2008introduction.

Test Formulation

Setup

Let $\nu$ be a probability measure on $\left( \mathbb{R}, \mathscr{B}(\mathbb{R}) \right)$. We first introduce the following assumptions.

assumFor every $\theta\in\Theta$, the function $x\mapsto \phi_P(x,\theta)$ is continuous.
assumThe probability measure $\nu$ on $\left( \mathbb{R}, \mathscr{B}(\mathbb{R}) \right)$ satisfies $\mu \ll \nu$, that is, if $\nu(B)=0$ for some $B\in\mathscr{B}(\mathbb{R})$, then $\mu(B)=0$.
assumThe set $\Theta$ is compact in $\mathbb{R}^{d_\theta}$.
assumFor every $\theta_0\in\Theta$ and every $\varepsilon>0$, there exists $\delta>0$ such that \begin{align*} \sup_{x\in\mathbb{R}} P\left[ (\psi_{x,\theta}-\psi_{x,\theta_0})^2 \right] <\varepsilon \end{align*} for all $\theta\in\Theta$ with $\left\Vert \theta-\theta_0 \right\Vert_2 <\delta$.

Assumption (ref) shows that we focus on moment restrictions that are continuous in $x$ for every $\theta\in\Theta$. Assumption (ref) requires the absolute continuity of the Lebesgue measure $\mu$ with respect to the probability measure $\nu$. For example, $\nu$ could be set as the probability measure corresponding to a normal distribution with a large variance.\footnote{See the discussion and simulation results in Section (ref).} Assumption (ref) is a common condition on the compactness of $\Theta$. Assumption (ref) can be understood as the continuity of $\psi_{x,\theta}$ with respect to $\theta$ under a certain metric.

Define a function space

align*[align* omitted — 219 chars of source]

In the definition of $\mathbb{D}_{\mathcal{L}0}$, the continuity of the map $\theta\mapsto \varphi(\cdot,\theta)$ is understood in the sense that for every $\theta_0\in \Theta$ and every $\varepsilon>0$, there exists $\delta>0$ such that

align*[align* omitted — 130 chars of source]

for all $\theta\in\Theta$ with $\left\Vert \theta-\theta_0 \right\Vert_2 <\delta$. Note that for every $x\in\mathbb{R}$ and all $\theta,\theta_0\in\Theta$, by Jensen's inequality,

align*[align* omitted — 114 chars of source]

Since $\nu$ is a probability measure, Assumption (ref) implies that $\phi_P\in\mathbb{D}_{\mathcal{L} 0}$.

The proposition below provides an equivalent characterization of the null hypothesis in (ref). We construct the test based on this equivalent characterization to avoid estimating the nuisance parameter $\theta$ under the null.

propIf Assumptions (ref)--(ref) hold, then the null hypothesis in (ref) is equivalent to \begin{align} \mathrm{H}_0: \inf_{\theta\in\Theta}\int_{\mathbb{R}} \left[\phi_P(x,\theta) \right]^2 \;\mathrm{d} \nu(x)=0. \end{align}

It is worth noting that different measures $\nu$ may deliver different power properties of the test. However, searching for the optimal $\nu$ to maximize power is challenging, as it may depend in a complicated manner on the DGP.

The measure $\nu(\mathbb{R})$ is assumed to be finite (Assumption (ref)) to obtain the theoretical results in the paper. In practice, we suggest setting $\nu$ to a normal probability measure with a large variance so that it does not heavily concentrate on some region of the real line, given no prior information about the DGP of the data. Other probability measures satisfying Assumption (ref) also work asymptotically for the proposed method. For finite samples, the simulation results in Section (ref) show that normal probability measures with different variances ($\mathcal{N}(0,1)$, $\mathcal{N}(0,5^2)$, $\mathcal{N}(0,10^2)$) perform well.

Test Statistic

We first restrict our attention to independent and identically distributed (i.i.d.)\ samples, and will extend the results to dependent data in Section (ref). Let $\widehat{P}_n$ be the empirical probability measure of the sample $\mathbf{Z}_n$, which assigns weight $1/n$ to each observation $Z_i$ with $i\in\{1,\ldots,n\}$. Then the sample analogue of $\phi_P$ is defined as

align*[align* omitted — 116 chars of source]

for every $(x,\theta)\in\mathbb{R}\times\Theta$. We present the exact function form of $\widehat{\phi}_n$ in every example.

exam[continues=CMR] With the known function $g$, it follows by definition that \begin{align*} \widehat{\phi}_n(x,\theta)=\widehat{P}_n(\psi_{x,\theta})=\frac{1}{n}\sum_{i=1}^n \psi_{x,\theta}(Z_i)=\frac{1}{n}\sum_{i=1}^n g(Y_i,\theta)\mathbbm{1}\{X_i\le x\} \end{align*} for every $(x,\theta)\in\mathbb{R}\times\Theta$.
exam[continues=symmetry] The cumulative distribution function $G$ can be estimated by the empirical distribution function $\widehat{G}_{n}$ such that for every $x\in \mathbb{R}$, \begin{align*} \widehat{G}_{n}(x)=\widehat{P}_n(\mathbbm{1}_{(-\infty,x]})=\frac{1}{n}\sum_{i=1}^{n} \mathbbm{1}_{(-\infty,x]}\left( Z_i \right) . \end{align*} Then \begin{align*} \widehat{\phi}_n(x,\theta)=\widehat{P}_n(\psi_{x,\theta})=\widehat{P}_n(\mathbbm{1}_{(-\infty,x]})+\widehat{P}_n(\mathbbm{1}_{(-\infty,2\theta-x]})-1= \widehat{G}_{n}(x)+\widehat{G}_{n}(2\theta-x)-1 \end{align*} for every $(x,\theta)\in\mathbb{R}\times\Theta$.
exam[continues=fit] The cumulative distribution function $G$ can be estimated by the empirical distribution function $\widehat{G}_{n}$ such that for every $x\in \mathbb{R}$, \begin{align*} \widehat{G}_{n}(x)=\widehat{P}_n(\mathbbm{1}_{(-\infty,x]})=\frac{1}{n}\sum_{i=1}^{n} \mathbbm{1}_{(-\infty,x]}\left( Z_i \right) . \end{align*} Then $\widehat{\phi}_n(x,\theta)=\widehat{P}_n(\psi_{x,\theta})=\widehat{P}_n(\mathbbm{1}_{(-\infty,x]})-G_0(x,\theta)=\widehat{G}_{n}(x)-G_0(x,\theta)$ for every $(x,\theta)\in\mathbb{R}\times\Theta$.
exam[continues=LST] For every $i\in\{1,\ldots, n\}$, the observation $Z_i=(X_i,Y_i)$. Let $\widehat{P}_n$ be the empirical distribution of $\{Z_i\}_{i=1}^n$, and $\widehat{H}_{n}$ be its empirical distribution function so that \begin{align*} \widehat{H}_{n}(x,y)=\frac{1}{n}\sum_{i=1}^{n} \mathbbm{1}_{(-\infty,x]\times(-\infty,y]}\left( X_i,Y_i \right) \end{align*} for all $(x,y)\in\mathbb{R}^2$. Let $\widehat{P}_{X,n}$ and $\widehat{P}_{Y,n}$ be the marginal distributions of $\widehat{P}_n$, i.e., the (marginal) empirical distributions of $\left\{ X_i \right\}_{i=1}^{n}$ and $\left\{ Y_i \right\}_{i=1}^{n}$, respectively. It follows that \begin{align*} \widehat{\phi}_n(x,\theta)=\widehat{P}_n(\psi_{x,\theta})=\widehat{P}_{X,n}(\mathbbm{1}_{(-\infty,x]})-\widehat{P}_{Y,n}(\mathbbm{1}_{(-\infty,(x-\theta_1)/\theta_2]}) \end{align*} for every $(x,\theta)\in\mathbb{R}\times\Theta$. The marginal distribution functions $F$ and $G$ can be estimated by the empirical distribution functions $\widehat{F}_n$ and $\widehat{G}_n$, respectively, where for every $x\in\mathbb{R}$, \begin{align*} &\widehat{F}_{n}(x)=\widehat{P}_{X,n}(\mathbbm{1}_{(-\infty,x]})=\frac{1}{n}\sum_{i=1}^{n} \mathbbm{1}_{(-\infty,x]}\left( X_i \right) and \\ &\widehat{G}_{n}(x)=\widehat{P}_{Y,n}(\mathbbm{1}_{(-\infty,x]})=\frac{1}{n}\sum_{i=1}^{n} \mathbbm{1}_{(-\infty,x]}\left( Y_i \right). \end{align*} This implies that $\widehat{\phi}_n(x,\theta)=\widehat{F}_{n} (x)-\widehat{G}_{n} [ (x-\theta_1)/\theta_2 ]$ for every $(x,\theta)\in\mathbb{R}\times\Theta$. We may also test this null hypothesis with two independent samples of different sizes. This case will be discussed in Appendix (ref), where we present the results for comparing multiple samples.

To obtain the asymptotic law of the stochastic process $\widehat{\phi}_n$, we need the following assumption on the function class $\Psi$.

assumThe function class $\Psi=\{\psi_{x,\theta}:(x,\theta)\in\mathbb{R}\times\Theta\}$ satisfies that \begin{align} \sup_{f\in\Psi} \left\vert f(z)-Pf \right\vert <\infty \end{align} for all $z\in\mathbb{R}^{d_z}$, and is $P$-Donsker in the sense that \begin{align} \sqrt{n}(\widehat{P}_n-P)\rightsquigarrow \mathbb{W} in \ell^\infty(\Psi) \end{align} as $n\to\infty$, where $\mathbb{W}$ is a $P$-Brownian bridge in $\ell^\infty(\Psi)$.

Lemma (ref) establishes the consistency of $\widehat{\phi}_n$ and the weak convergence of $\sqrt{n}( \widehat{\phi}_n-\phi_P )$ in $\ell^\infty(\mathbb{R}\times\Theta)$ as $n\to\infty$.

lemmaIf Assumptions (ref) and (ref) hold, then $(\widehat{\phi}_n-\phi_P)\in\ell^\infty(\mathbb{R}\times\Theta)$ for all $n\in\mathbb{Z}_+$. In addition, \begin{align*} \sup_{(x,\theta)\in\mathbb{R}\times\Theta} \left\vert \widehat{\phi}_n(x,\theta)-\phi_P(x,\theta) \right\vert \xrightarrow{\mathbb{P}} 0 and \sqrt{n}( \widehat{\phi}_n-\phi_P )\rightsquigarrow \mathbb{G}_0 in \ell^\infty(\mathbb{R}\times\Theta) \end{align*} as $n\to\infty$, where $\mathbb{G}_{0}$ is some tight random element which almost surely takes values in $\mathbb{D}_{\mathcal{L}0}$.

Define a function space

align*[align* omitted — 220 chars of source]

Define a map $\mathcal{L}$ on $\mathbb{D}_{\mathcal{L}}$ such that $\mathcal{L}(\varphi)= \inf_{\theta\in\Theta} \int_{\mathbb{R}} \left[ \varphi (x,\theta) \right]^2 \;\mathrm{d} \nu(x)$ for every $\varphi\in \mathbb{D}_{\mathcal{L}}$. Then under Assumptions (ref)--(ref), the null and the alternative hypotheses can be expressed as

align[align omitted — 122 chars of source]

To test the null hypothesis in (ref), we set the test statistic to $n \mathcal{L} ( \widehat{\phi}_n )$.

Next, we show that the map $\mathcal{L}$ is Hadamard directionally differentiable, but its Hadamard directional derivative is degenerate under $\mathrm{H}_0$.\footnote{See Definition (ref) for Hadamard directional differentiability.} Define $$\mathbb{D}_0=\left\{ \varphi\in \mathbb{D}_{\mathcal{L}0}:\mathcal{L}(\varphi)=0 \right\}.$$ The following lemma provides the Hadamard directional derivative of $\mathcal{L}$ and its first order degeneracy under $\mathrm{H}_0$.

lemmaIf Assumptions (ref) and (ref) hold, then $\mathcal{L}$ is Hadamard directionally differentiable at $\phi_P\in \mathbb{D}_{\mathcal{L}}$ tangentially to $\mathbb{D}_{\mathcal{L}0}$ with the Hadamard directional derivative \begin{align*} \mathcal{L}'_{\phi_P}(h)=2\inf_{\theta\in \Theta_0(\phi_P)} \int_{\mathbb{R}} \phi_P(x,\theta)h(x,\theta) \;\mathrm{d} \nu(x) for all h\in \mathbb{D}_{\mathcal{L}0}, \end{align*} where $\Theta_0(\phi_P)=\operatornamewithlimits{arg\,min}_{\theta\in\Theta}\int_{\mathbb{R}} \left[ \phi_P (x,\theta) \right]^2 \;\mathrm{d} \nu(x)$. Moreover, if $\phi_P\in\mathbb{D}_0$, then the derivative $\mathcal{L}'_{\phi_P}$ is well defined on the whole of $\ell^\infty(\mathbb{R}\times\Theta)$ with $\mathcal{L}'_{\phi_P}(h)=0$ for every $h\in\ell^\infty(\mathbb{R}\times\Theta)$.

The first order degeneracy of $\mathcal{L}$ under $\mathrm{H}_0$ implies that we may need to find the second order Hadamard directional derivative of $\mathcal{L}$.\footnote{See Definition (ref) for second order Hadamard directional differentiability.} We assume the following conditions to guarantee the existence of the second order Hadamard directional derivative of $\mathcal{L}$.

assumThe function $\phi_P$ is twice differentiable with respect to $\theta$, and the second partial derivative satisfies \begin{align} \int_{\mathbb{R}} \sup_{\theta\in\Theta} \left\Vert \left. \frac{\partial^2 \phi_P (z,\vartheta)}{\partial \vartheta \partial \vartheta^\mathsf{T}} \right\vert_{(z,\vartheta)=(x,\theta)} \right\Vert_{2}^2 \,\mathrm{d}\nu(x) <\infty, \end{align} where $\left\Vert \cdot \right\Vert_{2}$ denotes the $\ell^2$ operator norm of a matrix.
assumThe set $\Theta_0\equiv\{\theta\in\Theta:\int_{\mathbb{R}} \left[ \phi_P (x,\theta) \right]^2 \;\mathrm{d} \nu(x)=0\}\subset\mathrm{int}(\Theta)$, and there exist $\kappa\in(0,1]$, $\overline{\varepsilon}>0$, and $C>0$ such that for all $\varepsilon\in (0,\overline{\varepsilon})$, \begin{align} \inf_{\theta\in\Theta\setminus\Theta_0^\varepsilon} \left\{\int_\mathbb{R} \left[ \phi_P(x,\theta) \right]^2 \;\mathrm{d} \nu(x)\right\}^{1/2} \ge C \varepsilon^\kappa. \end{align}

We provide Assumptions (ref) and (ref) following the basic idea of chen2019inference. Assumption (ref) requires the boundedness of the second partial derivative of $\phi_P$ in the sense of (ref). Assumption (ref) requires that the set $\Theta_0$ is in the interior of $\Theta$ and it is well separated. The condition in (ref) is similar to the partial identification assumption used in chernozhukov2007estimation. It is worth noting that these conditions are sufficient but not necessary for our results, as also mentioned by chen2019inference. We impose such high level conditions for theoretical completeness. In Section (ref), we verify these assumptions for a conditional moment restriction model.

lemmaIf Assumptions (ref), (ref), (ref), and (ref) hold, and $\phi_P\in\mathbb{D}_0$, then the function $\mathcal{L}$ is second order Hadamard directionally differentiable at $\phi_P$ tangentially to $\mathbb{D}_{\mathcal{L}0}$ with the second order Hadamard directional derivative \begin{align*} \mathcal{L}”_{\phi_P}(h)=\inf_{\theta\in\Theta_0(\phi_P)} \inf_{v\in \mathbb{R}^{d_{\theta}}} \left\Vert \left[ \Phi'(\theta) \right]^\mathsf{T} v +\mathscr{H} (\theta) \right\Vert_{L^2(\nu)}^2 for all h\in \mathbb{D}_{\mathcal{L}0}, \end{align*} where $\Phi'(\theta): \mathbb{R}\to\mathbb{R}^{d_{\theta}}$ with \begin{align*} \Phi'(\theta)(x)=\left. \frac{\partial \phi_P (z,\vartheta)}{\partial \vartheta} \right\vert_{(z,\vartheta)=(x,\theta)} \quad for every (x,\theta)\in\mathbb{R}\times\Theta, \end{align*} and $\mathscr{H}:\Theta\to \ell^\infty(\mathbb{R})$ with $\mathscr{H}(\theta)(x)=h(x,\theta)$ for every $(x,\theta)\in\mathbb{R}\times\Theta$.
remarkLemma (ref) provides the explicit expression of the complicated second order Hadamard directional derivative of $\mathcal{L}$. We employ a numerical method that does not require exploring this function form.

With Lemma (ref), the asymptotic null distribution of the test statistic $\mathcal{L}( \widehat{\phi}_n )$ is obtained by applying the second order delta method.

propIf Assumptions (ref)--(ref) hold and $\mathrm{H}_0$ is true ($\phi_P\in\mathbb{D}_0$), then \begin{align*} n \mathcal{L}(\widehat{\phi}_n ) \rightsquigarrow \mathcal{L}”_{\phi_P}\left( \mathbb{G}_0 \right) as n\to \infty. \end{align*}

Bootstrap Procedure

The distribution of $\mathcal{L}''_{\phi_P}\left( \mathbb{G}_0 \right)$ in Proposition (ref) is unknown because both the function $\mathcal{L}''_{\phi_P}$ and the stochastic process $\mathbb{G}_0$ depend on the unknown underlying distribution $P$. Motivated by hong2018numerical and chen2019inference, we propose to approximate $\mathcal{L}''_{\phi_P}$ by a consistent estimator and approximate the distribution of $\mathbb{G}_0$ by bootstrap.\footnote{Bootstrap may not be the only method to approximate the distribution of $\mathbb{G}_0$ in our framework. Other consistent estimators of $\mathbb{G}_0$ might also suffice for the proposed approach.} We use the numerical second order Hadamard directional derivative $\widehat{\mathcal{L}}''_n$ to approximate $\mathcal{L}''_{\phi_P}$, which is defined as

align*[align* omitted — 132 chars of source]

for all $h\in \ell^\infty(\mathbb{R}\times\Theta)$, where $\left\{ \tau_n \right\}$ is a sequence of tuning parameters satisfying the assumption below.\footnote{As discussed in chen2019inference, the modified bootstrap in babu1984bootstrapping (Babu correction) is inappropriate when $\mathcal{L}$ is only second order Hadamard directionally differentiable but $\mathcal{L}''_{\phi_P}$ is not “continuous” in $\phi_P$. To ensure that our method can accommodate more general cases, we employ the bootstrap method of hong2018numerical and chen2019inference.}

assum$\left\{ \tau_n \right\}\subset\mathbb{R}_+$ is a sequence of scalars such that $\tau_n\downarrow 0$ and $\tau_n \sqrt{n}\to\infty$ as $n\to\infty$.

Assumption (ref) provides the rate at which $\tau_n\downarrow0$. Under this condition, we show that $\widehat{\mathcal{L}}''_n$ approximates $\mathcal{L}''_{\phi_P}$ well in the following lemma.

lemmaIf Assumptions (ref)--(ref) hold and $\mathrm{H}_0$ is true ($\phi_P\in\mathbb{D}_0$), then for every sequence $\left\{ h_n \right\}\subset \ell^\infty (\mathbb{R}\times\Theta)$ and every $h\in\mathbb{D}_{\mathcal{L}0}$ such that $h_n\to h$ in $\ell^\infty (\mathbb{R}\times\Theta)$ as $n\to\infty$, we have \begin{align*} \widehat{\mathcal{L}}”_n\left( h_n \right) \xrightarrow{\mathbb{P}} \mathcal{L}”_{\phi_P} (h) as n\to\infty. \end{align*}

We next approximate the distribution of $\mathbb{G}_0$ via bootstrap. The bootstrap sample $\mathbf{Z}_n^*=\{Z_i^*\}_{i=1}^n$ is i.i.d.\ drawn from the empirical distribution $\widehat{P}_n$ of the original sample $\mathbf{Z}_n$. Equivalently, $\mathbf{Z}_n^*$ is a random sample of size $n$, drawn from the set $\mathbf{Z}_n$ with replacement. Let $\widehat{P}_n^*$ be the empirical distribution of $\mathbf{Z}_n^*$. The the bootstrap version of $\widehat{\phi}_n$ is $\widehat{\phi}_n^*$ such that

align*[align* omitted — 122 chars of source]

for every $(x,\theta)\in\mathbb{R}\times\Theta$.

exam[continues=CMR] It follows by definition that \begin{align*} \widehat{\phi}_n^*(x,\theta)=\widehat{P}_n^*(\psi_{x,\theta})=\frac{1}{n}\sum_{i=1}^n \psi_{x,\theta}(Z_i^*)=\frac{1}{n}\sum_{i=1}^n g(Y_i^*,\theta)\mathbbm{1}\{X_i^*\le x\} \end{align*} for every $(x,\theta)\in\mathbb{R}\times\Theta$, where $Z_i^*=(X_i^*,Y_i^*)$.
exam[continues=symmetry] Define \begin{align*} \widehat{G}_n^*(x)=\widehat{P}_n^*(\mathbbm{1}_{(-\infty,x]})=\frac{1}{n}\sum_{i=1}^{n} \mathbbm{1}_{(-\infty,x]}\left( Z^*_i \right) \end{align*} for every $x\in\mathbb{R}$. Then \begin{align*} \widehat{\phi}_n^*(x,\theta)=\widehat{P}_n^*(\psi_{x,\theta})=\widehat{P}_n^*(\mathbbm{1}_{(-\infty,x]})+\widehat{P}_n^*(\mathbbm{1}_{(-\infty,2\theta-x]})-1= \widehat{G}_n^*(x)+\widehat{G}_n^*(2\theta-x)-1 \end{align*} for every $(x,\theta)\in \mathbb{R}\times\Theta$.
exam[continues=fit] Define \begin{align*} \widehat{G}_n^*(x)=\widehat{P}_n^*(\mathbbm{1}_{(-\infty,x]})=\frac{1}{n}\sum_{i=1}^{n} \mathbbm{1}_{(-\infty,x]}\left( Z^*_i \right) \end{align*} for every $x\in\mathbb{R}$. Then \begin{align*} \widehat{\phi}_n^*(x,\theta)=\widehat{P}_n^*(\psi_{x,\theta})=\widehat{P}_n^*(\mathbbm{1}_{(-\infty,x]})-G_0(x,\theta)=\widehat{G}_n^*(x)-G_0(x,\theta) \end{align*} for every $(x,\theta)\in \mathbb{R}\times\Theta$.
exam[continues=LST] Define $\widehat{P}_n^*$ as the empirical distribution of $\{Z_i^*\}_{i=1}^n$ with $Z_i^*=(X_i^*,Y_i^*)$. Let $\widehat{P}^*_{X,n}$ and $\widehat{P}^*_{Y,n}$ be the marginal distributions of $\widehat{P}_n^*$, i.e., the (marginal) empirical distributions of $\left\{ X_i^* \right\}_{i=1}^{n}$ and $\left\{ Y_i^* \right\}_{i=1}^{n}$, respectively. It follows that \begin{align*} \widehat{\phi}_n^*(x,\theta)=\widehat{P}_n^*(\psi_{x,\theta})=\widehat{P}_{X,n}^*(\mathbbm{1}_{(-\infty,x]})-\widehat{P}_{Y,n}^*(\mathbbm{1}_{(-\infty,(x-\theta_1)/\theta_2]}) \end{align*} for every $(x,\theta)\in\mathbb{R}\times\Theta$. Define $\widehat{F}_{n}^*$ and $\widehat{G}_{n}^*$ to be the (marginal) empirical distribution functions of $\left\{ X_i^* \right\}_{i=1}^{n}$ and $\left\{ Y_i^* \right\}_{i=1}^{n}$, respectively, such that for every $x\in\mathbb{R}$, \begin{align*} &\widehat{F}_{n}^*(x)=\widehat{P}_{X,n}^*(\mathbbm{1}_{(-\infty,x]})=\frac{1}{n}\sum_{i=1}^{n} \mathbbm{1}_{(-\infty,x]}\left( X_i^* \right) and \\ &\widehat{G}_{n}^*(x)=\widehat{P}_{Y,n}^*(\mathbbm{1}_{(-\infty,x]})=\frac{1}{n}\sum_{i=1}^{n} \mathbbm{1}_{(-\infty,x]}\left( Y_i^* \right). \end{align*} This implies that $\widehat{\phi}_n^*(x,\theta)=\widehat{F}_{n}^* (x)-\widehat{G}_{n}^*\left( (x-\theta_1)/\theta_2 \right)$ for every $(x,\theta)\in\mathbb{R}\times\Theta$.

The following lemma establishes the conditional weak convergence of $\sqrt{n}(\widehat{\phi}_n^*-\widehat{\phi}_n)$ in probability as $n\to\infty$.

lemmaIf Assumption (ref) holds, then as $n\to\infty$, \begin{align*} \sup_{\Gamma\in\mathrm{BL}_1\left( \ell^\infty(\mathbb{R}\times\Theta) \right)}\left\vert \mathbb{E} \left[ \left. \Gamma \left( \sqrt{n} \left( \widehat{\phi}_n^*-\widehat{\phi}_n \right) \right) \right\vert \mathbf{Z}_n \right]-\mathbb{E} \left[ \Gamma \left( \mathbb{G}_0 \right) \right] \right\vert \xrightarrow{\mathbb{P}} 0, \end{align*} and $\sqrt{n}( \widehat{\phi}_n^*-\widehat{\phi}_n )$ is asymptotically measurable, where $\mathbb{G}_0$ is defined as in Lemma (ref).

With the numerical estimator $\widehat{\mathcal{L}}''_n$ for $\mathcal{L}''_{\phi_P}$ and a suitable bootstrap approximation $\sqrt{n}( \widehat{\phi}_n^*-\widehat{\phi}_n )$ for $\mathbb{G}_0$ at hand, we can naturally approximate the distribution of $\mathcal{L}''_{\phi_P}\left( \mathbb{G}_0 \right)$ by the conditional distribution of the bootstrap test statistic $\widehat{\mathcal{L}}''_n \{ \sqrt{n}( \widehat{\phi}_n^*-\widehat{\phi}_n ) \}$ given the original samples. This is justified by the following proposition.

propIf Assumptions (ref)--(ref) hold and $\mathrm{H}_0$ is true ($\phi_P\in\mathbb{D}_0$), then \begin{align*} \sup_{\Gamma\in\mathrm{BL}_1\left( \mathbb{R} \right)} \left\vert \mathbb{E} \left[ \left. \Gamma \left( \widehat{\mathcal{L}}”_n \left[ \sqrt{n} \left( \widehat{\phi}_n^* -\widehat{\phi}_n \right) \right] \right) \right\vert \mathbf{Z}_n \right]-\mathbb{E} \left[ \Gamma \left( \mathcal{L}”_{\phi_P} \left( \mathbb{G}_0 \right) \right) \right] \right\vert \xrightarrow{\mathbb{P}} 0 \end{align*} as $n\to\infty$.

Asymptotic Properties

Now we construct the test for the null hypothesis $\mathrm{H}_0$. For a given level of significance $\alpha\in(0,1)$, define the bootstrap critical value

align*[align* omitted — 252 chars of source]

In practice, $\widehat{c}_{1-\alpha,n}$ may be approximated by the $1-\alpha$ empirical quantile of the $n_B$ independently generated bootstrap test statistics, with $n_B$ set to be as large as computationally feasible. We reject $\mathrm{H}_0$ if and only if $n\mathcal{L} ( \widehat{\phi}_n )>\widehat{c}_{1-\alpha,n}$. The following theorem shows that the proposed test is asymptotically size controlled and consistent.

thrySuppose that Assumptions (ref)--(ref) hold. \begin{enumerate}[label=(\roman*),nosep] • If $\mathrm{H}_0$ is true and the CDF of $\mathcal{L}''_{\phi_P}\left( \mathbb{G}_0 \right)$ is strictly increasing and continuous at its $1-\alpha$ quantile, then \begin{align*} \lim_{n\to\infty} \mathbb{P}\left( n \mathcal{L}( \widehat{\phi}_n )>\widehat{c}_{1-\alpha,n} \right) = \alpha. \end{align*} • If $\mathrm{H}_0$ is false, then \begin{align*} \lim_{n\to\infty} \mathbb{P}\left( n \mathcal{L}( \widehat{\phi}_n )>\widehat{c}_{1-\alpha,n} \right)=1. \end{align*} \end{enumerate}

Local Power

In this section, we consider the local power of the test following the discussion in chen2019inference. For each $n\in\mathbb{Z}_+$, let the sample $\mathbf{Z}_n=\{Z_i\}_{i=1}^n$ be distributed according to the joint law $P_n^n=\prod_{i=1}^n P_n$, where $P_n$ is a probability distribution on $(\mathbb{R}^{d_z}, \mathscr{B}(\mathbb{R}^{d_z}))$ with $P_n(B)=\mathbb{P}(Z_i\in B)$ for every Borel set $B$. That is, for each $n\in\mathbb{Z}_+$, the observations $Z_1,\ldots,Z_n$ are i.i.d.\ with distribution $P_n$. We suppose that the null hypothesis $\mathrm{H}_0$ is false for each $P_n$, that is, for all $\theta\in\Theta$, $P_n(\psi_{x,\theta})\ne 0$ for some $x\in\mathbb{R}$. Suppose that $P_n$ converges (in a way as described in the following assumption) to some probability measure $P$, and that $P$ satisfies $\mathrm{H}_0$, that is, for some $\theta\in\Theta$, $P(\psi_{x,\theta})= 0$ for all $x\in\mathbb{R}$.

assumThe probability distributions $P_n$ and $P$ satisfy that \begin{align} \lim_{n\to\infty}\int \left[ \sqrt{n}\left( \mathrm{d}P_n^{1/2} - \mathrm{d}P^{1/2} \right) -\frac{1}{2}v_0\,\mathrm{d}P^{1/2} \right]^2=0 \end{align} for some measurable function $v_0:\mathbb{R}^{d_z}\to\mathbb{R}$, where $\mathrm{d}P_n^{1/2}$ and $\mathrm{d}P^{1/2}$ denote the square roots of the densities of $P_n$ and $P$, respectively.

Our local power results rely on Assumption (ref), which is similar to (3.10.10) of van1996weak. The following proposition states formally the local power property of the test.

propSuppose that Assumptions (ref)--(ref) hold, $\sup_{f\in\Psi} |P(f)|<\infty$, and $\sup_{f\in\Psi} |P_n(f^2)|=O(1)$. Then $\sqrt{n}( \widehat{\phi}_n-\phi_P )\leadsto \mathbb{G}_0+\zeta_P$, where $\mathbb{G}_0$ is some tight random element, and $\zeta_P(x,\theta)= P(\psi_{x,\theta}v_0)$ for every $(x,\theta)\in\mathbb{R}\times\Theta$. Furthermore, if the CDF of $\mathcal{L}''_{\phi_P}(\mathbb{G}_0)$ is strictly increasing and continuous at its $1-\alpha$ quantile $c_{1-\alpha}$, then it follows that \begin{align*} \liminf_{n\to\infty}\mathbb{P}\left( n \mathcal{L}( \widehat{\phi}_n )>\widehat{c}_{1-\alpha,n} \right) \ge \mathbb{P}(\mathcal{L}”_{\phi_P}(\mathbb{G}_0+\zeta_P)>c_{1-\alpha}). \end{align*}

Proposition (ref) follows from Lemma C.1 of chen2019inference and provides lower bounds for the power of the test under local perturbations to the null.

Dependent Data

In this section, we consider the cases where the observations $\{Z_i\}_{i=1}^n$ may be dependent. For results established in Section (ref), it is worth noting that Lemmas (ref)--(ref), Propositions (ref)--(ref), and Theorem (ref) do not directly rely on the i.i.d.\ nature of the data observations, possibly given the consistency and weak convergence of $\widehat{\phi}_n$ (Lemma (ref)) and the conditional weak convergence of $\widehat{\phi}_n^*$ in probability (Lemma (ref)). Thus, to obtain the asymptotic properties of the proposed test in dependent samples, it suffices to establish the consistency and weak convergence of $\widehat{\phi}_n$ and the conditional weak convergence of $\widehat{\phi}_n^*$ in probability under dependency.

A sequence of $d_z$-dimensional random vectors, $\{Z_i:i\in\mathbb{Z}\}$, is said to be strictly stationary, if for all $\{i_1,\ldots,i_n\}\subset\mathbb{Z}$ and all $n\in\mathbb{Z}_+$, the joint distribution of $(Z_{i_1+k},\ldots,Z_{i_n+k})$ does not depend on $k$. For $-\infty\le s\le t \le \infty$, let $\mathscr{S}_{s}^{t}$ be the $\sigma$-field generated by $\{Z_s,\ldots, Z_t\}$. Following Equation (II) of volkonskii1959some and (1.1) of arcones1994central, the $\beta$-mixing coefficient $\beta_k$ of the sequence $\{Z_i:i\in\mathbb{Z}\}$ is defined as

align*[align* omitted — 215 chars of source]

and $\{Z_i:i\in\mathbb{Z}\}$ is said to be $\beta$-mixing if and only if $\beta_k\to 0$ as $k\to\infty$.

Throughout our discussion of cases with dependent data, we assume that the sample $\mathbf{Z}_n=\{Z_i:i=1,\ldots, n\}$ is a finite segment of the strictly stationary sequence $\{Z_i:i\in\mathbb{Z}\}$ in which the common marginal distribution of $Z_i$ is $P$. We impose the following assumptions.

assumThe class $\Psi=\{\psi_{x,\theta}:(x,\theta)\in\mathbb{R}\times\Theta\}$ is a VC-subgraph class of functions satisfying (ref) with $P(\overline{\psi}^p)<\infty$ for some $p\in(2,\infty)$, where $\overline{\psi}(z)=\sup_{f\in\Psi}|f(z)|$ for every $z\in\mathbb{R}^{d_z}$, and $\Psi$ is totally bounded under $\Vert\cdot\Vert_{L^2(P)}$.\footnote{See the definition of VC-subgraph class of functions in Section 2.6 of van1996weak.}

With $p$ specified as in Assumption (ref), we introduce the following condition for $\beta_k$.

assumThe sequence $\{Z_i:i\in\mathbb{Z}\}$ is $\beta$-mixing with coefficient $\beta_k=O(k^{-q})$ as $k\to\infty$ for some $q>p/(p-2)$.

Assumption (ref) emerges as one of the conditions in Theorem 2.1 of arcones1994central and Theorem 1 of radulovic1996bootstrap. Assumption (ref) corresponds to one of the conditions in Theorem 1 of radulovic1996bootstrap.

Let $\widehat{\phi}_n$ and $\phi_P$ be defined as in Section (ref). The lemma below establishes the consistency and weak convergence of $\widehat{\phi}_n$ as $n\to\infty$.

lemmaIf Assumptions (ref), (ref), (ref), and (ref) hold, then $(\widehat{\phi}_n-\phi_P)\in\ell^\infty(\mathbb{R}\times\Theta)$ for all $n\in\mathbb{Z}_+$. In addition, \begin{align*} \sup_{(x,\theta)\in\mathbb{R}\times\Theta} \left\vert \widehat{\phi}_n(x,\theta)-\phi_P(x,\theta) \right\vert \xrightarrow{\mathbb{P}} 0 and \sqrt{n}( \widehat{\phi}_n-\phi_P )\rightsquigarrow \mathbb{G}_0 in \ell^\infty(\mathbb{R}\times\Theta) \end{align*} as $n\to\infty$, where $\mathbb{G}_{0}$ is tight and almost surely takes values in $\mathbb{D}_{\mathcal{L}0}$.

To construct the bootstrap sample $\mathbf{Z}_n^*=\{Z_i^*\}_{i=1}^n$, we follow radulovic1996bootstrap and use the moving blocks bootstrap (MBB) procedure. Recall that the original sample is $\{Z_i\}_{i=1}^n$. Let $b\in\mathbb{Z}_+$ be the block size satisfying $b\to\infty$ and $b/n\to0$, and $k\in\mathbb{Z}_+$ be the number of blocks. Without loss of generality, we may assume that $k$ and $b$ satisfy $kb=n$.\footnote{In practice, $n/b$ may not always be an integer. In this case, we set $k=\lceil n/b \rceil$ and generate $kb>n$ bootstrap observations according to the algorithm described in the main text, and then keep the first $n$ observations as the bootstrap sample.} For $i\in\{1,\ldots, b-1\}$, we set $Z_{n+i}=Z_i$. Let the random variables $I_1,\ldots, I_k$ be i.i.d.\ from $\mathrm{Unif}\{1,\ldots, n\}$ and independent of the original sample. For all $\ell\in\{1,\ldots,k\}$ and $j\in\{1,\ldots,b\}$, set the bootstrap observation $Z_{(\ell-1)b+j}^*=Z_{I_\ell+j-1}$. That is, the bootstrap sample is

align*[align* omitted — 164 chars of source]

Let $\widehat{P}_n^*$ be the empirical distribution of $\mathbf{Z}_n^*$. The bootstrap version of $\widehat{\phi}_n$ is defined as

align*[align* omitted — 122 chars of source]

for every $(x,\theta)\in\mathbb{R}\times\Theta$.

We impose the assumption below on the block size $b$, which treats $b$ as a function of $n$, that is, $b=b(n)$. This assumption corresponds to one of the conditions in Theorem 1 of radulovic1996bootstrap.

assumThe block size $b$ is a function of the sample size $n$ such that $b=b(n)=O(n^r)$ as $n\to\infty$ for some $0<r<(p-2)/(2p-2)$.

The following lemma establishes the conditional weak convergence of $\sqrt{n}(\widehat{\phi}_n^*-\widehat{\phi}_n)$ in probability.

lemmaIf Assumptions (ref)--(ref) hold, then as $n\to\infty$, \begin{align*} \sup_{\Gamma\in\mathrm{BL}_1\left( \ell^\infty(\mathbb{R}\times\Theta) \right)}\left\vert \mathbb{E} \left[ \left. \Gamma \left( \sqrt{n} \left( \widehat{\phi}_n^*-\widehat{\phi}_n \right) \right) \right\vert \mathbf{Z}_n \right]-\mathbb{E} \left[ \Gamma \left( \mathbb{G}_0 \right) \right] \right\vert \xrightarrow{\mathbb{P}} 0, \end{align*} where $\mathbb{G}_0$ is defined as in Lemma (ref).

Given the modification to the construction of the bootstrap sample, the remaining steps of the test follow the procedure in Section (ref). For dependent data, the test is also asymptotically size controlled and consistent, as shown in Theorem (ref).

thrySuppose that Assumptions (ref)--(ref), (ref)--(ref), and (ref)--(ref) hold, and that $\sqrt{n}(\widehat{\phi}_n^*-\widehat{\phi}_n)$ is asymptotically measurable. \begin{enumerate}[label=(\roman*),nosep] • If $\mathrm{H}_0$ is true and the CDF of $\mathcal{L}''_{\phi_P}\left( \mathbb{G}_0 \right)$ is strictly increasing and continuous at its $1-\alpha$ quantile, then \begin{align*} \lim_{n\to\infty} \mathbb{P}\left( n \mathcal{L}( \widehat{\phi}_n )>\widehat{c}_{1-\alpha,n} \right) = \alpha. \end{align*} • If $\mathrm{H}_0$ is false, then \begin{align*} \lim_{n\to\infty} \mathbb{P}\left( n \mathcal{L}( \widehat{\phi}_n )>\widehat{c}_{1-\alpha,n} \right)=1. \end{align*} \end{enumerate}

Monte Carlo Experiments

In this section, we construct the Monte Carlo experiments based on the conditional moment restriction models with weak instrumental variables (IVs) in jun2009semiparametric. Let $y_i$ be a scalar outcome variable, $Y_i$ be a scalar endogenous variable, and $z_i$ be a scalar instrumental variable. The model of interest is

align[align omitted — 116 chars of source]

for a true structural parameter $\theta_0\in\Theta\subset\mathbb{R}$. We consider the null hypothesis

align*[align* omitted — 121 chars of source]

which is equivalent to

align*[align* omitted — 167 chars of source]

As noted by jun2009semiparametric, there are several specification tests for (ref) under strong point identification and the assumption that $\theta_0$ can be $\sqrt{n}$-consistently estimated under the null bierens1990consistent,zheng1996consistent,fan1996consistent,fan2000consistent. Since $\mathbb{E}_P[y_i-Y_i\theta_0|z_i]=0$ almost surely for some $\theta_0$ under the null, typical estimators of $\theta_0$ include two-stage least squares (2SLS) and semi-parametric methods. However, when instruments are weak, these estimators of $\theta_0$ may be undesirable staiger1997instrumental,stock2000gmm,jun2012testing, and thus two-step tests plugging in preliminary estimators of $\theta_0$ may not perform well.

jun2009semiparametric propose semi-parametric specification tests of conditional moment restrictions with weak instruments, which do not require a consistent first-step estimator. As shown in Theorems 1 and 2 of jun2009semiparametric, their tests yield limiting rejection probabilities no greater than the nominal significance level under the null. Based on their Example II, jun2009semiparametric study the finite sample performance of their tests with weak instruments via Monte Carlo experiments. We first follow jun2009semiparametric and consider two cases under the null. In the first case (Case 1), $\mathbb{E}_P[z_iY_i]=0$, that is, the rank condition fails and $z_i$ is not a valid instrument for $Y_i$ when estimating $\theta_0$ by 2SLS in two-step tests, which may be seen as an extreme case of weak instruments. Thus, the 2SLS estimator of $\theta_0$ that uses $z_i$ as the instrument for $Y_i$ is unreliable. In the second case (Case 2), $\mathbb{E}_P[Y_i|z_i]\to 0$ almost surely as $n\to\infty$, that is, all measurable functions $f$ of $z_i$ with $\mathbb{E}_P[|f(z_i)Y_i|]<\infty$ may be weak instruments for $Y_i$ when estimating $\theta_0$ by 2SLS in two-step tests because $\mathbb{E}_P[f(z_i)Y_i]$ may converge to $0$. As discussed in jun2012testing, semi-parametric estimators of $\theta_0$ may also break down when $\mathbb{E}_P[Y_i|z_i]$ decays too fast in $n$. In addition, we consider a third case (Case 3), which is an extreme case of Case 2: $\mathbb{E}_P[Y_i|z_i]= 0$ almost surely.

As demonstrated in Tables 1 and 2 of jun2009semiparametric, their tests improve greatly upon two-step plug-in methods in the presence of weak instruments, while they are often conservative, which is in line with their theoretical results. The proposed test in this paper is asymptotically exactly size controlled and consistent under certain conditions, regardless of the strength of instruments. We numerically present these properties through Monte Carlo experiments, where the DGPs are designed for conditional moment restriction models with weak instruments as in the above cases.

Now we introduce the designs of our simulations. For i.i.d.\ samples, we follow the design of jun2009semiparametric:

align*[align* omitted — 89 chars of source]

where $\{(u_i,v_i,z_i):i=1,\ldots,n\}$ are i.i.d.\ with

align*[align* omitted — 211 chars of source]

The aforementioned three cases are realized in the following manner:

itemize• Case 1: $\rho=0.5$, $\lambda=1$, and $g(z)=z^2-1$. The moment $\mathbb{E}_P[z_iY_i]=0$. • Case 2: $\rho=-0.99$, $\lambda=0.07\sqrt{200/n}$, and $g(z)=z$. The conditional moment $\mathbb{E}_P[Y_i|z_i]\to 0$ almost surely as $n\to\infty$. • Case 3: $\rho=-0.5$ and $\lambda=0$. The conditional moment $\mathbb{E}_P[Y_i|z_i]=0$ almost surely.

For each case, we consider four DGPs characterized by the values of $\delta$:

itemize• DGP (0): $\delta=0$. The null is true. • DGP (1): $\delta=0.2$. The null is false. • DGP (2): $\delta=0.6$. The null is false. • DGP (3): $\delta=1$. The null is false.

We also consider dependent data. For every DGP introduced above, we construct the dependent-data counterpart by generating $\{z_i:i=1,\ldots,n\}$ as

align*[align* omitted — 52 chars of source]

where $\{\varepsilon_{i}:i=1,\ldots,n\}$ are i.i.d.\ $\mathcal{N}(0,1)$ and independent of $\{(u_i,v_i):i=1,\ldots,n\}$.

remarkFor illustration of the high level assumptions in Section (ref), we consider Case 1 with $\delta=0$. With $\theta_{0}=1$, $\mathrm{H}_{0}$ is true and thus \[ \mathbb{E}_{P}\left[ \left( y_{i}-Y_{i}\theta_{0}\right) 1\left\{ z_{i}\leq x\right\} \right] =0 \] for all $x$. We have that for all $\theta$, \begin{align*} \phi_{P}\left( x,\theta\right) & =\mathbb{E}_{P}\left[ \left( y_{i}-Y_{i}\theta\right) 1\left\{ z_{i}\leq x\right\} \right] \\ & =\mathbb{E}_{P}\left[ y_{i}1\left\{ z_{i}\leq x\right\} \right] -\mathbb{E}_{P}\left[ Y_{i}1\left\{ z_{i}\leq x\right\} \right] \theta. \end{align*} It follows that for all $\theta$, \begin{align*} & \int_{\mathbb{R}}\phi_{P}\left( x,\theta\right) ^{2}\mathrm{d}\nu\left( x\right) \\ =&\,\int_{\mathbb{R}}\mathbb{E}_{P}\left[ y_{i}1\left\{ z_{i}\leq x\right\} \right] ^{2}\mathrm{d}\nu\left( x\right) -2\int_{\mathbb{R}}\mathbb{E} _{P}\left[ y_{i}1\left\{ z_{i}\leq x\right\} \right] \mathbb{E}_{P}\left[ Y_{i}1\left\{ z_{i}\leq x\right\} \right] \mathrm{d}\nu\left( x\right) \theta\\ &+\int_{\mathbb{R}}\mathbb{E}_{P}\left[ Y_{i}1\left\{ z_{i}\leq x\right\} \right] ^{2}\mathrm{d}\nu\left( x\right) \theta^{2}\geq0. \end{align*} The value $\theta_{0}=1$ satisfies $\int_{\mathbb{R}}\phi_{P}\left( x,\theta_{0}\right) ^{2}\mathrm{d}\nu\left( x\right) =0$, so we have \[ \Theta_{0}=\left\{ \theta_{0}\right\} ,\theta_{0}=\frac{\int_{\mathbb{R} }\mathbb{E}_{P}\left[ y_{i}1\left\{ z_{i}\leq x\right\} \right] \mathbb{E}_{P}\left[ Y_{i}1\left\{ z_{i}\leq x\right\} \right] \mathrm{d}\nu\left( x\right) }{\int_{\mathbb{R}}\mathbb{E}_{P}\left[ Y_{i}1\left\{ z_{i}\leq x\right\} \right] ^{2}\mathrm{d}\nu\left( x\right) }=1. \] For every $\varepsilon>0$, \begin{align*} & \int_{\mathbb{R}}\phi_{P}\left( x,\theta_{0}-\varepsilon\right) ^{2}\mathrm{d}\nu\left( x\right) \\ =&\,\int_{\mathbb{R}}\mathbb{E}_{P}\left[ y_{i}1\left\{ z_{i}\leq x\right\} \right] ^{2}\mathrm{d}\nu\left( x\right) -2\int_{\mathbb{R}}\mathbb{E} _{P}\left[ y_{i}1\left\{ z_{i}\leq x\right\} \right] \mathbb{E}_{P}\left[ Y_{i}1\left\{ z_{i}\leq x\right\} \right] \mathrm{d}\nu\left( x\right) \left( \theta_{0}-\varepsilon\right) \\ & +\int_{\mathbb{R}}\mathbb{E}_{P}\left[ Y_{i}1\left\{ z_{i}\leq x\right\} \right] ^{2}\mathrm{d}\nu\left( x\right) \left( \theta_{0}-\varepsilon \right) ^{2}\\ =&\,2\int_{\mathbb{R}}\mathbb{E}_{P}\left[ y_{i}1\left\{ z_{i}\leq x\right\} \right] \mathbb{E}_{P}\left[ Y_{i}1\left\{ z_{i}\leq x\right\} \right] \mathrm{d}\nu\left( x\right) \varepsilon-2\int_{\mathbb{R} }\mathbb{E}_{P}\left[ Y_{i}1\left\{ z_{i}\leq x\right\} \right] ^{2}\mathrm{d}\nu\left( x\right) \theta_{0}\varepsilon\\ & +\int_{\mathbb{R}}\mathbb{E}_{P}\left[ Y_{i}1\left\{ z_{i}\leq x\right\} \right] ^{2}\mathrm{d}\nu\left( x\right) \varepsilon^{2}\\ =&\,\int_{\mathbb{R}}\mathbb{E}_{P}\left[ Y_{i}1\left\{ z_{i}\leq x\right\} \right] ^{2}\mathrm{d}\nu\left( x\right) \varepsilon^{2}. \end{align*} This implies that \[ \left\{ \int_{\mathbb{R}}\phi_{P}\left( x,\theta_{0}-\varepsilon\right) ^{2}\mathrm{d}\nu\left( x\right) \right\} ^{1/2}=\left\{ \int_{\mathbb{R} }\mathbb{E}_{P}\left[ Y_{i}1\left\{ z_{i}\leq x\right\} \right] ^{2}\mathrm{d}\nu\left( x\right) \right\} ^{1/2}\varepsilon. \] In this case, Assumptions (ref) and (ref) hold. The asymptotic limit of the test statistic is \begin{align} \mathcal{L}_{\phi_{P}}^{\prime\prime}\left( \mathbb{G}_{0}\right) =\inf_{v\in \mathbb{R} }\int _{\mathbb{R}}\left( \mathbb{G}_{0}\left( x,\theta_{0}\right) -\mathbb{E} _{P}\left[ Y_{i}1\left\{ z_{i}\leq x\right\} \right] v\right) ^{2}\mathrm{d}\nu\left( x\right) . \end{align} Theorem (ref)(i) requires that the CDF of $\mathcal{L}_{\phi_{P}}^{\prime\prime }\left( \mathbb{G}_{0}\right) $ in (ref) is strictly increasing and continuous at its $1-\alpha$ quantile.

The sample size is set to $n\in\{100,200,400,800\}$. We set the tuning parameter $\tau_n$ as $\tau_n=\sqrt{\ln(n)/n}$, $n^{-2/5}$, $n^{-1/3}$, $n^{-1/4}$, $n^{-1/5}$, and $n^{-1/6}$, which all satisfy Assumption (ref). For dependent data, the moving blocks bootstrap involves an additional tuning parameter $b(n)$. We set $b(n)=n^{1/6}$, $n^{1/5}$, $n^{1/4}$, and $n^{1/3}$. Recall that the test statistic involves an integration with respect to a measure $\nu$ and an infimum. The integration is approximated by an equally weighted average on the grid $\{-3,-2.998, -2.996, \ldots, 3\}$ of $x$, and the infimum is achieved by a search on the grid $\{0.7, 0.702, 0.704, \ldots, 1.3\}$ of $\theta$. Furthermore, we apply the warp-speed method giacomini2013warp to implement all the Monte Carlo experiments. Specifically, for each DGP and sample size, we generate $1000$ samples and compute one original statistic $n\mathcal{L}(\widehat{\phi}_n)$ and one bootstrap statistic $\widehat{\mathcal{L}}''_n[\sqrt{n}(\widehat{\phi}_n^*-\widehat{\phi}_n)]$ for each sample. The critical value $\widehat{c}_{1-\alpha,n}$ is approximated by the $(1-\alpha)$-empirical quantile of the $1000$ bootstrap statistics, and the rejection rate is computed by comparing the $1000$ original statistics with the critical value $\widehat{c}_{1-\alpha,n}$.

We present some main simulation results in the following and leave the remaining results to Section (ref) of the Online Supplementary Appendix. Tables (ref)--(ref) and (ref)--(ref) show the rejection rates for different DGPs, tuning parameters, and nominal significance levels with measure $\nu$ being the probability measure of $\mathcal{N}(0,10^2)$. Tables (ref)--(ref) display the rejection rates for Case 1 with the measure $\nu$ being the probability measure of $\mathcal{N}(0,1)$ or $\mathcal{N}(0,5^2)$. The results are stable for different choices of $\tau_n$, $b(n)$, and $\nu$. Most of the rejection rates under the null are close to the nominal significance levels. The rejection rates under the alternatives increase to one as the sample size $n$ increases. For dependent samples, the rejection rates under the null may exceed the significance level $\alpha$ for some $t_n$, $b(n)$, and $\nu$ as shown, for example, in Table (ref). As we increase the sample sizes, the results become closer to $\alpha$, as shown in Table (ref).

table[table omitted — 1,974 chars of source]
table[table omitted — 1,359 chars of source]
table[table omitted — 1,691 chars of source]
table[table omitted — 1,706 chars of source]
table[table omitted — 1,706 chars of source]
table[table omitted — 1,706 chars of source]

Performance Improvement in Conditional Moment Restriction Models with Weak Instruments

Note that Cases 1 and 2 ($n=200$) with $\delta=0$ (under the null hypothesis) are identical to the DGPs in Tables 1 and 2 of jun2009semiparametric, respectively. Thus, the results in Table (ref) with $n=100$ and Table (ref) with $n=200$ can be compared with those in Tables 1 and 2 of jun2009semiparametric, respectively. We present the comparisons in Tables (ref) and (ref) below, where $\widehat{T}_2(\widehat{\theta}_{*})$ and $\widehat{T}_\mathrm{k}(\widehat{\theta}_{*})$ with $\widehat{\theta}_{*}\in \{\widehat{\theta}_{\mathrm{2SLS}}, \widehat{\theta}_{\mathrm{SP}}\}$ are two-step plug-in test statistics computed by using either a 2SLS or a semi-parametric estimator of $\theta_0$ as described in jun2009semiparametric, and $\widehat{T}_1(\widehat{\theta}_{\mathrm{CUE1}})$, $\widehat{T}_2(\widehat{\theta}_{\mathrm{CUE2}})$, and $\widehat{T}_\mathrm{k}(\widehat{\theta}_{\mathrm{CUEk}})$ are the test statistics proposed by jun2009semiparametric.\footnote{The function $\widehat{T}_\mathrm{k}(\cdot)$ is proposed by zheng1996consistent, and the test statistic $\widehat{T}_\mathrm{k}(\widehat{\theta}_{\mathrm{CUEk}})$ based on minimization of $\widehat{T}_\mathrm{k}(\cdot)$ follows the idea of jun2009semiparametric.} Our test uses $\tau_n=n^{-1/4}$ for illustration. The plug-in method suffers from substantial size distortion. The tests of jun2009semiparametric improve upon the plug-in approach, but could be conservative as shown in their theoretical results. The proposed method achieves rejection rates closer to the nominal significance levels compared to the results of jun2009semiparametric. These numerical observations provide supporting evidence for the theoretical results in the paper.

table[table omitted — 1,080 chars of source]
table[table omitted — 1,078 chars of source]

Conclusion

This paper provides a unified framework for inference on moment restriction models with nuisance parameters. We employ a new characterization that does not require the estimation of nuisance parameters, along with a numerical delta method to construct the test. The test is asymptotically size controlled and consistent. We conduct extensive Monte Carlo simulations to illustrate the finite sample properties of the proposed test. The numerical results show that the proposed method may achieve improvement in testing conditional moment restriction models with weak instruments.

\setcounter{page}{1} \setcounter{equation}{0} \setcounter{footnote}{0}

center[center omitted — 229 chars of source]

The online supplementary appendix consists of five sections. Section (ref) provides auxiliary lemmas. Section (ref) verifies the assumptions for the examples in the main text. Section (ref) extends the results for location-scale transformation to general parametric transformations on multiple CDFs. Section (ref) contains the proofs of all main results. Section (ref) provides additional simulation results.