EconBase
← Back to paper

Optimal testing in a class of nonregular models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

37,768 characters · 8 sections · 32 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Optimal testing in a class of nonregular models

\address{Department of Economics, University of Wisconsin, Madison, 1180 Observatory Drive, Madison, WI 53706-1393, USA.} \email{[email removed]} \address{Department of Economics, London School of Economics, Houghton Street, London, WC2A 2AE, UK.} \email{[email removed]}

abstractThis paper studies optimal hypothesis testing for nonregular econometric models with parameter-dependent support. We consider both one-sided and two-sided hypothesis testing and develop asymptotically uniformly most powerful tests based on a limit experiment. Our two-sided test becomes asymptotically uniformly most powerful without imposing further restrictions such as unbiasedness, and can be inverted to construct a confidence set for the nonregular parameter. Simulation results illustrate desirable finite sample properties of the proposed tests.

Introduction

This paper studies optimal hypothesis testing of a class of nonregular econometric models in which the boundary of the support of the observed data depends on some parameter of interest. Such nonregular models, which typically imply discontinuous likelihood functions and nonstandard convergence rates of estimators, have been often studied in the econometrics literature; see, flinn1982new, smith1985maximum, christensen1991exact, donald1993maximum, hong1998maximum, donald2002superconsistent, hirano2003asymptotic, and chernozhukov2004likelihood, among others. In contrast to most existing papers that focus on point estimation, this paper is concerned with optimal (composite) hypothesis testing for nonregular models instead of point estimation.

For testing a simple null hypothesis against a simple alternative one, Neyman-Pearson's fundamental lemma yields an optimal power property of the likelihood ratio test even for the case of parameter-dependent support. However, the optimality result is no longer available for general testing problems with composite hypotheses. On the other hand, for regular statistical models, it is known that standard testing methods (such as the likelihood ratio, Wald, and score tests) can achieve certain asymptotic optimal power properties for testing general composite hypotheses (see, Chapter 15 of lehmann2022testing). An open question is whether we can establish an analogous asymptotic optimality result for testing composite hypotheses on nonregular parameters in the case of parameter-dependent support, and this paper addresses this question in a positive way.

In this paper, we consider one-sided and two-sided hypothesis testing for parametric models with parameter-dependent support and develop an asymptotically uniformly most powerful (AUMP) test based on a limit experiment. Interestingly, our two-sided test attains the AUMP property without imposing further restrictions such as unbiasedness, and can be inverted to construct a confidence set for the nonregular parameter. For clarity we first present the main results under a benchmark setup in Section (ref), where there is no covariate or nuisance parameter. Then we extend our optimality results to a general setup in Section (ref) that involves covariates and nuisance parameters. For the general case, we need some independent auxiliary sample to estimate the nuisance parameters, which is typically obtained by sample splitting. Our simulation results in Section (ref) illustrate desirable finite sample properties of the proposed tests.

The most closely related papers to ours are hirano2003asymptotic and chernozhukov2004likelihood. hirano2003asymptotic studied efficient point estimation of parameter-dependent support models by extending the limit of experiments argument, and showed that the Bayes estimator is asymptotically efficient under a minimax criterion but the maximum likelihood estimator is generally inefficient. To the best of our knowledge, this paper is the first one that studies optimal hypothesis testing for nonregular models with parameter-dependent support. In contrast to the Bayes estimator in hirano2003asymptotic that involves priors on parameters, our optimal testing methods are developed based on the limiting likelihood ratio process without priors. chernozhukov2004likelihood also investigated nonstandard asymptotic properties of estimation and testing methods for parameter-dependent support models. They established asymptotic optimality of the Bayes estimators in terms of the asymptotic average risk, and showed that the Wald test and Bayes posterior quantiles are valid for inference. However, they did not discuss optimality of the testing methods. In addition to these papers, chen2018monte (chen2018monte, Appendix C) examined partially identified models with parameter-dependent support, and chen2025identification adapted their quasi-Bayesian likelihood ratio statistics to partially identified auction models. In our simulation study, we demonstrate that the proposed test exhibits more reasonable performance than the Wald type test by chernozhukov2004likelihood.

Benchmark case

This section presents our main results for a simple nonregular model with a scalar parameter and no covariate. This model, covering the uniform distribution as a canonical example, provides a useful benchmark to highlight our developments. We propose one-sided tests for this model and derive the asymptotic properties of these tests, including the AUMP. Based on the one-sided tests, we construct an optimal two-sided test with an unequal-tailed way.

Let $Y\in\mathcal{Y}\subset\mathbb{R}$ follow the parametric model

equation[equation omitted — 74 chars of source]

where $\mathbb{I}\{\cdot\}$ is the indicator function, $\theta\in\Theta\subset\mathbb{R}$ is the true value of a scalar parameter, and conditions on the functions $f$ and $g$ are specified below. For this model, we consider the one-sided testing problem

equation[equation omitted — 216 chars of source]

for a given $\theta_{0}\in\Theta$. Hereafter, we focus on the case of $\nabla_{\theta}g(\theta_{0})>0$. The case of $\nabla_{\theta}g(\theta_{0})<0$ is analyzed in the same manner. Based on the reparametrization $h=n(\theta-\theta_{0})$, this testing problem is written as

equation[equation omitted — 162 chars of source]

We wish to test $H_{0}^{-}$ or $H_{0}^{+}$ based on an independent and identically distributed (iid) sample $Y^{n}=(Y_{1},\ldots,Y_{n})$ of $Y$. Although our results can be easily extended to cover a more general reparametrization $h=n(\theta-\theta_{0})+\zeta_{n}$ with $\zeta_{n}=o(1)$, we focus on the case of $\zeta_{n}=0$ for simplicity.

Our testing procedure is constructed based on the limiting likelihood ratio process. Since the joint density of $Y^{n}$ is written as \[ \mathrm{d}P_{\theta}^{n}(y^{n})=\mathbb{I}\{y_{(1)}\ge g(\theta)\}\prod_{i=1}^{n}f(y_{i}|\theta), \] where $y_{(1)}=\min\{y_{1},\ldots,y_{n}\}$, the likelihood ratio process on a parameter space $\mathcal{H}\subset\mathbb{R}$ is \[ Z_{n}(h,\bar{h}):=\frac{\mathrm{d}P_{\theta_{0}+\frac{\bar{h}}{n}}^{n}}{\mathrm{d}P_{\theta_{0}+\frac{h}{n}}^{n}}(Y^{n})=\frac{\mathbb{I}\{Y_{(1)}\ge g(\theta_{0}+\bar{h}/n)\}}{\mathbb{I}\{Y_{(1)}\ge g(\theta_{0}+h/n)\}}\prod_{i=1}^{n}\frac{f(Y_{i}|\theta_{0}+\bar{h}/n)}{f(Y_{i}|\theta_{0}+h/n)}, \] for $h,\bar{h}\in\mathcal{H}$. To characterize asymptotic properties of this process, we impose the following assumptions.

asm$\quad$ \begin{description} • $\{Y_{i}\}_{i=1}^{n}$ is an iid sample of $Y\in\mathcal{Y}\subset\mathbb{R}$ with the Lebesgue density in ((ref)). The parameter space $\Theta$ is convex. • $f(y|\theta)$ is twice continuously differentiable in $\theta$ for all $y$. In some open neighborhood $\mathcal{N}$ of $\theta_{0}$, $f(y|\theta)$ and $\nabla_{\theta}f(y|\theta)$ are continuous in $y$ for $\theta\in\mathcal{N}$, there exists a constant $C$ such that $0<f(y|\theta)<C<\infty$ for all $y$ and $\theta\in\mathcal{N}$, and \[ \begin{aligned} & \int\sup_{\theta\in\mathcal{N}}\left\Vert \nabla_{\theta}f(y|\theta)\right\Vert \mathbb{I}\{y\geq g(\theta)\}dy<\infty,\\ & \int\sup_{\theta,\bar{\theta}\in\mathcal{N}}\frac{\left\Vert \nabla_{\theta}f(y|\bar{\theta})\right\Vert ^{2}}{f(y|\bar{\theta})^{2}}\mathbb{I}\{y\geq g(\theta)\}f(y|\theta)dy<\infty,\\ & \int\sup_{\theta,\bar{\theta}\in\mathcal{N}}\frac{\left\Vert \nabla_{\theta\theta}f(y|\bar{\theta})\right\Vert }{f(y|\bar{\theta})}\mathbb{I}\{y\geq g(\theta)\}f(y|\theta)dy<\infty. \end{aligned} \] $g(\theta)$ is continuously differentiable in $\theta$, $\nabla_{\theta}g(\theta_{0})>0$, and $\sup_{\theta\in\mathcal{N}}\left\Vert \nabla_{\theta}g(\theta)\right\Vert <\infty$. \end{description}

Assumption (ref) (i) is standard, and Assumption (ref) (ii) contains smoothness and boundedness conditions on the functions $f$ and $g$ in ((ref)). By adapting hirano2003asymptotic (hirano2003asymptotic, Theorem 2) to our setup, Assumption (ref) guarantees weak convergence (denoted by “$\rightsquigarrow$”) of the likelihood ratio process

equation[equation omitted — 212 chars of source]

for every finite $I\subset\mathcal{H}$, where $\lambda=\{f(g(\theta_{0})|\theta_{0})\nabla_{\theta}g(\theta_{0})\}^{-1}$ and \[ D_{h,\bar{h}}:=\mathbb{I}\{W_{h}>\bar{h}\}\text{ with }W_{h}\sim f_{W}(w|h)=\frac{1}{\lambda}e^{-(w-h)/\lambda}\mathbb{I}\{w>h\}. \] Note that $\lambda$ is a known constant in the present setup. In contrast to the standard likelihood ratio process, which is locally asymptotically normal, $Z_{n}(h,\bar{h})$ converges to a limit of experiments whose randomness is given by the binary variable $D_{h,\bar{h}}$. This is due to lack of differentiability in quadratic mean of the density $\mathrm{d}P_{\theta}^{n}(y^{n})$. Since $D_{h,\bar{h}}$ is discrete, the limiting likelihood ratio process $Z(h,\bar{h})$ is discontinuous in the sense that $\Pr\{Z(h,\bar{h})\leq z\}$ is not continuous at $z=0$ and $e^{(\bar{h}-h)/\lambda}$. Thus, we cannot use the conventional asymptotic theory based on the quadratic expansion of the likelihood ratio process to evaluate asymptotic size and power properties of the likelihood ratio test.

Even though the likelihood ratio process $\{Z_{n}(h,\bar{h})\}_{\bar{h}\in I}$ exhibits such nonregularity, it should be noted the limiting likelihood ratio process $\{Z(h,\bar{h})\}_{\bar{h}\in I}$ satisfies the monotone likelihood ratio property, which is defined as follows.

defn[Monotone likelihood ratio] (shao2003mathematical, Definition 6.2) Suppose that the distribution of $W$ is in $\mathcal{P}=\{P_{h}:h\in\mathcal{H}\}$, a parametric family indexed by a real-valued $h$, and that $\mathcal{P}$ is dominated by a $\sigma$-finite measure $\mu$. The family $\mathcal{P}$ is said to have monotone likelihood ratio in $S(W)$ (a real-valued statistic) if and only if, for any $h_{1}<h_{2}$, $\mathrm{d}P_{h_{2}}/\mathrm{d}P_{h_{1}}(w)$ is a nondecreasing function of $S(w)$ for values $w$ at which at least one of $\mathrm{d}P_{h_{1}}(w)$ and $\mathrm{d}P_{h_{2}}(w)$ is positive.

A key feature of distributions with monotone likelihood ratio is the existence of the uniformly most powerful (UMP) test. Relying on the limiting likelihood ratio process is essential since the finite sample likelihood ratio does not exhibit a monotone likelihood ratio property in general.

lem(shao2003mathematical, Theorem 6.2.1) Suppose that a random variable $U_{h}$ has a distribution in $\mathcal{P}=\{P_{h}:h\in\mathcal{H}\subset\mathbb{R}\}$ that has monotone likelihood ratio in $S(U_{h})$. Consider the problem of testing $H_{0}:h\leq h_{0}$ against $H_{1}:h>h_{0}$, where $h_{0}$ is a given constant. Then there exists a UMP test of size $\alpha$ given by \begin{equation} \phi(U_{h})=\begin{cases} 1 & if S(U_{h})>c\\ \kappa & if S(U_{h})=c\\ 0 & if S(U_{h})<c \end{cases}, \end{equation} where $c$ and $\kappa$ are determined by $E_{h_{0}}[\phi(U_{h})]=\alpha$.

Here $\phi(U_{h})=1$ and $0$ mean rejection and acceptance of $H_{0}$, respectively, and $\phi(U_{h})=\kappa$ means rejection with probability $\kappa$. Therefore, to derive an AUMP test for $H_{0}^{-}:h\leq0$, we can still invoke the asymptotic representation lemma below to argue that the sample counterpart of the monotone likelihood ratio test with $S(W_{h})=W_{h}$ is AUMP.

lem(van2000asymptotic, Theorem 15.1) Let the sequence of experiments $\mathcal{E}_{n}=\{P_{n,h}:h\in\mathcal{H}\}$ converge to a dominated experiment $\mathcal{E}=\{P_{h}:h\in\mathcal{H}\}$. Suppose that a sequence of power functions $\pi_{n}$ of tests in $\mathcal{E}_{n}$ converges pointwise, i.e., $\pi_{n}(h)\rightarrow\pi(h)$ for every $h$ and some function $\pi$. Then $\pi$ is a power function in the limit experiment, i.e., there exists a test $\phi$ in $\mathcal{E}$ with $\pi(h)=E_{h}[\phi(X)]$ for every $h$.

To derive an AUMP test for the other one-sided test $H_{0}^{+}:h\geq0$, we utilize the fact that the limit experiment can be represented as a function of the random variable $W_{h}$. Accordingly, we apply the Neyman-Pearson lemma (e.g., Theorem 3.2.1 (ii) of lehmann2022testing) to $W_{h}$.

Hereafter, we first formalize this argument for one-sided tests on $H_{0}^{-}:h\leq0$ and $H_{0}^{+}:h\geq0$ (Section (ref)), and then extend the argument to two-sided testing for $H_{0}:h=0$ (Section (ref)).

One-sided tests

First, we consider testing $H_{0}^{-}:h\leq0$ against $H_{1}^{-}:h>0$ at the significance level $\alpha\in(0,1)$. By taking a sample counterpart of $\phi(W_{h})$ in ((ref)) with the limiting process $W_{h}$ in ((ref)), our test is constructed as

equation[equation omitted — 303 chars of source]

This test achieves an asymptotic optimal property in Definition $\text{\ref{def:AUMP}}$ below, introduced by choi1996asymptotically.

defn[Asymptotically uniformly most powerful test] For testing $H_{0}:\theta\leq\theta_{0}$ against $H_{1}:\theta>\theta_{0}$ (or $H_{0}:\theta\geq\theta_{0}$ against $H_{1}:\theta<\theta_{0}$ or $H_{0}:\theta=\theta_{0}$ against $H_{1}:\theta\neq\theta_{0}$), a sequence of tests $\{\phi_{n}\}$ is called asymptotically uniformly most powerful (AUMP) in $\mathcal{H}$ at asymptotic level $\alpha$ if $\limsup_{n}E_{\theta_{0}+h/n}[\phi_{n}]\leq\alpha$ for every $h\leq0$ (or $h\geq0$ or $h=0$) in $\mathcal{H}$ and for any other sequence of test functions $\{\psi_{n}\}$ satisfying $\limsup_{n}E_{\theta_{0}+h/n}[\psi_{n}]\leq\alpha$ for every $h\leq0$ (or $h\geq0$ or $h=0$) in $\mathcal{H}$, \[ \liminf_{n}E_{\theta_{0}+h/n}[\phi_{n}]\geq\limsup_{n}E_{\theta_{0}+h/n}[\psi_{n}], \] for every $h>0$ (or $h<0$ or $h\neq0$) in $\mathcal{H}$.

The following asymptotic optimality result is established using the monotone likelihood ratio property of the limit experiment, as demonstrated in Appendix.

thmSuppose that Assumption (ref) holds for the true local parameter $h\in\mathcal{H}$. Then the test $\phi_{n}^{-}(Y^{n})$ is AUMP in $\mathcal{H}$ at level $\alpha$ for testing $H_{0}^{-}:\theta\leq\theta_{0}$ against $H_{1}^{-}:\theta>\theta_{0}$.

Next, we consider another one-sided testing $H_{0}^{+}:h\geq0$ against $H_{1}^{+}:h<0$. The basic idea is same as the previous case. We propose the following test:

equation[equation omitted — 375 chars of source]

and the asymptotic optimality of this test is obtained as follows by applying the Neyman-Pearson lemma to $W_{h}$.

thmSuppose that Assumption (ref) holds holds for the true local parameter $h\in\mathcal{H}$. Then the test $\phi_{n}^{+}(Y^{n})$ is AUMP in $\mathcal{H}$ at level $\alpha$ for testing $H_{0}^{+}:\theta\geq\theta_{0}$ against $H_{1}^{+}:\theta<\theta_{0}$.

Since we assume $\nabla_{\theta}g(\theta_{0})>0$ and $\lambda>0$, we have $g(\theta_{0})<g\left(\theta_{0}+\frac{\lambda}{n}\log\left(\frac{1}{1-\alpha}\right)\right)$ eventually. However, this inequality can be violated in finite samples. We can show that any test $\psi_{n}^{+}$ rejecting the null with probability one if $Y_{(1)}<g(\theta_{0})$ and rejecting the null with probability at most $\alpha$ if $Y_{(1)}\geq g(\theta_{0})$ is AUMP at level $\alpha$ for testing $H_{0}^{+}:\theta\geq\theta_{0}$ against $H_{1}^{+}:\theta<\theta_{0}$. We recommend using ($\text{\ref{eq:phi+}}$) since it avoids randomization.

Two-sided test

This subsection considers two-sided testing $H_{0}:h=0$ against $H_{1}:h\neq0$. By combining the optimal one-sided tests derived in the last subsection, we propose the following (unequal-tailed) two-sided test:

equation[equation omitted — 351 chars of source]

Indeed, this test is shown to be AUMP.

thmSuppose that Assumption (ref) holds for the true local parameter $h\in\mathcal{H}$. Then the proposed test $\phi_{n}(Y^{n})$ is AUMP in $\mathcal{H}$ at level $\alpha$ for testing $H_{0}:\theta=\theta_{0}$ against $H_{1}:\theta\neq\theta_{0}$.

Importantly, the proposed test achieves the AUMP property without imposing additional restrictions such as unbiasedness. This feature is shared by a finite sample UMP two-sided test for the uniform distribution (e.g., lehmann2022testing, lehmann2022testing, Problem 3.2, p.105).

Since the two-sided test $\phi_{n}(Y^{n})$ does not randomize, we can easily construct a $100(1-\alpha)$% confidence set by the test inversion: \[ CS=\left\{ \theta\in\Theta:g(\theta)\leq Y_{(1)}\leq g\left(\theta+\frac{\lambda}{n}\log\left(\frac{1}{\alpha}\right)\right)\right\} . \] This confidence set also has a pointwise optimal property (asymptotically uniformly most accurate).

General case

In this section, as in hirano2003asymptotic, we generalize the benchmark model to accommodate discrete covariates and nuisance parameters:

equation[equation omitted — 94 chars of source]

where $\theta\in\Theta\subset\mathbb{R}$ is a scalar parameter of interest, $\gamma\in\mathbb{R}^{d}$ is a $d$-dimensional vector of (regular) nuisance parameters, $Y$ is a scalar dependent variable, and $X$ is an $m$-dimensional vector of discrete covariates with support $\mathcal{X}=\{a_{1},\dots,a_{L}\}$. This section considers the one-sided testing problem

align*[align* omitted — 296 chars of source]

for a given $\theta_{0}\in\mathbb{R}$ with the asymptotic level of significance $\alpha$. By reparametrization $h=n(\theta-\theta_{0})$, this testing problem is written as

align*[align* omitted — 240 chars of source]

Let $(Y^{n},X^{n})=((Y_{1},\ldots,Y_{n}),(X_{1},\ldots,X_{n}))$ be an iid sample of $(Y,X)\in\mathbb{R}\times\mathcal{X}$. To extend our benchmark results in the last section, we consider the plug-in likelihood ratio process

equation[equation omitted — 298 chars of source]

for $h\in\mathcal{H}$ with a parameter space $\mathcal{H}\subset\mathbb{R}$, where $\hat{\gamma}_{n}$ and $\hat{h}_{n}$ are some estimators of the nuisance parameters $\gamma$ and the user-specified alternative value $\bar{h}$, respectively, based on auxiliary data independent from the main sample $(Y^{n},X^{n})$. In contrast to the benchmark case, where the threshold for testing takes the form $g(\theta_{0}+\bar{h}/n)$, the optimal values of $\bar{h}$ for our test depend on certain population quantities that must be estimated (see Theorems (ref) and (ref) below). Thus we introduce an estimator $\hat{h}_{n}$ for $\bar{h}$ in this general case. Typically we split the sample (say, $(Y^{2n},X^{2n})$) into the main sample $(Y^{n},X^{n})$ and auxiliary one $((Y_{n+1},\ldots,Y_{2n}),(X_{n+1},\ldots,X_{2n}))$ to obtain $\hat{\gamma}_{n}$ and $\hat{h}_{n}$.

In this section, we impose the following assumptions.

asm$\quad$ \begin{description} • $\{Y_{i},X_{i}\}_{i=1}^{n}$ is an iid sample of $\mathcal{Y}\times\mathcal{X}$, where $\mathcal{Y}\subset\mathbb{R}$ and $\mathcal{X}=\{a_{1},\dots,a_{L}\}$ with the conditional density in ((ref)). The parameter space $\Theta\times\Gamma$ of $(\theta,\gamma)$ is convex. • $\{\hat{h}_{n}\}$ is a random sequence independent from $(Y^{n},X^{n})$ satisfying $\sqrt{n}(\hat{h}_{n}-\bar{h})=O_{P_{\theta_{0}+\frac{h}{n},\gamma}}(1)$ for some user specified constant $\bar{h}$. $\{\hat{\gamma}_{n}\}$ is a random sequence independent from $(Y^{n},X^{n})$ satisfying $\sqrt{n}(\hat{\gamma}_{n}-\gamma)=O_{P_{\theta_{0}+\frac{h}{n},\gamma}}(1)$. • Let $\beta=(\theta,\gamma^{\prime})^{\prime}$ and $\beta_{0}=(\theta_{0},\gamma^{\prime})^{\prime}$. $f(y|x,\beta)$ is twice continuously differentiable in $\beta$ for all $y$ and $x$, and $g(x,\theta)$ is continuously differentiable in $\theta$ for all $x$. In some open neighborhood $\mathcal{N}$ of $\beta_{0}$, $f(y|x,\beta)$ and $\nabla_{\beta}f(y|x,\beta)$ are continuous in $y$ for $\beta\in\mathcal{N}$, there exists a constant $C$ such that $0<f(y|x,\beta)<C<\infty$ for all $y$, $x$, and $\beta\in\mathcal{N}$, and for each $j=1,\ldots,L$, \[ \begin{aligned} & \int\sup_{\beta\in\mathcal{N}}\left\Vert \nabla_{\beta}f(y\mid a_{j},\beta)\right\Vert \mathbb{I}\{y\geq g(a_{j},\theta)\}dy<\infty,\\ & \int\sup_{\bar{\beta},\beta\in\mathcal{N}}\frac{\left\Vert \nabla_{\beta}f(y\mid a_{j},\bar{\beta})\right\Vert ^{2}}{f(y\mid a_{j},\bar{\beta})^{2}}\mathbb{I}\{y\geq g(a_{j},\theta)\}f(y\mid a_{j},\beta)dy<\infty,\\ & \int\sup_{\bar{\beta},\beta\in\mathcal{N}}\frac{\left\Vert \nabla_{\beta\beta}f(y\mid a_{j},\bar{\beta})\right\Vert }{f(y\mid a_{j},\bar{\beta})}\mathbb{I}\{y\geq g(a_{j},\theta)\}f(y\mid a_{j},\beta)dy<\infty. \end{aligned} \] $\nabla_{\theta}g(a_{j},\theta_{0})>0$, $\sup_{\beta\in\mathcal{N}}\left\Vert \nabla_{\theta}g(a_{j},\theta)\right\Vert <\infty$, and $\Pr\{X=a_{j}\}>0$ for each $j=1,\ldots,L$. \end{description}

As specified in Assumption (ref) (i), we focus on the case where covariates are discrete variables as in hirano2003asymptotic. Assumption (ref) (ii) requires that the nuisance parameters $\gamma$ and the alternative value $\bar{h}$ to construct our test statistics below can be estimated at the $\sqrt{n}$ rate. For example, based on an auxiliary sample independent from $(Y^{n},X^{n})$, $\gamma$ can be estimated by the maximum likelihood and $\bar{h}$ can be estimated by the method of moments. This assumption enables us to focus on a neighborhood to establish the weak convergence in Lemma $\text{\ref{lem:weak_conv_nui}}$ below. Since $\hat{h}_{n}$ and $\hat{\gamma}_{n}$ are independent from the main sample $(Y^{n},X^{n})$, we can argue the weak convergence in a straightforward way even after conditioning on the concentration to the neighborhood. Assumption (ref) (iii) lists boundedness and smoothness conditions on the functions $f$ and $g$.

To present the limiting distribution of the plug-in likelihood ratio process $\{Z_{n}(h,\hat{h}_{n},\hat{\gamma}_{n})\}$, we introduce further notations. Define

eqnarray*[eqnarray* omitted — 278 chars of source]

$G_{j}=\nabla_{\theta}g(a_{j},\theta_{0})$, and $\lambda_{j}=\{\Pr(X=a_{j})f(g(a_{j},\theta_{0})|a_{j},\beta_{0})\}^{-1}$. Also let $(W_{h,1},\cdots,W_{h,L})$ be mutually independent random variables that follow \[ W_{h,j}\sim f_{W_{j}}(w\mid h,\gamma)=e^{-(w-G_{j}h)/\lambda_{j}}\mathbb{I}\{w>G_{j}h\}/\lambda_{j}. \] The weak convergence of the plug-in likelihood ratio process is established as follows.

lemUnder Assumption (ref), it holds \begin{equation} \{Z_{n}(h,\hat{h}_{n},\hat{\gamma}_{n})\}_{\bar{h}\in I}\rightsquigarrow\{Z(h,\bar{h})\}_{\bar{h}\in I}:=\{e^{(\bar{h}-h)/\lambda}D_{h,\bar{h}}\}_{\bar{h}\in I}\quad under \ensuremath{P_{\theta_{0}+\frac{h}{n},\gamma}}, \end{equation} for every finite $I\subset\mathcal{H}$, where $\lambda=(\sum_{j=1}^{L}G_{j}/\lambda_{j})^{-1}$ and $D_{h,\bar{h}}=\prod_{j=1}^{L}\mathbb{I}\{W_{h,j}>G_{j}\bar{h}\}$. Moreover, we can show the convergences for the components as \begin{equation} \prod_{i=1}^{n}\mathbb{I}\{Y_{i}\ge g(X_{i},\theta_{0}+\hat{h}_{n}/n)\}\rightsquigarrow D_{h,\bar{h}}\quadunder P_{\theta_{0}+\frac{h}{n}}, \end{equation} and \begin{equation} \prod_{i=1}^{n}\frac{f(Y_{i}|X_{i},\theta_{0}+\hat{h}_{n}/n,\hat{\gamma}_{n})}{f(Y_{i}|X_{i},\theta_{0}+h/n,\hat{\gamma}_{n})}\stackrel{p}{\rightarrow}e^{(\bar{h}-h)/\lambda}\quadunder P_{\theta_{0}+\frac{h}{n}}, \end{equation} for any $\bar{h}\in\mathcal{H}$.

This lemma is different from the weak convergence in ($\text{\ref{eq:Z}}$) for the benchmark case in the following aspects. First, the process $\{Z_{n}(h,\hat{h}_{n},\hat{\gamma}_{n})\}$ contains the estimated parameters $\hat{h}_{n}$ and $\hat{\gamma}_{n}$, which also covers the case of deterministic parameter sequences. Second, due to presence of the discrete covariates $X\in\{a_{1},\ldots,a_{L}\}$, the limiting process $Z(h,\bar{h})$ involves an $L$-dimensional random vector $(W_{h,1},\cdots,W_{h,L})$. Third, the limiting process $Z(h,\bar{h})$ depends on the nuisance parameters $\gamma$.

Hereafter, we separately consider testing $H_{0}^{-}:h\leq0,\gamma\in\mathbb{R}^{d}$ against $H_{1}^{-}:h>0,\gamma\in\mathbb{R}^{d}$ and $H_{0}^{+}:h\geq0,\gamma\in\mathbb{R}^{d}$ against $H_{1}^{+}:h<0,\gamma\in\mathbb{R}^{d}$. Let $\hat{\lambda}_{nj}$ and $\hat{\lambda}_{n}$ be consistent estimators of $\lambda_{j}$ and $\lambda$, respectively. For example, based on an auxiliary sample $\{X_{n+1},\ldots,X_{2n}\}$, $\lambda_{j}$ can be estimated by $\hat{\lambda}_{nj}=\{n^{-1}\sum_{i=n+1}^{2n}\mathbb{I}\{X_{i}=a_{j}\}f(g(a_{j},\theta_{0})|a_{j},\theta_{0},\hat{\gamma}_{n})\}^{-1}$, and then $\lambda$ can be estimated by $\hat{\lambda}_{n}=(\sum_{j=1}^{L}G_{j}/\hat{\lambda}_{nj})^{-1}$.

One-sided tests

First, we consider one-sided testing $H_{0}^{-}:h\leq0,\gamma\in\mathbb{R}^{d}$ against $H_{1}^{-}:h>0,\gamma\in\mathbb{R}^{d}$ at the significance level $\alpha\in(0,1)$. By extending the one-sided test in ((ref)), our test is defined as

equation[equation omitted — 283 chars of source]

where $\hat{h}_{n}^{-}=\hat{\lambda}_{n}\log\left(\frac{1}{\alpha}\right)$. The idea to construct this test is essentially the same as the benchmark case in Section (ref). The main difference is that $\lambda$ is unknown and needs to be estimated by plugging-in the consistent estimator $\hat{\lambda}_{n}$ that is independent from the data. The next theorem shows that this test achieves an asymptotic optimal property.

thmSuppose that Assumption (ref) holds for the true local parameter $h\in\mathcal{H}$. Let $\{\hat{h}_{n}^{-}\}$ be any random sequence independent from the data such that $\sqrt{n}(\hat{h}_{n}^{-}-\bar{h}^{-})=O_{P_{\theta_{0}+\frac{h}{n},\gamma}}(1)$ for $\bar{h}^{-}=\lambda\log\left(\frac{1}{\alpha}\right)$ and $\{\hat{\lambda}_{n1},\cdots,\hat{\lambda}_{nL}\}$ be any random sequence independent from the data such that $\sqrt{n}(\hat{\lambda}_{nj}-\lambda_{j})=O_{P_{\theta_{0}+\frac{h}{n},\gamma}}(1)$ for each $j=1,\cdots,L$. Then the test $\phi_{n}^{-}(\hat{h}_{n}^{-},Y^{n},X^{n})$ defined with $\hat{h}_{n}^{-}$ and $\hat{\lambda}_{n}$ constructed by $\{\hat{\lambda}_{n1},\cdots,\hat{\lambda}_{nL}\}$ independent from the data is AUMP in $\mathcal{H}$ at level $\alpha$ for testing $H_{0}^{-}$ against $H_{1}^{-}$.

Next, we consider another one-sided testing $H_{0}^{+}:h\ge0,\gamma\in\mathbb{R}^{d}$ against $H_{1}^{+}:h<0,\gamma\in\mathbb{R}^{d}$. In this case, our test is defined as

eqnarray[eqnarray omitted — 370 chars of source]

where $\hat{h}_{n}^{+}=\hat{\lambda}_{n}\log\left(\frac{1}{1-\alpha}\right)$. Similar to Theorem (ref), asymptotic optimality of this test is obtained as follows.

thmSuppose that Assumption (ref) holds for the true local parameter $h\in\mathcal{H}$. Let $\{\hat{h}_{n}^{+}\}$ be any random sequence independent from the data such that $\sqrt{n}(\hat{h}_{n}^{+}-\bar{h}^{+})=O_{P_{\theta_{0}+\frac{h}{n},\gamma}}(1)$ for $\bar{h}^{+}=\lambda\log\left(\frac{1}{1-\alpha}\right)$ and and $\{\hat{\lambda}_{n1},\cdots,\hat{\lambda}_{nL}\}$ be any random sequence independent from the data such that $\sqrt{n}(\hat{\lambda}_{nj}-\lambda_{j})=O_{P_{\theta_{0}+\frac{h}{n},\gamma}}(1)$ for each $j=1,\cdots,L$. Then the test $\phi_{n}^{+}(\hat{h}_{n}^{+},Y^{n},X^{n})$ defined with $\hat{\lambda}_{n}$ constructed by $\{\hat{\lambda}_{n1},\cdots,\hat{\lambda}_{nL}\}$ independent from the data is AUMP in $\mathcal{H}$ at level $\alpha$ for testing $H_{0}^{+}$ against $H_{1}^{+}$.

Two-sided test

This subsection considers two-sided testing $H_{0}:h=0,\gamma\in\mathbb{R}^{d}$ against $H_{1}:h\neq0,\gamma\in\mathbb{R}^{d}$. By combining the optimal one-sided tests derived in the last subsection, we propose the following (unequal-tailed) two-sided test:

equation[equation omitted — 394 chars of source]
thmSuppose that Assumption (ref) holds for the true local parameter $h\in\mathcal{H}$. Let $\{\hat{h}_{n}^{-}\}$ and $\{\hat{\lambda}_{n1},\cdots,\hat{\lambda}_{nL}\}$ be any random sequences defined in Theorem $\ref{thm:opt-}$. Then the test $\phi_{n}(\hat{h}_{n}^{-},Y^{n},X^{n})$ defined with $\hat{\lambda}_{n}$ constructed by $\{\hat{\lambda}_{n1},\cdots,\hat{\lambda}_{nL}\}$ independent from the data is AUMP in $\mathcal{H}$ at level $\alpha$ for testing $H_{0}$ against $H_{1}$.

Based on this two-sided test, we can construct an asymptotically optimal $100(1-\alpha)$% confidence set for $\theta$ by the test inversion: \[ CS=\left\{ \theta\in\Theta:Y_{i}\leq g(X_{i},\theta+\hat{h}_{n}^{-}/n)\text{ for some }i\text{ and }Y_{i}\geq g(X_{i},\theta)\text{ for all }i\right\} . \]

Simulation

In this section, we investigate the finite-sample performance of the proposed test through a simulation study. We consider one-sided testing problems: $H_{0}^{-}:h\leq0$ against $H_{1}^{-}:h>0$ and $H_{0}^{+}:h\geq0$ against $H_{1}^{+}:h<0$. The proposed test, $\phi_{n}^{-}(Y^{n})$ and $\phi_{n}^{+}(Y^{n})$, are compared with the Wald test based on the maximum likelihood estimator and the asymptotic distribution derived by chernozhukov2004likelihood, hereafter referred to as the CH test.

In the CH test, we reject $H_{0}^{-}$ if $n(\hat{\theta}-\theta_{0})>q_{1-\alpha}(Z^{\theta})$, and reject $H_{0}^{+}$ if $n(\hat{\theta}-\theta_{0})<0$, where $\hat{\theta}$ denotes the maximum likelihood estimator. In our benchmark setup in ((ref)), under $P_{\theta+h/n}$, this estimator satisfies $n(\hat{\theta}-(\theta+h/n))\rightsquigarrow Z^{\theta}$, where $q_{1-\alpha}(Z^{\theta})$ is the $(1-\alpha)$-th quantile of $Z^{\theta}$, and \[ Z^{\theta}=\frac{\mathrm{Exp}(1)}{f(g(\theta)|\theta)\nabla_{\theta}g(\theta)}. \]

As the data generating process, we consider a truncated normal distribution $N(\theta,1)$ restricted to the range $[\theta-1.25,\infty)$, i.e., the benchmark model in ((ref)) with $f(y|\theta)=\frac{\phi(y-\theta)}{1-\Phi(-1.25)}$ and $g(\theta)=\theta-1.25$. We set the significance level at $\alpha=0.05$ and consider two sample sizes $n\in\{20,40\}$. The number of Monte Carlo replications is set to 2000.

Figures (ref) and (ref) present the power curves of the proposed test and CH test, and the power envelope for each one-sided test. The power envelopes, derived in the proofs of Theorems (ref) and (ref) in Appendix, represent the asymptotically optimal values.\footnote{The power envelope is given by $E_{h}[\phi^{-}(W)]$ for the range $h\geq0$ and $E_{h}[\phi^{+}(W)]$ for the range $h<0$ using the notation introduced in Appendix.}

Figure (ref) shows that the proposed test exhibits reasonable power across a range of values for the true parameter value $h>0$, performing closer to the power envelope, while the CH test demonstrates lower power. In contrast, Figure (ref) shows that the CH test exhibits over-rejection under the null $h=0$, with size $13.0\%$ for $n=20$ and $6.2\%$ for $n=40$. The proposed test controls sizes effectively, with values of $4.4\%$ for $n=20$ and $4.3\%$ for $n=40$. These simulation results suggest that the proposed test offers better finite-sample performance.

figure[figure omitted — 385 chars of source]
figure[figure omitted — 385 chars of source]