Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
37,768 characters · 8 sections · 32 citation commands
Optimal testing in a class of nonregular models
\address{Department of Economics, University of Wisconsin, Madison, 1180 Observatory Drive, Madison, WI 53706-1393, USA.} \email{[email removed]} \address{Department of Economics, London School of Economics, Houghton Street, London, WC2A 2AE, UK.} \email{[email removed]}
This paper studies optimal hypothesis testing of a class of nonregular econometric models in which the boundary of the support of the observed data depends on some parameter of interest. Such nonregular models, which typically imply discontinuous likelihood functions and nonstandard convergence rates of estimators, have been often studied in the econometrics literature; see, flinn1982new, smith1985maximum, christensen1991exact, donald1993maximum, hong1998maximum, donald2002superconsistent, hirano2003asymptotic, and chernozhukov2004likelihood, among others. In contrast to most existing papers that focus on point estimation, this paper is concerned with optimal (composite) hypothesis testing for nonregular models instead of point estimation.
For testing a simple null hypothesis against a simple alternative one, Neyman-Pearson's fundamental lemma yields an optimal power property of the likelihood ratio test even for the case of parameter-dependent support. However, the optimality result is no longer available for general testing problems with composite hypotheses. On the other hand, for regular statistical models, it is known that standard testing methods (such as the likelihood ratio, Wald, and score tests) can achieve certain asymptotic optimal power properties for testing general composite hypotheses (see, Chapter 15 of lehmann2022testing). An open question is whether we can establish an analogous asymptotic optimality result for testing composite hypotheses on nonregular parameters in the case of parameter-dependent support, and this paper addresses this question in a positive way.
In this paper, we consider one-sided and two-sided hypothesis testing for parametric models with parameter-dependent support and develop an asymptotically uniformly most powerful (AUMP) test based on a limit experiment. Interestingly, our two-sided test attains the AUMP property without imposing further restrictions such as unbiasedness, and can be inverted to construct a confidence set for the nonregular parameter. For clarity we first present the main results under a benchmark setup in Section (ref), where there is no covariate or nuisance parameter. Then we extend our optimality results to a general setup in Section (ref) that involves covariates and nuisance parameters. For the general case, we need some independent auxiliary sample to estimate the nuisance parameters, which is typically obtained by sample splitting. Our simulation results in Section (ref) illustrate desirable finite sample properties of the proposed tests.
The most closely related papers to ours are hirano2003asymptotic and chernozhukov2004likelihood. hirano2003asymptotic studied efficient point estimation of parameter-dependent support models by extending the limit of experiments argument, and showed that the Bayes estimator is asymptotically efficient under a minimax criterion but the maximum likelihood estimator is generally inefficient. To the best of our knowledge, this paper is the first one that studies optimal hypothesis testing for nonregular models with parameter-dependent support. In contrast to the Bayes estimator in hirano2003asymptotic that involves priors on parameters, our optimal testing methods are developed based on the limiting likelihood ratio process without priors. chernozhukov2004likelihood also investigated nonstandard asymptotic properties of estimation and testing methods for parameter-dependent support models. They established asymptotic optimality of the Bayes estimators in terms of the asymptotic average risk, and showed that the Wald test and Bayes posterior quantiles are valid for inference. However, they did not discuss optimality of the testing methods. In addition to these papers, chen2018monte (chen2018monte, Appendix C) examined partially identified models with parameter-dependent support, and chen2025identification adapted their quasi-Bayesian likelihood ratio statistics to partially identified auction models. In our simulation study, we demonstrate that the proposed test exhibits more reasonable performance than the Wald type test by chernozhukov2004likelihood.
This section presents our main results for a simple nonregular model with a scalar parameter and no covariate. This model, covering the uniform distribution as a canonical example, provides a useful benchmark to highlight our developments. We propose one-sided tests for this model and derive the asymptotic properties of these tests, including the AUMP. Based on the one-sided tests, we construct an optimal two-sided test with an unequal-tailed way.
Let $Y\in\mathcal{Y}\subset\mathbb{R}$ follow the parametric model
where $\mathbb{I}\{\cdot\}$ is the indicator function, $\theta\in\Theta\subset\mathbb{R}$ is the true value of a scalar parameter, and conditions on the functions $f$ and $g$ are specified below. For this model, we consider the one-sided testing problem
for a given $\theta_{0}\in\Theta$. Hereafter, we focus on the case of $\nabla_{\theta}g(\theta_{0})>0$. The case of $\nabla_{\theta}g(\theta_{0})<0$ is analyzed in the same manner. Based on the reparametrization $h=n(\theta-\theta_{0})$, this testing problem is written as
We wish to test $H_{0}^{-}$ or $H_{0}^{+}$ based on an independent and identically distributed (iid) sample $Y^{n}=(Y_{1},\ldots,Y_{n})$ of $Y$. Although our results can be easily extended to cover a more general reparametrization $h=n(\theta-\theta_{0})+\zeta_{n}$ with $\zeta_{n}=o(1)$, we focus on the case of $\zeta_{n}=0$ for simplicity.
Our testing procedure is constructed based on the limiting likelihood ratio process. Since the joint density of $Y^{n}$ is written as \[ \mathrm{d}P_{\theta}^{n}(y^{n})=\mathbb{I}\{y_{(1)}\ge g(\theta)\}\prod_{i=1}^{n}f(y_{i}|\theta), \] where $y_{(1)}=\min\{y_{1},\ldots,y_{n}\}$, the likelihood ratio process on a parameter space $\mathcal{H}\subset\mathbb{R}$ is \[ Z_{n}(h,\bar{h}):=\frac{\mathrm{d}P_{\theta_{0}+\frac{\bar{h}}{n}}^{n}}{\mathrm{d}P_{\theta_{0}+\frac{h}{n}}^{n}}(Y^{n})=\frac{\mathbb{I}\{Y_{(1)}\ge g(\theta_{0}+\bar{h}/n)\}}{\mathbb{I}\{Y_{(1)}\ge g(\theta_{0}+h/n)\}}\prod_{i=1}^{n}\frac{f(Y_{i}|\theta_{0}+\bar{h}/n)}{f(Y_{i}|\theta_{0}+h/n)}, \] for $h,\bar{h}\in\mathcal{H}$. To characterize asymptotic properties of this process, we impose the following assumptions.
Assumption (ref) (i) is standard, and Assumption (ref) (ii) contains smoothness and boundedness conditions on the functions $f$ and $g$ in ((ref)). By adapting hirano2003asymptotic (hirano2003asymptotic, Theorem 2) to our setup, Assumption (ref) guarantees weak convergence (denoted by “$\rightsquigarrow$”) of the likelihood ratio process
for every finite $I\subset\mathcal{H}$, where $\lambda=\{f(g(\theta_{0})|\theta_{0})\nabla_{\theta}g(\theta_{0})\}^{-1}$ and \[ D_{h,\bar{h}}:=\mathbb{I}\{W_{h}>\bar{h}\}\text{ with }W_{h}\sim f_{W}(w|h)=\frac{1}{\lambda}e^{-(w-h)/\lambda}\mathbb{I}\{w>h\}. \] Note that $\lambda$ is a known constant in the present setup. In contrast to the standard likelihood ratio process, which is locally asymptotically normal, $Z_{n}(h,\bar{h})$ converges to a limit of experiments whose randomness is given by the binary variable $D_{h,\bar{h}}$. This is due to lack of differentiability in quadratic mean of the density $\mathrm{d}P_{\theta}^{n}(y^{n})$. Since $D_{h,\bar{h}}$ is discrete, the limiting likelihood ratio process $Z(h,\bar{h})$ is discontinuous in the sense that $\Pr\{Z(h,\bar{h})\leq z\}$ is not continuous at $z=0$ and $e^{(\bar{h}-h)/\lambda}$. Thus, we cannot use the conventional asymptotic theory based on the quadratic expansion of the likelihood ratio process to evaluate asymptotic size and power properties of the likelihood ratio test.
Even though the likelihood ratio process $\{Z_{n}(h,\bar{h})\}_{\bar{h}\in I}$ exhibits such nonregularity, it should be noted the limiting likelihood ratio process $\{Z(h,\bar{h})\}_{\bar{h}\in I}$ satisfies the monotone likelihood ratio property, which is defined as follows.
A key feature of distributions with monotone likelihood ratio is the existence of the uniformly most powerful (UMP) test. Relying on the limiting likelihood ratio process is essential since the finite sample likelihood ratio does not exhibit a monotone likelihood ratio property in general.
Here $\phi(U_{h})=1$ and $0$ mean rejection and acceptance of $H_{0}$, respectively, and $\phi(U_{h})=\kappa$ means rejection with probability $\kappa$. Therefore, to derive an AUMP test for $H_{0}^{-}:h\leq0$, we can still invoke the asymptotic representation lemma below to argue that the sample counterpart of the monotone likelihood ratio test with $S(W_{h})=W_{h}$ is AUMP.
To derive an AUMP test for the other one-sided test $H_{0}^{+}:h\geq0$, we utilize the fact that the limit experiment can be represented as a function of the random variable $W_{h}$. Accordingly, we apply the Neyman-Pearson lemma (e.g., Theorem 3.2.1 (ii) of lehmann2022testing) to $W_{h}$.
Hereafter, we first formalize this argument for one-sided tests on $H_{0}^{-}:h\leq0$ and $H_{0}^{+}:h\geq0$ (Section (ref)), and then extend the argument to two-sided testing for $H_{0}:h=0$ (Section (ref)).
First, we consider testing $H_{0}^{-}:h\leq0$ against $H_{1}^{-}:h>0$ at the significance level $\alpha\in(0,1)$. By taking a sample counterpart of $\phi(W_{h})$ in ((ref)) with the limiting process $W_{h}$ in ((ref)), our test is constructed as
This test achieves an asymptotic optimal property in Definition $\text{\ref{def:AUMP}}$ below, introduced by choi1996asymptotically.
The following asymptotic optimality result is established using the monotone likelihood ratio property of the limit experiment, as demonstrated in Appendix.
Next, we consider another one-sided testing $H_{0}^{+}:h\geq0$ against $H_{1}^{+}:h<0$. The basic idea is same as the previous case. We propose the following test:
and the asymptotic optimality of this test is obtained as follows by applying the Neyman-Pearson lemma to $W_{h}$.
Since we assume $\nabla_{\theta}g(\theta_{0})>0$ and $\lambda>0$, we have $g(\theta_{0})<g\left(\theta_{0}+\frac{\lambda}{n}\log\left(\frac{1}{1-\alpha}\right)\right)$ eventually. However, this inequality can be violated in finite samples. We can show that any test $\psi_{n}^{+}$ rejecting the null with probability one if $Y_{(1)}<g(\theta_{0})$ and rejecting the null with probability at most $\alpha$ if $Y_{(1)}\geq g(\theta_{0})$ is AUMP at level $\alpha$ for testing $H_{0}^{+}:\theta\geq\theta_{0}$ against $H_{1}^{+}:\theta<\theta_{0}$. We recommend using ($\text{\ref{eq:phi+}}$) since it avoids randomization.
This subsection considers two-sided testing $H_{0}:h=0$ against $H_{1}:h\neq0$. By combining the optimal one-sided tests derived in the last subsection, we propose the following (unequal-tailed) two-sided test:
Indeed, this test is shown to be AUMP.
Importantly, the proposed test achieves the AUMP property without imposing additional restrictions such as unbiasedness. This feature is shared by a finite sample UMP two-sided test for the uniform distribution (e.g., lehmann2022testing, lehmann2022testing, Problem 3.2, p.105).
Since the two-sided test $\phi_{n}(Y^{n})$ does not randomize, we can easily construct a $100(1-\alpha)$% confidence set by the test inversion: \[ CS=\left\{ \theta\in\Theta:g(\theta)\leq Y_{(1)}\leq g\left(\theta+\frac{\lambda}{n}\log\left(\frac{1}{\alpha}\right)\right)\right\} . \] This confidence set also has a pointwise optimal property (asymptotically uniformly most accurate).
In this section, as in hirano2003asymptotic, we generalize the benchmark model to accommodate discrete covariates and nuisance parameters:
where $\theta\in\Theta\subset\mathbb{R}$ is a scalar parameter of interest, $\gamma\in\mathbb{R}^{d}$ is a $d$-dimensional vector of (regular) nuisance parameters, $Y$ is a scalar dependent variable, and $X$ is an $m$-dimensional vector of discrete covariates with support $\mathcal{X}=\{a_{1},\dots,a_{L}\}$. This section considers the one-sided testing problem
for a given $\theta_{0}\in\mathbb{R}$ with the asymptotic level of significance $\alpha$. By reparametrization $h=n(\theta-\theta_{0})$, this testing problem is written as
Let $(Y^{n},X^{n})=((Y_{1},\ldots,Y_{n}),(X_{1},\ldots,X_{n}))$ be an iid sample of $(Y,X)\in\mathbb{R}\times\mathcal{X}$. To extend our benchmark results in the last section, we consider the plug-in likelihood ratio process
for $h\in\mathcal{H}$ with a parameter space $\mathcal{H}\subset\mathbb{R}$, where $\hat{\gamma}_{n}$ and $\hat{h}_{n}$ are some estimators of the nuisance parameters $\gamma$ and the user-specified alternative value $\bar{h}$, respectively, based on auxiliary data independent from the main sample $(Y^{n},X^{n})$. In contrast to the benchmark case, where the threshold for testing takes the form $g(\theta_{0}+\bar{h}/n)$, the optimal values of $\bar{h}$ for our test depend on certain population quantities that must be estimated (see Theorems (ref) and (ref) below). Thus we introduce an estimator $\hat{h}_{n}$ for $\bar{h}$ in this general case. Typically we split the sample (say, $(Y^{2n},X^{2n})$) into the main sample $(Y^{n},X^{n})$ and auxiliary one $((Y_{n+1},\ldots,Y_{2n}),(X_{n+1},\ldots,X_{2n}))$ to obtain $\hat{\gamma}_{n}$ and $\hat{h}_{n}$.
In this section, we impose the following assumptions.
As specified in Assumption (ref) (i), we focus on the case where covariates are discrete variables as in hirano2003asymptotic. Assumption (ref) (ii) requires that the nuisance parameters $\gamma$ and the alternative value $\bar{h}$ to construct our test statistics below can be estimated at the $\sqrt{n}$ rate. For example, based on an auxiliary sample independent from $(Y^{n},X^{n})$, $\gamma$ can be estimated by the maximum likelihood and $\bar{h}$ can be estimated by the method of moments. This assumption enables us to focus on a neighborhood to establish the weak convergence in Lemma $\text{\ref{lem:weak_conv_nui}}$ below. Since $\hat{h}_{n}$ and $\hat{\gamma}_{n}$ are independent from the main sample $(Y^{n},X^{n})$, we can argue the weak convergence in a straightforward way even after conditioning on the concentration to the neighborhood. Assumption (ref) (iii) lists boundedness and smoothness conditions on the functions $f$ and $g$.
To present the limiting distribution of the plug-in likelihood ratio process $\{Z_{n}(h,\hat{h}_{n},\hat{\gamma}_{n})\}$, we introduce further notations. Define
$G_{j}=\nabla_{\theta}g(a_{j},\theta_{0})$, and $\lambda_{j}=\{\Pr(X=a_{j})f(g(a_{j},\theta_{0})|a_{j},\beta_{0})\}^{-1}$. Also let $(W_{h,1},\cdots,W_{h,L})$ be mutually independent random variables that follow \[ W_{h,j}\sim f_{W_{j}}(w\mid h,\gamma)=e^{-(w-G_{j}h)/\lambda_{j}}\mathbb{I}\{w>G_{j}h\}/\lambda_{j}. \] The weak convergence of the plug-in likelihood ratio process is established as follows.
This lemma is different from the weak convergence in ($\text{\ref{eq:Z}}$) for the benchmark case in the following aspects. First, the process $\{Z_{n}(h,\hat{h}_{n},\hat{\gamma}_{n})\}$ contains the estimated parameters $\hat{h}_{n}$ and $\hat{\gamma}_{n}$, which also covers the case of deterministic parameter sequences. Second, due to presence of the discrete covariates $X\in\{a_{1},\ldots,a_{L}\}$, the limiting process $Z(h,\bar{h})$ involves an $L$-dimensional random vector $(W_{h,1},\cdots,W_{h,L})$. Third, the limiting process $Z(h,\bar{h})$ depends on the nuisance parameters $\gamma$.
Hereafter, we separately consider testing $H_{0}^{-}:h\leq0,\gamma\in\mathbb{R}^{d}$ against $H_{1}^{-}:h>0,\gamma\in\mathbb{R}^{d}$ and $H_{0}^{+}:h\geq0,\gamma\in\mathbb{R}^{d}$ against $H_{1}^{+}:h<0,\gamma\in\mathbb{R}^{d}$. Let $\hat{\lambda}_{nj}$ and $\hat{\lambda}_{n}$ be consistent estimators of $\lambda_{j}$ and $\lambda$, respectively. For example, based on an auxiliary sample $\{X_{n+1},\ldots,X_{2n}\}$, $\lambda_{j}$ can be estimated by $\hat{\lambda}_{nj}=\{n^{-1}\sum_{i=n+1}^{2n}\mathbb{I}\{X_{i}=a_{j}\}f(g(a_{j},\theta_{0})|a_{j},\theta_{0},\hat{\gamma}_{n})\}^{-1}$, and then $\lambda$ can be estimated by $\hat{\lambda}_{n}=(\sum_{j=1}^{L}G_{j}/\hat{\lambda}_{nj})^{-1}$.
First, we consider one-sided testing $H_{0}^{-}:h\leq0,\gamma\in\mathbb{R}^{d}$ against $H_{1}^{-}:h>0,\gamma\in\mathbb{R}^{d}$ at the significance level $\alpha\in(0,1)$. By extending the one-sided test in ((ref)), our test is defined as
where $\hat{h}_{n}^{-}=\hat{\lambda}_{n}\log\left(\frac{1}{\alpha}\right)$. The idea to construct this test is essentially the same as the benchmark case in Section (ref). The main difference is that $\lambda$ is unknown and needs to be estimated by plugging-in the consistent estimator $\hat{\lambda}_{n}$ that is independent from the data. The next theorem shows that this test achieves an asymptotic optimal property.
Next, we consider another one-sided testing $H_{0}^{+}:h\ge0,\gamma\in\mathbb{R}^{d}$ against $H_{1}^{+}:h<0,\gamma\in\mathbb{R}^{d}$. In this case, our test is defined as
where $\hat{h}_{n}^{+}=\hat{\lambda}_{n}\log\left(\frac{1}{1-\alpha}\right)$. Similar to Theorem (ref), asymptotic optimality of this test is obtained as follows.
This subsection considers two-sided testing $H_{0}:h=0,\gamma\in\mathbb{R}^{d}$ against $H_{1}:h\neq0,\gamma\in\mathbb{R}^{d}$. By combining the optimal one-sided tests derived in the last subsection, we propose the following (unequal-tailed) two-sided test:
Based on this two-sided test, we can construct an asymptotically optimal $100(1-\alpha)$% confidence set for $\theta$ by the test inversion: \[ CS=\left\{ \theta\in\Theta:Y_{i}\leq g(X_{i},\theta+\hat{h}_{n}^{-}/n)\text{ for some }i\text{ and }Y_{i}\geq g(X_{i},\theta)\text{ for all }i\right\} . \]
In this section, we investigate the finite-sample performance of the proposed test through a simulation study. We consider one-sided testing problems: $H_{0}^{-}:h\leq0$ against $H_{1}^{-}:h>0$ and $H_{0}^{+}:h\geq0$ against $H_{1}^{+}:h<0$. The proposed test, $\phi_{n}^{-}(Y^{n})$ and $\phi_{n}^{+}(Y^{n})$, are compared with the Wald test based on the maximum likelihood estimator and the asymptotic distribution derived by chernozhukov2004likelihood, hereafter referred to as the CH test.
In the CH test, we reject $H_{0}^{-}$ if $n(\hat{\theta}-\theta_{0})>q_{1-\alpha}(Z^{\theta})$, and reject $H_{0}^{+}$ if $n(\hat{\theta}-\theta_{0})<0$, where $\hat{\theta}$ denotes the maximum likelihood estimator. In our benchmark setup in ((ref)), under $P_{\theta+h/n}$, this estimator satisfies $n(\hat{\theta}-(\theta+h/n))\rightsquigarrow Z^{\theta}$, where $q_{1-\alpha}(Z^{\theta})$ is the $(1-\alpha)$-th quantile of $Z^{\theta}$, and \[ Z^{\theta}=\frac{\mathrm{Exp}(1)}{f(g(\theta)|\theta)\nabla_{\theta}g(\theta)}. \]
As the data generating process, we consider a truncated normal distribution $N(\theta,1)$ restricted to the range $[\theta-1.25,\infty)$, i.e., the benchmark model in ((ref)) with $f(y|\theta)=\frac{\phi(y-\theta)}{1-\Phi(-1.25)}$ and $g(\theta)=\theta-1.25$. We set the significance level at $\alpha=0.05$ and consider two sample sizes $n\in\{20,40\}$. The number of Monte Carlo replications is set to 2000.
Figures (ref) and (ref) present the power curves of the proposed test and CH test, and the power envelope for each one-sided test. The power envelopes, derived in the proofs of Theorems (ref) and (ref) in Appendix, represent the asymptotically optimal values.\footnote{The power envelope is given by $E_{h}[\phi^{-}(W)]$ for the range $h\geq0$ and $E_{h}[\phi^{+}(W)]$ for the range $h<0$ using the notation introduced in Appendix.}
Figure (ref) shows that the proposed test exhibits reasonable power across a range of values for the true parameter value $h>0$, performing closer to the power envelope, while the CH test demonstrates lower power. In contrast, Figure (ref) shows that the CH test exhibits over-rejection under the null $h=0$, with size $13.0\%$ for $n=20$ and $6.2\%$ for $n=40$. The proposed test controls sizes effectively, with values of $4.4\%$ for $n=20$ and $4.3\%$ for $n=40$. These simulation results suggest that the proposed test offers better finite-sample performance.