Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
116,673 characters · 20 sections · 80 citation commands
Locally regular and efficient tests in non-regular semiparametric models
\thispagestyle{empty}
\onehalfspacing \setcounter{page}{1}
It is often considered desirable that estimators are “locally regular” in that they exhibit the same limiting behaviour under the true parameter as they do under sequences of “local alternatives” which cannot be consistently distinguished from the true parameter.\footnote{Precise definitions will be given below. See BKRW98, vdV98, for example, for textbook treatments.} Unfortunately, there are many models in which locally regular estimators do not exist.\footnote{See e.g. C86, C92, N90, RB90 for some examples.} One necessary condition is given by C86: if the efficient information for a scalar parameter is 0, then no locally regular estimator of that parameter exists. Similarly, singularity of the efficient information matrix implies the non-existence of locally regular estimators of Euclidean parameters. Models in which this may occur are called “non-regular”. Many widely used models are non-regular (at least at certain parameter values): examples include single index models, instrumental variables, errors-in-variables, mixed proportional hazards, discrete choice models and sample selection models.
In this paper, I demonstrate that locally regular tests exist in a broad class of non-regular models, despite the non-existence of locally regular estimators. In particular, I show that a class of tests based on the C$(\alpha)$ idea of N59, N79 are locally regular.
These tests are based on a quadratic form of moment conditions evaluated under the null hypothesis. The key [C($\alpha$)] idea which ensures the local regularity is that the moment conditions must be (asymptotically) orthogonal to the collection of score functions for all nuisance parameters. Such moment conditions can always be constructed from any initial moment conditions by an orthogonal projection.
A key advantage of these C($\alpha$) tests is that they do not (asymptotically) overreject under (semiparametric) weak identification asymptotics, i.e. under local alternatives to a point of identification failure.\footnote{The semiparametric weak identification asymptotics used are those of Kaji21 (see also AM22), suitably generalised to permit non-i.i.d models. } The local regularity of these tests ensures that if the test is asymptotically of level $\alpha$ under any fixed parameter consistent with the null, it is also asymptotically of level $\alpha$ under any sequence of local alternatives consistent with the null, i.e. under (semiparametric) weak identification asymptotics. In addition to the well-studied case where weak identification stems from potential identification failure due to a finite dimensional nuisance parameter, the results in this paper also cover the case where identification failure is due to an infinite dimensional nuisance parameter and thus provide a generally applicable approach to weak identification robust inference in semiparametric models.\footnote{ These C($\alpha$) tests also behave well in other non-standard settings, such as when nuisance functions are estimated under shape constraints; see Section (ref) for a discussion. } Even in the case where the identification failure due to a finite dimensional nuisance parameter, the resulting weak identification robust tests appear to be new in the literature.\footnote{For instance, in the case of homoskedastic linear IV, the test that results from the construction in this paper does not coincide with any of the “usual” weak instrument robust tests (e.g. AR, LM, K, CLR). Demonstration of this is available from the author.} The tests proposed here are derived directly from an asymptotic orthogonality condition. As such they are close in spirit to the identification robust test of K05 which also requires an orthogonalisation, albeit with respect to different objects and in a different Hilbert space.
Achieving local regularity does not come at the expense of (local asymptotic) power. I characterise power bounds for tests in non-regular models and show that the C($\alpha$) tests proposed in this paper acheive these power bounds provided the moment conditions are chosen optimally. These power bounds contain those for regular models as a special case. Moreover, the conditions required for attainment of the power bounds are weaker than those in the literature.\footnote{In particular, in regular models the attainment result is well known if either (a) the observations are i.i.d. vdV98 or (b) the information operator (as defined in CHS96) is boundedly invertible CHS96. The result in this paper does not require either of these conditions.}
Following the theoretical development, I give details of its application to two examples: (i) a single index model which may be weakly identified when the link function is too flat and (ii) an instrumental variables (IV) model which may be weakly identified when the (nonparametric) first stage is too close to a constant function. Simulation experiments based on these examples demonstrate that the proposed tests enjoy good finite sample performance.
The application to IV may also be of interest for empirical researchers concerned about weak instruments. If the instruments are mean independent of the errors, then the test proposed here is robust to weak identification and can be substantially more powerful than tests assuming a linear first stage. This imposes no cost if the true first stage is (approximately) linear: the power of the proposed test is comparable to optimal tests based on a linear first stage. The practical use of these tests is demonstrated in two IV applications with possibly weak instruments.
This paper is connected to three main strands of the literature: the first is that concerned with general results on estimation and testing in semiparametric models. Much of this is now textbook material: see e.g. N90, CHS96, BKRW98, vdV98. The second is the literature on C($\alpha$) tests. These were introduced by N59, N79 and have seen many useful applications, most recently as a way to handle machine learning or otherwise high dimensional first steps CHS15,BEvK20, CEINR22. In this paper, the same structure which ensures good performance in such settings is used for a different purpose -- to construct tests which remain robust in non-regular settings. Lastly, the literature on robust testing in non -- regular or otherwise non -- standard settings is closely related to this paper AG09,RS12,EMW15, McC17. In particular, the locally regular tests derived in this paper are especially useful in cases of weak identification and therefore this paper is closely related to the literature on weak identification robust inference SS97, D97, SW00, K05, AC12, AM15, AM16b. More specifically, this paper is most closely related to the recent work on semiparametric weak identification Kaji21, AM22 and extends the notion of semiparametric weak identification considered therein to non -- i.i.d. models.\footnote{Failure of local identification and singularity of the information matrix are closely linked in parametric models, see R71. In the semiparametric case, parameters may be identified but nevertheless have a singular efficient information matrix. The relationship between the efficient information matrix and identification is considered by E22.}
The goal considered throughout this paper is to construct hypothesis tests of $\mathrm{H}_0: \theta = \theta_0$ against $\mathrm{H}_1: \theta \neq \theta_0$ in the sequence of models $\mathcal{P}_n = \{P_{n, \gamma}: \gamma\in \Gamma\}$ where $\gamma = (\theta, \eta)\in \Gamma = \Theta\times\mathcal{H}$ for some open $\Theta\subset \mathbb{R}^{d_\theta}$ and $\mathcal{H}$ an arbitrary set. Each $\mathcal{P}_n$ consists of probability measures on a measurable space $(\mathcal{W}_n, \mathcal{B}(\mathcal{W}_n))$ and is dominated by a $\sigma$-finite measure $\nu_n$.\footnote{Typically the index $n$ is sample size and $\mathcal{W}_n$ is the space in which a sample of size $n$ takes its values. This is the situation considered in Section (ref) as well as in the examples in Section (ref). }
Let $H_{\gamma}=\mathbb{R}^{d_\theta}\times B_{\gamma}$ be a subset of a linear space containing 0, and suppose that $\{P_{n, \gamma, h}: h\in H_{\gamma}\}\subset \mathcal{P}_n$ are such that $P_{n, \gamma} = P_{n, \gamma, 0}$. Elements of $H_\gamma$ will be written as $h = (\tau, b) \in \mathbb{R}^{d_\theta}\times B_{\gamma}$.\footnote{In most examples, $H_\gamma$ will be a linear space. The more general situation as considered here is nevertheless important to allow for, for example, Euclidean nuisance parameters subject to boundary constraints. In such a setting, if the constraint is binding at $\gamma$, then $\gamma$ can only be perturbed in certain directions if $P_{n, \gamma, h}$ is to remain within the model.} The measures $P_{n, \gamma, h}$ should be viewed as local perturbations of the measure $P_{n, \gamma}$ in a “direction” $h\in H_\gamma$. These local perturbations can be split in two groups: the perturbations $P_{n, \gamma, h}$ with $h\in H_{\gamma, 0}\coloneqq \{(0, b): b\in B_{\gamma}\}$ correspond to the null hypothesis $\mathrm{H}_0: \theta = \theta_0$ and those with $h\in H_{\gamma, 1}\coloneqq \{h = (\tau, b): 0 \neq \tau\in \mathbb{R}^{d_\theta}, b\in B_{\gamma}\}$ to the alternative $\mathrm{H}_1: \theta \neq \theta_0$. As such, $P_{n, \gamma, h}$ for $h\in H_{\gamma, 0}$ will be referred to as local perturbations consistent with the null hypothesis, whilst $P_{n, \gamma, h}$ for $h\in H_{\gamma, 1}$ are local alternatives. The subsequent analysis is local with the parameter $\gamma$ being considered fixed at a $\gamma$ consistent with $\mathrm{H}_0$. As such, to lighten the notation, dependence on $\gamma$ will be mostly left implicit: I write $P_{n, h}$ for $P_{n, \gamma, h}$, $H$ for $H_\gamma$, $H_{i}$ for $H_{\gamma, i}$ ($i=0, 1$) and similarly for other objects. I also use the abbreviation $P_n\coloneqq P_{n, 0}$.
I use the single-index model as a running example throughout the paper.\footnote{Technical details for this example are deferred to Sections (ref) and (ref).}
The key technical condition under which the theory is developed is local asymptotic normality (LAN; see e.g. vdV98 or LCY00). Define the log-likelihood ratios
The requirement that $\Delta_{n}h$ converges in $d_2$ is equivalent to requiring that it converges weakly and $(\Delta_{n}h)_{n\in \mathbb{N}}$ is uniformly square $P_{n}$-integrable BKRW98. This implies that $\sigma(h)= \lim_{n\to\infty} \|\Delta_{n}h\|^2$.
That is, the finite sample (local) power function of the test, $\uppi_n$ converges under each $P_{n, h}$ to a function $\uppi$ which may depend on $\tau$ (and, implicitly, $\gamma$) but not on $b$, the parameter which describes local deviations from the nuisance parameter $\eta$.\footnote{Cf. the definition of a (locally) regular estimator in vdV98.} If a sequence of tests does not satisfy (ref) it is (locally) non-regular.
Local regularity of test sequences as in (ref) is a pointwise concept. It is also of interest to consider a uniform version.
If $H$ is a (pseudo-)metric space and $K$ is a compact set, for the convergence in (ref) to hold uniformly on $K$ it is necessary and sufficient to show that the sequence of functions $\uppi_n$ is asymptotically equicontinuous on $K$.\footnote{ The same is true if $K$ is totally bounded. See e.g. D21, p. 123, for the definition of asymptotic equicontinuity.}
Directly working with the power functions $\uppi_n$ to show their asymptotic equicontinuity is complicated in many cases. It is, however, often possible to show results which imply this property. For instance, the functions $h\mapsto P_{n, h}$ being asymptotically equicontinuous in $d_{TV}$ implies the required asymptotic equicontinuity of the power functions. Despite being (much) stronger, this often holds.\footnote{ See the discussion following Remark (ref) below. }
\paragraph{Weak identification asymptotics and local regularity}
In many models there are parameter values, $\gamma$, at which locally regular estimators do not exist. Points where the parameter of interest, $\theta$, is un- or under-identified provide an important class of examples. Moreover, as is well known from the literature on weak identification, even if $\theta$ is identified at $\gamma$, finite sample inference may be poor if $\gamma$ is too close to a point of identification failure relative to the amount of information contained in the sample. Such behaviour has been widely studied in models where the part of $\gamma$ causing the identification failure is finite dimensional AC12, AM15.
There are also many examples where weak identification may occur due to the value of infinite-dimensional nuisance parameters. Kaji21 and AM22 use a differentiability in quadratic mean (DQM) condition to define semiparametric weak identification asymptotics in i.i.d. models. In particular, they consider sequences $P_{n, h}^n$ which satisfy
for a point $P_0$ where the parameter of interest is unidentified. In the i.i.d. case, (ref) implies the LAN expansion in Assumption (ref) with $\Delta_nh = \frac{1}{\sqrt{n}}\sum_{i=1}^n f(W_i)$ vdV98.\footnote{ If Assumption (ref) holds with $\Delta_n$ having this form, the converse is also true. } Working with Assumption (ref) in place of (ref) broadens the applicability of this class of semiparametric weak identification asymptotics to non-i.i.d. models. It is clear from Definition (ref) that a locally regular test sequence will have asymptotic null rejection probability (NRP) which does not exceed the nominal level under weak identification asymptotics $P_{n, h}$.\footnote{Of course, a (non-regular) test sequence may have asymptotic NRP which depends on $b$ and yet is bounded by the nominal level under $P_{n, h}$ for all $h = (0, b)\in H_0$ and / or have an asymptotic power function which depends on $b$ for $h = (\tau, b)\in H_1$. Restricting attention to locally regular test sequences may be justified by the power optimality results of Section (ref).}
I now give two examples of semiparametric models where the parameter of interest $\theta$ may be un- or under-identified depending on the value of an infinite dimensional nuisance parameter.\footnote{A further example is the linear simultaneous equations model in LM21.} The first is the running example.
In Examples (ref) and (ref), at the points of identification failure, no locally regular estimator exists, however locally regular C($\alpha$) tests are developed in Section (ref).\footnote{These examples consider i.i.d. data for simplicity. See HLM22 for an example of a locally regular C($\alpha$) test of the form proposed in this paper for the potentially un- / under-identified parameter in a structural vector autoregressive model.}
To construct locally regular tests of $\mathrm{H}_0: \theta = \theta_0$ against $\mathrm{H}_1: \theta\neq \theta_0$, I use a generalisation of the class of C($\alpha$) tests introduced by N59, N79 to characterise optimal tests in regular parametric models. These tests are a based on a quadratic form of (estimators of) a vector of $d_\theta$ moment conditions $g_n\in L_2(P_n)$ which satisfy the following requirements.
Built-in to Assumption (ref) is a requirement of asymptotic orthogonality of $g_n$ and the scores for the nuisance parameters $\eta$. This generalises the analogous condition in N59, N79 and is key to the local regularity of C$(\alpha)$ tests.
Given any $d_\theta$ moment conditions $f_{n}\in L_2^0(P_{n})$, moment conditions which satisfy an exact version of the orthogonality condition (ref) may be obtained as
An important special case of this construction is with $f_{n}$ the score function for $\theta$, i.e. $f_{n} = \dot{\ell}_{n}$ such that $\tau^\prime \dot{\ell}_{n}= \Delta_{n}(\tau, 0)$ for each $\tau\in \mathbb{R}^{d_\theta}$. The function
is called the efficient score function. This yields a power optimal choice of moment conditions satisfying (ref) as shown in Section (ref) below.
To construct the test statistic, I assume that consistent estimators of $g_{n}$, $V^\dagger$ (the Moore-Penrose pseudo-inverse of $V$) and $r\coloneqq \operatorname{rank}(V)$ are available, given $\theta$.
Verification of Assumption (ref)(ref) typically proceeds by model specific arguments. That $g_n$ is (asymptotically) orthogonal to $\{\Delta_n(0, b):b\in B\}$ often helps in establishing this consistency (cf. CEINR22). One generally applicable approach to obtain an estimator which satisfies Assumption (ref)(ref) is to take an initial estimator which is consistent for $V$, threshold its eigenvalues at an appropriate rate and then take the pseudo-inverse.\footnote{See Section S5 of LM21-S for full details of this approach. Other regularisation schemes are also possible (see e.g. DV15).} If one uses the estimator $\hat{\Lambda}_{n, \theta} \coloneqq \hat{V}_{n, \theta}^\dagger$ where $\hat{V}_{n, \theta}\xrightarrow{P_{n}}V$ and $\hat{r}_{n, \theta}\coloneqq \operatorname{rank}(\hat{V}_{n, \theta})$ then condition (ref) holds if and only if condition (ref) holds A87. Nevertheless, as emphasised by the notation, it is not necessary that the estimate $\hat{\Lambda}_{n, \theta}$ be the pseudo-inverse of an initial estimate.
Given the estimators of Assumption (ref), the C($\alpha$)-style test statistic is
The C($\alpha$) -- style test $\psi_{n, \theta_0}$ of $\mathrm{H}_0$ against $\mathrm{H}_1$ at level $\alpha$ is:
where $c_n$ is the $1-\alpha$ quantile of a $\chi^2_{\hat{r}_n}$ random variable.
\paragraph{Local regularity} Assumptions (ref) -- (ref) suffice for local regularity of $\psi_{n, \theta_0}$.
Theorem (ref) immediately shows that $\psi_{n, \theta_0}$ is locally regular (cf. (ref)). The asymptotic orthogonality in (ref) is key to this result. If, instead, $\lim_{n\to\infty} \left\langle {\Delta_n h}\, ,\, {g_{n}^\prime} \right\rangle = \tau^\prime \Sigma_{21}^\prime + c(b)$ with $c(b)\neq 0$, then (by Le Cam's third Lemma) the limiting distribution of $g_n$ under $P_{n, h}$ would be $\mathcal{N}(\Sigma_{21}\tau + c(b), V)$ and hence the limiting power function of the test sequence would not be free of $b$.
\paragraph{Uniform local regularity}
The local regularity given by (ref) may be “upgraded” to local uniform regularity (Definition (ref)) under various conditions. Here I consider the case where $H$ posseses a (pseudo-)metric structure (e.g. if $H$ is a subset of a (semi-)normed linear space).\footnote{If $H$ posseses a (finite) measure structure and the functions $h =(\tau, b)\mapsto \uppi_{n}(\tau, b)$ are measurable then $\psi_{n, \theta_0}$ is locally uniformly regular except on a “small” subset of $H$ by Egorov's Theorem. See Section (ref) for details.} In this case, for $\psi_{n, \theta_0}$ to be locally uniformly regular on a compact (or totally bounded) $K\subset H$ it is necessary and sufficient that the functions $h =(\tau, b)\mapsto \uppi_{n}(\tau, b)$ are asymptotically equicontinuous.
I now give a sufficient condition for the asymptotic equicontinuity required by Corollary (ref).
In the parametric i.i.d. case LAN is often verified by establishing a DQM condition, e.g. equation (7.1) in vdV98. This is sufficient for the ULAN expansion in Assumption (ref) to hold vdV98. Semiparametric generalisations of this result are available (e.g. combine Proposition (ref) and Lemma (ref)).\footnote{LM21 and HLM22 verify this asymptotic equicontinuity property in i.i.d. and time series semiparametric examples respectively.}
The condition in Lemma (ref) is natural given its link with the ULAN condition. Neverthelesss, it is (much) stronger than necessary for the condition required by Corollary (ref); see Lemma (ref) for a weaker sufficient condition.
The preceding section established the local regularity of the tests $\psi_{n, \theta_0}$ based on (estimates of) moment functions $g_{n}$ satisfying certain asymptotic orthogonality conditions. Thus far, nothing has been said about the choice of $g_{n}$ beyond these orthogonality requirements. The choice of the functions $g_{n}$ determines the power of the corresponding test. As such, they ought to be chosen such that the resulting test has good power against alternatives of interest.
One natural choice is the efficient score function (ref). It is well known that tests based on the efficient score function have certain optimality properties in regular models when (a) the observations are i.i.d. vdV98 or (b) when the information operator for $\eta$ is boundedly invertible CHS96. I show that this optimality persists in non-regular models and does not require (a) or (b).
The results in this section are derived using the limits of experiments framework of Le Cam LC86, vdV98. In particular, I show that the local experiments consisting of the measures $P_{n, h}$ for $h\in H$ converge weakly to a limit experiment which has a close relationship to a Gaussian shift experiment on the Hilbert space formed by taking the quotient of $H$ under the seminorm induced by the variance function $\sigma(h)$. The connection between these experiments is sufficiently tight that power bounds derived in the latter transfer to the former.\footnote{That the local experiments do not converge to the mentioned Gaussian shift experiment is essentially a purely technical point: the Gaussian shift experiment is defined on a different parameter space to the local experiments, whilst (weak) convergence of experiments (in the sense of LC86) is defined for experiments with the same parameter space.}
\paragraph{The limit experiment} For this development $H$ is required to be linear and I will therefore assume that $B$ (hence $H$) is a linear space. Under LAN, there exists a positive semi-definite symmetric bilinear form $\left\langle {\cdot}\, ,\, {\cdot} \right\rangle_{K}$ on $H= \mathbb{R}^{d_\theta} \times B$ such that $\sigma(h) = \left\langle {h}\, ,\, {h} \right\rangle_{K}$. This can be seen as a by-product of the following Lemma.
For $h, g\in H$, setting $\left\langle {h}\, ,\, {g} \right\rangle_{K} \coloneqq K(h, g)$ gives a positive semi-definite symmetric bilinear form. Let $\|\cdot\|_{K}$ denote the seminorm induced by $\left\langle {\cdot}\, ,\, {\cdot} \right\rangle_{K}$ on $H$.
Define $\mathbb{H}$ as the quotient of $H$ by the subspace on which $\|\cdot\|_{K}$ vanishes:
which is an inner product space when equipped with the natural inner product induced by $\left\langle {\cdot}\, ,\, {\cdot} \right\rangle_{K}$, which I also denote by $\left\langle {\cdot}\, ,\, {\cdot} \right\rangle_{K}$. An element of $\mathbb{H}$ corresponding to representative element $h\in H$ will be denoted by $[h]$.\footnote{ Analogous comments apply to the related space $\mathbb{H}_1$, defined below. In both cases, to avoid an excess of parentheses / brackets, if $h = (\tau, b)$ I will write either $[h]$ or $[\tau, b]$, rather than $[(\tau, b)]$. }
The (weak) limit of the sequence of experiments consisting of the measures $P_{n, h}$ can be obtained by standard results on weak convergence of experiments.
Under the assumption that $\mathbb{H}$ is separable, the experiment $\mathcal{E}$ is equivalent to a Gaussian shift on $(\mathbb{H}, \left\langle {\cdot}\, ,\, {\cdot} \right\rangle_{K})$, in the sense given by Proposition (ref) below.
\paragraph{The efficient information matrix}
Power bounds for tests of $\mathrm{K}_0: h\in H_0$ against $\mathrm{K}_1: h\notin H_1$ can be expressed in terms of the efficient information matrix, $\tilde{\mathcal{I}}_{\gamma}$, so named because in the i.i.d. setting it is the covariance matrix of the efficient score function for a single observation. Here I provide an alternative definition of this matrix which applies more generally, and reduces to the classical definition in the i.i.d. case (as shown in Lemma (ref)).
Let $\|\tau\| \coloneqq \inf_{b\in B}\|(\tau, b)\|_{K}$, which defines a semi-norm on $\mathbb{R}^{d_\theta}$. Equipping the quotient $\mathbb{H}_1 \coloneqq {\mathbb{R}^{d_\theta}} \,/\, {\{\tau \in \mathbb{R}^{d_\theta} :\|\tau\| = 0\}}$ with the natural norm induced by $\|\cdot\|$ (which I also denote by $\|\cdot\|$) turns it into a normed space. Define the linear map $\pi_1:\mathbb{H} \to\mathbb{H}_1$ as $\pi_1([\tau, b]) \coloneqq [\tau]$. As $\pi_1$ is continuous it may be uniquely extended to a continuous function defined on $\overline{\mathbb{H}}$, the completion of $\mathbb{H}$; this extension will also be called $\pi_1$. Since $\pi_1$ is continuous, $\ker \pi_1\subset \overline{\mathbb{H}} $ is closed. Let $\Pi$ be the orthogonal projection onto $\ker \pi_1$ and define $\Pi^\perp\coloneqq I - \Pi$, the orthogonal projection onto $[\ker \pi_1]^\perp$. Let $e_i$ be the $i$-th canonical basis vector in $\mathbb{R}^{d_\theta}$ and define the efficient information matrix $\tilde{\mathcal{I}}_{\gamma}$ as the $d_\theta \times d_\theta$ matrix with $i,j$-th entry $\tilde{\mathcal{I}}_{ij}$ given by\footnote{Lemma (ref) gives an alternative expression for $\tilde{\mathcal{I}}_{\gamma}$ based on the Gaussian process $\Delta$ of Lemma (ref).}
The following Theorem records the power bound for (locally asymptotically) unbiased two-sided tests of a scalar $\theta$. In this case the matrix $\tilde{\mathcal{I}}_{\gamma}$ has rank either 0 or 1. Theorem (ref) handles both cases simultaneously.
Theorem (ref) implies that the power bound of Theorem (ref) is achieved by the test $\psi_{n, \theta_0}$ provided $\Sigma_{21}V^\dagger \Sigma_{21}^\prime = \tilde{\mathcal{I}}_{\gamma}$ and $r=1$.
When $d_\theta >1$ there is an intermediate case where $0<\operatorname{rank}(\tilde{\mathcal{I}}_{\gamma})<d_\theta$. Here I permit $0<\operatorname{rank}(\tilde{\mathcal{I}}_{\gamma})\le d_\theta$ and establish a maximin power bound for (potentially) non -- regular models, which contains the (regular) full rank case as a special case.\footnote{Section (ref) shows that the most stringent test (in the sense of W43) in the limit experiment has the same power function as the maximin test, and no sequence of asymptotically level $\alpha$ tests can correspond to a test in the limit experiment with smaller regret.}
By Theorem (ref), the power bound on the right hand side of (ref) is achieved by $\psi_{n, \theta_0}$ provided $\Sigma_{21}V^\dagger \Sigma_{21}^\prime = \tilde{\mathcal{I}}_{\gamma}$ and $\operatorname{rank}(V) = \operatorname{rank}(\tilde{\mathcal{I}}_{\gamma}) = r\ge1$. In order that the test be asymptotically maximin over a compact subset $K_a$ of $\{h = (\tau, b)\in H: \tau^\prime \tilde{\mathcal{I}}_{\gamma} \tau \ge a\}$, with $a = \inf\{\tau^\prime \tilde{\mathcal{I}}_{\gamma} \tau = a: h\in K_a\}$, some uniformity is required.\footnote{ The pseudometric $d$ in Corollary (ref) need not be related to the seminorm $\|\cdot\|_K$.}
A sufficient condition for the asymptotic equicontinuity required for the second part of Corollary (ref) based on an asymptotic equicontinuity in total variation requirement was given as Lemma (ref) in the previous section.\footnote{As noted there Lemma (ref) provides some weaker sufficient conditions; in the present context, condition (ref) of Lemma (ref) is not required (cf. Remark (ref)).}
If the efficient information matrix $\tilde{\mathcal{I}}_{\gamma}$ is zero, no test with correct asymptotic size has non -- trivial asymptotic power against any sequence of local alternatives.
There are a number of important aspects to highlight regarding the interpretation of the power bounds obtained in the preceding subsections.
\paragraph{Optimality in multivariate testing problems}
Just as in the classical finite -- dimensional case, the multivariate optimality results just presented should not be taken in an absolute sense. Nevertheless they seem reasonable if the researcher does not have directions against which they wish to direct power a priori. If there are alternatives of particular interest, one could construct a locally regular test by utilising the same moment conditions $g_{n}$ but weighting them differently BRS06.
\paragraph{The intermediate case with $1\le\operatorname{rank}(\tilde{\mathcal{I}}_{\gamma})<d_\theta$}
A key benefit of the multivariate power results is that they apply equally to non-regular models, i.e. cases where $\tilde{\mathcal{I}}_{\gamma}$ is rank deficient. This scenario can occur for various reasons. Firstly the model may not identify all parameters of interest $\theta$ (i.e. underidentification). Secondly some of the elements of $\theta$ may be weakly identified (i.e. weak underidentification). The power results above apply in either of these cases.
There are a number of other papers which provide inference results in similarly rank deficient settings RCBR00, HM19, AG19, ABS23; none of these papers consider optimal testing.
\paragraph{Alternative approximations}
In the case where $\operatorname{rank}(\tilde{\mathcal{I}}_{\gamma}) = 0$, Proposition (ref) reveals that the LAN approximation in Assumption (ref) is, in a certain sense, the wrong approximation: it does not provide any useful way of (asymptotically) comparing tests. Other approximations might provide valuable comparisons. Alternative approximations have been explored in, for example, the IV model M09 and semiparametric GMM models by AM22, AM23. For example, in the IV case M09 considers alternatives which are at a fixed distance from the true parameter, rather than in a shrinking $\sqrt{n}$-neighbourhood. Whether such an approach can be developed for the class of models considered here is an interesting question for future work.
Provided that the $L_2$ distance between $g_{n}$ and $\tilde{\ell}_{n}$ (as defined in (ref)) vanishes, $\psi_{n, \theta_0}$ attains the power bounds established in the preceding subsections. For regular models this result is well known in two special cases: (a) the i.i.d. case (cf. Section 25.6 in vdV98; Lemma (ref)) and (b) when the information operator $\mathsf{B}$ in Remark (ref) is positive-definite with $\mathsf{B}_{22}$, the information operator for $\eta$, boundedly invertible CHS96. Here I provide a general version of this result which applies to both regular and non-regular models and does not require (a) or (b). In particular, I show that $\Sigma_{21}V^\dagger \Sigma_{21}^\prime = \tilde{\mathcal{I}}_{\gamma}$, which suffices given Theorem (ref) and the power bounds in Theorems (ref) and (ref).
In this section I give conditions which are sufficient for some of the foregoing Assumptions in the in the benchmark case for semiparametric theory: where the observations are i.i.d. and the model is “smooth”.
In the i.i.d. setting, it is well known that quadratic mean differentiability of the square root of the density $p = \frac{\mathrm{d}P}{\mathrm{d}\nu}$ of $P \coloneqq P_0$ is sufficient for LAN. In particular, if
for a measurable $Ah: \mathcal{W}\to \mathbb{R}$, then with $\Delta_{n}h \coloneqq \frac{1}{\sqrt{n}} \sum_{i=1}^n [Ah](W_i)$ the remainder term $R_{n}$ in the LAN expansion satisfies $R_{n}(h_n)\xrightarrow{P}0$ vdVW96. This can be used to establish either the LAN expansion required by Assumption (ref) by taking $h_n = h$ for each $n\in\mathbb{N}$ or the ULAN expansion as in Assumption (ref) by considering sequences $h_n\to h$. Sufficient conditions for (ref) are well known vdV98.
In this case, the scores $Ah$ typically take the form
where $\dot{\ell}_{\gamma}$ is a $d_\theta$-vector of functions in $L_2^0(P)$ (typically the partial derivatives of $\theta \mapsto \log p_{\gamma}$ at $\gamma$) and $D:\overline{\operatorname{lin}}\ B \to L_2^0(P)$ a bounded linear map. Showing that (ref) holds (with $h_n =h$) is typically the most straightforward way to verify the LAN expansion required by Assumption (ref). If $A: \overline{\operatorname{lin}}\ H \to L_2(P)$ is a bounded linear map, then the remainder of Assumption (ref) also follows directly.\footnote{A version of Lemma (ref) for ULAN (Assumption (ref)) is Lemma (ref) in the supplementary material.}
When the data are i.i.d., the the joint convergence of $(\Delta_{n}h, g_{n}^\prime)$ as in Assumption (ref) is particularly straightforward to verify. As noted in the discussion around (ref), the required orthogonality condition can be ensured by performing an orthogonal projection. Assumption (ref) then follows straightforwardly. In the i.i.d. setting typically $g_{n}$ will have the form $g_{n}(W^{(n)}) = \mathbb{G}_n g$.
I now illustrate the application of the theoretical results to the single index and IV models and conduct simulation studies to investigate finite sample performance of the proposed approach. In this section I work under high level conditions to avoid repeating standard regularity conditions; lower level sufficient conditions are given in section (ref) of the supplementary material.
Consider the single index model of Example (ref). I now formalise the development given in Section (ref). The model parameters are $\gamma = (\theta, \eta)$ where $\eta= (f, \zeta)$ and the density of one observation with respect to a $\sigma$-finite measure $\tilde{\nu}$ is $p_{\gamma}$ as in (ref); $P_{\gamma}$ denotes the corresponding probability measure. The parameters $\gamma$ are restricted by the following Asssumption. Let $\mathscr{X}$ be the support of $X$, $\mathscr{D}$ a convex open set containing $\{x_1 + x_2^\prime \theta: \theta\in \Theta,x\in \mathscr{X}\}$ and $C_b^1(\mathscr{D})$ the class of real functions which are bounded and continuously differentiable with bounded derivative on $\mathscr{D}$.
That $p_{\gamma}$ is a valid probability density holds automatically (with $\tilde{\nu} = \nu$) when $\epsilon|X$ is continuously distributed.
\paragraph*{Local Asymptotic Normality} Consider local perturbations $P_{\gamma + \varphi_n(h)}$ for
$B_{1}$ is the set which indexes the perturbations to $f$ and consists of a subset of the continuously differentiable functions $b_1:\mathscr{D}\to\mathbb{R}$. $B_{2}$ indexes the perturbations to $\zeta$ and consists of a subset of the functions $b_2:\mathbb{R}^{1+K}\to \mathbb{R}$ which are continuously differentiable in their first argument and satisfy
The precise form of $\varphi_{n, 2}$ is left unspecified. It is required only that the local perturbations satisfy the LAN property below.\footnote{Examples of $\varphi_{n, 2}$ and $B$ for which Assumption (ref) holds are given in Section (ref). }
\paragraph*{The moment conditions} The test statistic is based on $g_{n}\coloneqq \mathbb{G}_n g$ for $g$ given in (ref). This satisfies Assumption (ref) under Assumptions (ref) & (ref) and (ref) below.
\paragraph*{A feasible test}
To form a feasible test $g_n$ must be replaced by an estimator $\hat{g}_{n, \theta}$. Let this have the form $\hat{g}_{n, \theta}(W^{(n)}) \coloneqq \frac{1}{\sqrt{n}}\sum_{i=1}^n \hat{g}_{n, \theta, i}$, for $\hat{g}_{n, \theta, i}$ defined as in (ref). To keep the notation concise let $Z_{3}\coloneqq f$, $Z_4\coloneqq f'$, $Z_0 \coloneqq Z_{1} / Z_2$ and correspondingly $\hat{Z}_{0,n, i}\coloneqq \hat{Z}_{1, n, i} / \hat{Z}_{2, n, i}$. Let $\check{V}_{n, \theta}\coloneqq \frac{1}{n}\sum_{i=1}^n \hat{g}_{n, \theta, i}\hat{g}_{n, \theta, i}^\prime$ and If $V$ is known to have full rank then let $\hat{V}_{n, \theta} \coloneqq \check{V}_{n, \theta}$, $\hat{\Lambda}_{n, \theta} \coloneqq \hat{V}_{n, \theta}^{-1}$ and $\hat{r}_{n, \theta} = \operatorname{rank}(V)$. Form the estimator $\hat{V}_{n, \theta}$ according to the construction in Section S5 of LM21-S using a truncation rate $\upnu_n$. $\hat{\Lambda}_{n, \theta}$ is then taken to be $\hat{V}_{n, \theta}^\dagger$ and $\hat{r}_{n, \theta}\coloneqq \operatorname{rank}(\hat{V}_{n, \theta})$. Under the following condition, these estimators satisfy the conditions of Assumption (ref).
The rate conditions in Assumption (ref) can be satisfied by e.g. (sample -- split) series estimators under standard conditions; see e.g. BCCK15.
A consequence of Assumption (ref) and Propositions (ref) and (ref) is that the test $\psi_{n, \theta_0}$ formed as in (ref) is locally regular by Theorem (ref).
\paragraph*{Simulation study}
I take $K=1$ and test $H_0: \theta=\theta_0 = 1$ at a nominal level of 5%. Each study reports the results of 5000 monte carlo replications with a sample size of $n\in \{400, 600, 800\}$. I report empirical rejection frequencies for the $\psi_{n, \theta_0}$ test along with a Wald test based on an estimator in the style of I93.
I consider two different classes of link function. The first sets $f(v) =f_j(v) = 5\exp (-v^2 / 2c_j^2)$ (“exponential”); the second $f(v) =f_j(v) = 25\left(1 + \exp(-v /c_j)\right)^{-1}$ (“logistic”). The values of $c_j$ considered are recorded in Table (ref).
In each case, as $c_j$ increases, the derivative of $f$ flattens out, moving towards a point with $f'=0$, at which $\theta$ is unidentified.\footnote{The functions $f$ and $f'$ are plotted in Figures (ref) and (ref).} I draw covariates as $X= (Z_1, 0.2 Z_1 + 0.4 Z_2 + 0.8)$, where each $Z_k\sim U(-1, 1)$ is independent. The error term is drawn either as $\epsilon = \upsilon / \sqrt{3/2}$ with $\upsilon \sim t(6)$ (“homoskedastic”) or $\epsilon \sim \mathcal{N}(0, 1 + \sin(X_1)^2)$ (“heteroskedastic”).
I compute the test $\psi_{n, \theta_0}$ as described on p. (ref), with $\upomega(X) = 1$. The functions $f,\, f'$ and $Z_{1}$ are estimated via sample split smoothing splines.\footnote{I use the base R function smooth.spline with 20 knots. In this setting $Z_2(V_\theta) = 1$ is known.} The truncation parameter $\upnu$ is set to $10^{-3}$. I additionally compute a Wald test in the style of I93, using the same non-parametric estimates as for $\hat{g}_{n, \theta}$.\footnote{Given $\hat{f}$, $\hat\theta = \operatorname*{arg\,min}_{\theta\in \Theta_\star} \frac{1}{n}\sum_{i=1}^n (Y_i - \hat{f}(V_{\theta, i}))^2$, for $\Theta_\star = [-10, 10]$. The asymptotic variance is estimated by $\hat{\sigma}^2 / \frac{1}{n}\sum_{i=1}^n \left(\widehat{f'}(V_{\hat\theta, i}) \left[X_2 - \hat{Z}(V_{\hat\theta, i})\right]\right)^2$, for $\hat{\sigma}^2 = \frac{1}{n}\sum_{i=1}^n (Y_i - \hat{f}(V_{\hat\theta, i}))^2$.}
The empirical rejection frequencies of these procedures are recorded in Table (ref): $\psi_{n, \theta_0}$ rejects at close to the nominal 5% for all simulation designs considered whilst the Wald test over -- rejects in most of the simulation designs considered. Figure (ref) contains power plots of the $\psi_{n, \theta_0}$ test ($n=800$). For almost flat link functions there is very identifying information and hence very little power available. As the link function moves away from the point of identification failure ($f'=0$), the available power increases and is captured by the $\psi_{n, \theta_0}$ test.
In the IV model of Example (ref), $n$ i.i.d. copies of $W = (Y, X, Z)$ are observed where
Let $d_W\coloneqq d_{\theta} + d_Z + 1$. With $\pi(Z)\coloneqq \mathbb{E}[X|Z]$ and $\upsilon = X - \pi(Z)$,
If $\pi(Z)$ is constant the instruments $Z$ provide no information about $\theta$. Weak identification in this model can be very different from in the IV model with a linear first stage: there are many data configurations in which $\mathrm{Var}(\mathbb{E}[X|Z])$ may be “large” whilst $\mathbb{E}[XZ^\prime]\mathbb{E}[ZZ^\prime]^{-1}\approx 0$. In such situations, tests which can exploit such non-linear identifying information can provide substantially more power than tests which (implicitly) use a linear first stage. In this section I develop a $\psi_{n, \theta_0}$ test which can capture such identifying information whilst remaining robust to weak identification.\footnote{This does not contradict optimality results that are known for, e.g., the AR test M09, CHJ09 as these results assume a linear first stage. }$^,$ \footnote{An alternative approach to capturing this non-linear identifying information (whilst remaining robust to weak instruments) is to use a large number of transformations of the instruments, $f_1(Z), \ldots, f_M(Z)$, in a linear first stage, combined with a testing procedure which remains robust in the presence of many weak instruments. In the simulation study below, I compare the $\psi_{n, \theta_0}$ test to this approach, using the test of MS21.}
Let $\zeta$ denote the density of $\xi \coloneqq (\epsilon, \upsilon^\prime, Z^\prime)$ with respect to a $\sigma$-finite measure $\nu$. The parameters of the IV model are $\gamma = (\theta, \eta)$ with the nuisance parameters collected in $\eta = (\beta, \pi, \zeta)$. The density of one observation is
with respect to a $\sigma$-finite measure $\tilde{\nu}$ and $P_{\gamma}$ denotes the corresponding measure. The model parameters are restricted as follows.
Assumption (ref) imposes the existence of certain moments and the (IV) conditional mean restriction. That $p_{\gamma}$ is a valid probability density holds automatically (with $\nu = \tilde{\nu}$) when $U|Z$ is continuously distributed.
\paragraph*{Local Asymptotic Normality} Consider local perturbations $P_{\gamma + \varphi_n(h)}$ for
with $ B \coloneqq \mathbb{R}^{d_\beta} \times B_{1}\times B_{2}$. $B_{1}$ is a subset of the bounded functions $b_1:\mathbb{R}^{d_Z}\to \mathbb{R}^{d_\theta}$ and $B_{2}$ a subset of the functions $b_2:\mathbb{R}^{d_W}\to \mathbb{R}$ which are bounded and continuously differentiable in its first $1 + d_\theta$ components with bounded derivative and such that
The precise forms of $\varphi_{n, 1}, \varphi_{n, 2}$ are left unspecified. It is required only that the local perturbations satisfy LAN.\footnote{Examples of $\varphi_{n, 1}, \varphi_{n, 2}$ and $B$ for which Assumption (ref) holds are given in Section (ref). }
\paragraph*{The moment conditions} The test will be based on moment conditions related to the efficient score function for $\theta$, $\tilde{\ell}_{\gamma}$. This is given in the following Lemma.\footnote{The last two conditions in (ref) hold if $\lim_{|u_i|\to\infty} |u_i|\zeta(u, z)=0$ for $i=1, \ldots, d_\alpha$. }
For simplicity, I will use the moment functions
$g$ belongs to the orthocomplement of $\{Db: b\in B\}$ and coincides with the efficient score function when $J(Z) = \mathbb{E}[UU^\prime]$ a.s. (i.e. under homoskedasticity).\footnote{Nevertheless, homoskedasticity is not assumed and the results below hold under heteroskedasticity. For full efficiency one could base the test on (ref). This is left for future work. }
\paragraph*{A feasible test}
Suppose that $\hat{\beta}_n$ and $\hat{\pi}_{n, i}(Z_i)$ are estimators of $\beta$ and $\pi(Z_i)$ respectively. Let the $i$-th residual in (ref) based on $\theta = \theta_0$ and $\hat\beta_n$ be $\hat{\epsilon}_{n, i} \coloneqq Y_i - X_i^\prime \theta - Z_{1, i}^\prime\hat{\beta}_n$. Let $\hat{s}_n\coloneqq \frac{1}{n}\sum_{i=1}^n \hat{\epsilon}_{n, i}^2$ and define
and
Based on $\check{V}_{n, \theta}$, form $\hat{V}_{n, \theta}$ according to the construction in Section S5 of LM21-S using a truncation rate $\upnu_n$, set $\hat{\Lambda}_{n, \theta}\coloneqq \hat{V}_{n, \theta}^\dagger$ and $\hat{r}_{n, \theta}\coloneqq \operatorname{rank}(\hat{V}_{n, \theta})$.
The following assumption provides sufficient high-level conditions on the estimators $\hat\beta_n$ and $\hat{\pi}_{n, i}(Z_i)$ such that Assumption (ref) holds. These conditions are compatible with $\hat{\pi}_{n, i}$ being a leave-one-out series estimator.\footnote{The discretisation of $\hat\beta_n$ is a technical device which permits the proof to go through under weaker conditions LCY00. This can be arranged given a $\sqrt{n}$ -- consistent initial estimator, by replacing its value with the closest point in the set $\mathscr{S}_n$.}$^,\,$\footnote{See e.g. BCCK15 for sufficient conditions for (ref) and Section (ref) for a discussion of (ref).}
There is no requirement on the rate $\delta_n$ in (ref), (ref) beyond $\delta_n = o(1)$.
A consequence of Assumption (ref) and Propositions (ref) and (ref) is that the test $\psi_{n, \theta_0}$ formed as in (ref) is locally regular by Theorem (ref).
\paragraph{Simulation study} I test $H_0: \theta=\theta_0 = 0$ at a nominal level of 5%. Each study reports the results of 5000 monte carlo replications with a sample size of $n\in \{200, 400, 600\}$.\footnote{The power surfaces in Design 1 are computed with 2500 replications.} Two simulation designs are considered.
Design 1 is a bivariate, just identified design. Here $d_\theta = 2$ and $Z_2$ is drawn from a zero-mean multivariate normal distribution with covariance matrix $\left[
\right]$. The error terms $\epsilon, \upsilon$ are drawn from a zero-mean multivariate normal such that each has variance 1 and the covariances are $\mathrm{Cov}(\epsilon, \upsilon_i) = 0.9$ and $\mathrm{Cov}(\upsilon_1,\upsilon_2) = 0.7$. $Z_1=1$ with $\beta = 1$ and $\pi(Z) = \pi(Z_2) = (\pi_1(Z_{2,1}), \pi_2(Z_{2, 2}))^\prime$ with each $\pi_i$ ($i=1, 2$) being one of the exponential or logistic functions $f_j$ in Table \ref{tbl:SIM-index-fcns}. The exponential form is a prototypical function shape for which the linear projection of $X$ on $Z$ will provide essentially no identifying information; for the logistic form this linear projection should perform well.\footnote{These functions are plotted in Figures \ref{fig:SIM-gaussian-fs} and \ref{fig:SIM-logistic-fs}. The separation $\pi(Z_2) = (\pi_1(Z_{2,1}), \pi_2(Z_{2, 2}))^\prime$ is assumed unknown and is not imposed in the estimation of $\pi$. }$^,\, $\footnote{Results for the case where $\pi(Z_2)$ is linear are very similar to the “approximately linear” logistic case and are therefore unreported.}
I consider the $\psi_{n, \theta_0}$ test developed above, with a leave-one-out series estimator of $\pi$ based on (tensor product) Legendre polynomials. I consider both fixing the number of polynomials at $k=3$ in each of the univariate series which form the tensor product basis and choosing $k\in \{3, 4, 5, 6, 7\}$ using information criteria. $\upnu$ is set to $10^{-2}$. I additionally consider the AR49 (AR) test.\footnote{The AR test is computed with $Z_2$ as instruments, after partialling out $Z_1$. }$^,$\footnote{I do not consider alternative weak instrument robust tests based on a linear first stage (e.g. LM, CLR) in this design as the AR test is known to be optimal when the model is just-identified.}
The empirical rejection frequencies under the null are shown in Tables (ref) and (ref). The parameter $j$ controls the level of identification: the larger is $j$ the closer $\pi_j$ is to a constant function. In each specification all the considered tests reject at close to the nominal level. Power surfaces for the $\psi_{n, \theta_0}$ and AR tests are shown in figures (ref) -- (ref). As can be seen in these figures, the $\psi_{n, \theta_0}$ test is able to detect deviations from the null when $\pi_j$ has the exponential form, unlike the AR test.\footnote{ One could consider an AR test using e.g. some basis functions $f_1(Z_{2}), \ldots, f_K(Z_{2})$ however as noted in MS21, the AR statistic is not well behaved for large $K$. The jackknife AR test of MS21 applies only to the case where $d_\theta=1$. } For the logistic form, the power of the two tests is similar. Unsurprisingly, neither test provides non-trivial power when identification is very weak.
Design 2 is a univariate, over identified model with heteroskedastic errors. $Z_1$, $\beta$ and $Z_2$ are as in Design 1, and $\pi(Z_2) = (\pi_1(Z_{2, 1}) + \pi_2(Z_{2, 2}))/2$ where the $\pi_i$ have one of the exponential or logistic forms of Table (ref).\footnote{This functional form is treated as unknown and not imposed in the estimation of $\pi$.} I draw $(\tilde{\epsilon}, \tilde{\upsilon})$ from a zero-mean multivariate normal distribution with unit variances, covariance $0.95$ and set $(\epsilon, \upsilon)^\prime = \left[
\right](\tilde{\epsilon}, \tilde{\upsilon})^\prime$.
The $\psi$ tests are computed in the same manner as in Design 1. I also compute the AR, LM and CLR tests based on $Z_2$ (with $Z_1$ partialled out) as well as the many weak instrument robust jackknife AR test of MS21. MS$_1$ uses $Z_2$ as instruments; MS$_2$ uses the (tensor product) of Legendre polynomials used to estimate $\pi$ as instruments.
The empirical rejection frequencies under the null are shown in Tables (ref) -- (ref). As in Design 1, the parameter $j$ controls the level of identification: the larger is $j$ the closer $\pi_j$ is to a constant function and hence $\theta$ unidentified. In each specification all the considered tests reject close to the nominal level; the MS$_2$ test is somewhat oversized for smaller $n$. The power of these tests is plotted in Figures (ref) -- (ref); the $\psi_{n, \theta_0}$ tests are denoted by $k=3$, AIC and BIC, corresponding to how $\pi$ is estimated. For the design with both $\pi_i$ exponential, the $\psi_{n, \theta_0}$ test clearly delivers the highest power whenever there is non-trivial power available; of the other tests, only MS$_2$ delivers non-trivial power in this specification. For the case with both $\pi_i$ logistic, all tests except MS$_2$ perform similarly, with MS$_2$ offering lower power. The same holds for the final specification, where $\pi_1$ is exponential and $\pi_2$ logistic.
In this section I re-analyse two IV studies with potentially weak instruments by inverting the $\psi_{n, \theta_0}$ test developed in section (ref) to construct weak-instrument robust confidence intervals (CIs). In each case the $\psi_{n, \theta_0}$ test is able to exploit non-linearities to yield substantial reductions in CI length relative to AR CIs.
H14 studies the long term effect of skilled immigration on productivity using a natural experiment in which the skilled but religiously persecuted French protestants (Hugenots) fled and settled in Prussia.\footnote{That the Hugenots were more skilled (on average) than the Prussian population seems to be broadly accepted cf. pp. 85-86, 93-95 in H14.} In the notation of Example (ref), $Y$ is log output in textile manufacturing, $X$ is the proportion of Hugenots in each town and $Z_1$ contains various control variables, see H14 for details. H14 argues that “by the order of centralized ruling by the king and his agents Huguenots were channeled into Prussian towns in order to compensate for severe population losses during the Thirty Years' War”, motivating the instrument $Z_2$: the percentage population losses during the war. In particular, three different measurements of this population loss are used and refered to as specifications (1), (2) and (3) hereafter.\footnote{Specifications (1) & (2) are those considered in the left and and right hand parts of Table 4 of H14; Specification (3) is that considered in the left hand part of Table 5.}
As noted in H14, this instrument may be weak: in each case the first stage F statistic is “small”.\footnote{It is less than the cutoff of 10 suggested by SS97 for homoskedastic IV.} I implement the $\psi_{n, \theta_0}$ test using a leave-one-out series estimator of $\pi$ of the form
where the $i$ subscript on the estimated coefficients indicates they have been estimated on all observations except for the $i$-th. $p_{K}$ is a vector of a constant and the first $K$ Legendre polynomials. I choose $K\in \{1, 2, 3, 4\}$ and whether to include $Z_1$ in the model for $\pi$ by using BIC: all specifications include $Z_1$ and $K=4$.
Table (ref) reports 2SLS estimates of $\theta$ along with 2SLS (Wald) CIs, AR CIs and CIs found by inverting the $\psi_{n, \theta_0}$ test.\footnote{The inversion is performed over a grid of 5000 equally spaced points from -1 to 7.} The resulting CIs provide a similar interpretation as that based on the AR CIs: the effect of (skilled) Hugenot immigration was positive on textile output. However, the $\psi_{n, \theta_0}$ based CIs are smaller than the AR CIs, achieving approximately a 40% - 50% reduction in length.
A11 estimates the effect of racial segregation ($X$) on poverty and inequality ($Y$, measured respectively by the poverty rate and log gini coefficient for black / white city residents), instrumenting segregation by a “railroad division index” (RDI, $Z_2$), “a variation on a Herfindahl index that measures the dispersion of a city's land into subunits” via the layout of railroad tracks.\footnote{Section 3 and Appendix A of A11 provides evidence that the choice of railroad placement was not related to local social or economic concerns.} The first and second stages also include an intercept and control for railroad track length ($Z_1$).
The instrument may be weak: I calculate the first stage F statistic to be 2.307.\footnote{A11 refers to Column 1 of Table 1 when discussing the first stage F statistic. The values in this table imply a first stage F of $(0.357 / 0.088)^2 \approx 16.458$. This “discrepancy” arises from different default choices of robust covariance estimate in R's sandwich package (HC3; my calculation) and STATA's robust command (HC1; A11).} I implement the $\psi_{n, \theta_0}$ test using a leave-one-out series estimator of $\pi$ of the form (ref). I choose $K\in \{1, 2, 3, 4\}$ and whether to include $Z_1$ in the model for $\pi$ by using BIC: this excludes $Z_1$ and chooses $K=2$.
Table (ref) reports 2SLS estimates of $\theta$ along with 2SLS (Wald) CIs, AR CIs and CIs found by inverting the $\psi_{n, \theta_0}$ test.\footnote{The inversion is performed over a grid of 5000 equally spaced points from -1 to 2.} The resulting CIs provide a similar interpretation as that based on the AR CIs: racial segregation increases poverty and inequality within the Black community and decreases poverty and inequality within the White community. The $\psi_{n, \theta_0}$ based CIs are shorter than the AR CIs, achieving a reduction in length varying from 5% to around 38%.
In this paper I establish that C($\alpha$)-style tests are locally regular under mild conditions, including in non-regular cases where locally regular estimators do not exist. As a consequence, these tests do not overreject under semiparametric weak identification asymptotics. Additionally I generalise the classical local asymptotic power bounds for LAN models to the case where the efficient information matrix has positive, but potentially deficient, rank, such that these results also apply in cases of underidentification (or weak underidentification). I show that, if the C($\alpha$) test is based on the efficient score function, it attains these power bounds. This (attainment) result improves on results known in the literature in two ways: (i) it applies also to non-regular models and (ii) it does not require the data to be i.i.d. nor the information operator to be boundedly invertible. A simulation study based on two examples shows that the asymptotic theory provides an accurate approximation to the finite sample performance of the proposed tests.