EconBase
← Back to paper

Locally Regular and Efficient Tests in Non-Regular Semiparametric Models

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

116,644 characters

Locally regular and efficient tests in non-regular semiparametric models




\title{\sc \Large
Locally regular and efficient tests in non-regular semiparametric models
}

\author{{Adam Lee\thanks{BI Norwegian Business School, [email removed].
Previous versions of this paper were titled ``Robust and Efficient Inference For Non-Regular Semiparametric Models''.
I have benefitted from discussions with and comments / questions from Majid Al-Sadoon,
 Isiah Andrews,
 Christian Brownlees, Bjarni G. Einarsson, Juan Carlos Escanciano, Kirill Evdokimov, Lukas Hoesch,
Geert Mesters,
Vladislav Morozov, Jonas Moss,
Whitney Newey,
Katerina Petrova, Francesco Ravazzolo, Barbara Rossi, Andr\'{e} B. M. Souza, Emil Aas Stoltenberg, Philipp Tiozzo and participants at various conferences and seminars.
All errors are my own.}}
	}
\date{\today
}



\maketitle
\thispagestyle{empty}

\begin{abstract}
\onehalfspacing
\noindent


This paper considers hypothesis testing in semiparametric models which may be non-regular.
I show that C($\alpha$) style tests are locally regular under mild conditions, including in cases where locally regular estimators do not exist, such as models which are (semiparametrically) weakly identified.
I characterise the appropriate limit experiment in which to study local (asymptotic) optimality of tests in the non-regular case and generalise classical power bounds to this case. I give conditions under which these power bounds are attained by the proposed C($\alpha$) style tests. The application of the theory to a single index model and an instrumental variables model is worked out in detail.








\noindent

\bigskip \noindent \textit{JEL classification}: C10, C12, C14, C21, C39

\bigskip \noindent \textit{Keywords}:
Hypothesis testing, local asymptotics, uniformity, semiparametric models, weak identification, boundary, regularisation, single-index,
instrumental variables.

\end{abstract}

\clearpage

\onehalfspacing
\setcounter{page}{1}
\section{Introduction}\label{sec:intro}

It is often considered desirable that estimators are ``locally regular'' in that they exhibit the same limiting behaviour under the true parameter as they do
 under sequences of ``local alternatives'' which cannot be consistently distinguished from the true parameter.\footnote{Precise definitions will be given below. See \cite{BKRW98, vdV98}, for example, for textbook treatments.}
Unfortunately, there are many
models in which locally regular estimators do not exist.\footnote{See e.g. \cite{C86, C92,  N90, RB90} for some examples.} One necessary condition is given by \cite{C86}: if the efficient information for a scalar parameter is 0, then no locally regular estimator of that parameter exists. Similarly, singularity of the efficient information matrix implies the non-existence of locally regular estimators of Euclidean parameters. Models in which this may occur are called ``non-regular''. Many widely used models are non-regular (at least at certain parameter values): examples include single index models, instrumental variables, errors-in-variables, mixed proportional hazards, discrete choice models and sample selection models.

In this paper, I demonstrate that locally regular \emph{tests} exist in a broad class of non-regular models, despite the non-existence of locally regular estimators. In particular, I show that a class of tests based on the C$(\alpha)$ idea of \cite{N59, N79} are locally regular.

These tests are based on a quadratic form of moment conditions evaluated under the null hypothesis. The key [C($\alpha$)] idea which ensures the local regularity is that the moment conditions must be (asymptotically) orthogonal to the collection of score functions for all nuisance parameters. Such moment conditions can always be constructed from any initial moment conditions by an orthogonal projection.

A key advantage of these C($\alpha$) tests is that they do not (asymptotically) overreject under (semiparametric) weak identification asymptotics, i.e. under local alternatives to a point of identification failure.\footnote{The semiparametric weak identification asymptotics used are those of \cite{Kaji21} (see also \citealp{AM22}), suitably generalised to permit non-i.i.d models.
} The local regularity of these tests ensures that if the test is asymptotically of level $\alpha$ under any fixed parameter consistent with the null, it is also asymptotically of level $\alpha$ under any sequence of local alternatives consistent with the null, i.e. under (semiparametric) weak identification asymptotics. In addition to the well-studied case where weak identification stems from potential identification failure due to a finite dimensional nuisance parameter, the results in this paper also cover the case where identification failure is due to an infinite dimensional nuisance parameter and thus provide a generally applicable approach to weak identification robust inference in semiparametric models.\footnote{
    These C($\alpha$) tests also behave well in other non-standard settings, such as when nuisance functions are estimated under shape constraints; see Section \ref{sm:shape-constraints} for a discussion.
} Even in the case where the identification failure due to a finite dimensional nuisance parameter, the resulting weak identification robust tests appear to be new in the literature.\footnote{For instance, in the case of homoskedastic linear IV, the test that results from the construction in this paper does not coincide with any of the ``usual'' weak instrument robust tests (e.g. AR, LM, K, CLR). Demonstration of this is available from the author.}
 The tests proposed here are derived directly from an asymptotic orthogonality condition. As such they are close in spirit to the identification robust test of \cite{K05} which also requires an orthogonalisation, albeit with respect to different objects and in a different Hilbert space.








Achieving local regularity does \emph{not} come at the expense of (local asymptotic) power. I characterise  power bounds for tests in non-regular models and show
that the C($\alpha$) tests proposed in this paper acheive these power bounds provided the moment conditions are chosen optimally. These power bounds contain those for regular models as a special case. Moreover, the conditions required for attainment of the power bounds are weaker than those in the literature.\footnote{In particular, in regular models the attainment result is well known if either (a) the observations are i.i.d. \cite[cf.][Chapter 25]{vdV98} or (b) the information operator (as defined in \citealp[][p. 846]{CHS96}) is boundedly invertible \citep{CHS96}. The result in this paper does not require either of these conditions.}

Following the theoretical development, I give details of its application to two examples: (i) a single index model which may be weakly identified when the link function is too flat and (ii) an instrumental variables (IV) model which may be weakly identified when the (nonparametric) first stage is too close to a constant function. Simulation experiments based on these examples demonstrate that the proposed tests enjoy good finite sample performance.

The application to IV may also be of interest for empirical researchers concerned about weak instruments. If the instruments are mean independent of the errors, then the test proposed here is robust to weak identification and can be substantially more powerful than tests assuming a linear first stage. This imposes no cost if the true first stage is (approximately) linear:  the power of the proposed test is comparable to optimal tests based on a linear first stage. The practical use of these tests is demonstrated in two IV applications with possibly weak instruments.
















This paper is connected to three main strands of the literature: the first is that concerned with general results on estimation and testing in semiparametric models. Much of this is now textbook material: see e.g. \cite{N90, CHS96, BKRW98, vdV98}. The second is the literature on C($\alpha$) tests. These were introduced by \cite{N59, N79} and have seen many useful applications, most recently as a way to handle machine learning or otherwise high dimensional first steps \citep[see e.g.][]{CHS15,BEvK20, CEINR22}. In this paper, the same structure which ensures good performance in such settings is used for a different purpose -- to construct tests which remain robust in non-regular settings. Lastly, the literature on robust testing in non -- regular or otherwise non -- standard settings is closely related to this paper \cite[e.g.][]{AG09,RS12,EMW15, McC17}. In particular, the  locally regular tests derived in this paper are especially useful in cases of weak identification and therefore this paper is closely related to the literature on weak identification robust inference \citep[e.g.][]{SS97, D97, SW00, K05, AC12, AM15, AM16b}. More specifically, this paper is most closely related to the recent work on semiparametric weak identification \citep{Kaji21, AM22} and extends the notion of semiparametric weak identification considered therein to non -- i.i.d. models.\footnote{Failure of local identification and singularity of the information matrix are closely linked in parametric models, see \cite{R71}. In the semiparametric case, parameters may be identified but nevertheless have a singular efficient information matrix. The relationship between the efficient information matrix and identification is considered by \cite{E22}.}








































































\section{Locally regular testing}\label{sec:heuristic}

\subsection{The local setup}

The goal considered throughout this paper is to construct hypothesis tests of $\mathrm{H}_0: \theta = \theta_0$ against $\mathrm{H}_1: \theta \neq \theta_0$ in the sequence of models $\mathcal{P}_n = \{P_{n, \gamma}: \gamma\in \Gamma\}$ where $\gamma = (\theta, \eta)\in \Gamma = \Theta\times\mathcal{H}$ for some open $\Theta\subset \mathbb{R}^{d_\theta}$ and $\mathcal{H}$ an arbitrary set. Each $\mathcal{P}_n$ consists of probability measures on a measurable space $(\mathcal{W}_n, \mathcal{B}(\mathcal{W}_n))$ and is dominated by a $\sigma$-finite measure $\nu_n$.\footnote{Typically the index $n$ is sample size and $\mathcal{W}_n$ is the space in which a sample of size $n$ takes its values. This is the situation considered in Section \ref{ssec:smooth-iid} as well as in the examples in Section \ref{sec:examples}.
}

Let $H_{\gamma}=\mathbb{R}^{d_\theta}\times B_{\gamma}$ be a subset of a linear space containing 0, and suppose that $\{P_{n, \gamma, h}: h\in H_{\gamma}\}\subset \mathcal{P}_n$ are such that $P_{n, \gamma} = P_{n, \gamma, 0}$. Elements of $H_\gamma$ will be written as $h = (\tau, b) \in \mathbb{R}^{d_\theta}\times B_{\gamma}$.\footnote{In most examples, $H_\gamma$ will be a linear space. The more general situation as considered here is nevertheless important to allow for, for example, Euclidean nuisance parameters subject to boundary constraints. In such a setting, if the constraint is binding at $\gamma$, then $\gamma$ can only be perturbed in certain directions if $P_{n, \gamma, h}$ is to remain within the model.} The measures $P_{n, \gamma, h}$ should be viewed as local perturbations of the measure $P_{n, \gamma}$ in a ``direction'' $h\in H_\gamma$. These local perturbations can be split in two groups: the perturbations $P_{n, \gamma, h}$ with  $h\in H_{\gamma, 0}\coloneqq \{(0, b): b\in B_{\gamma}\}$ correspond to the null hypothesis $\mathrm{H}_0: \theta = \theta_0$ and those with $h\in H_{\gamma, 1}\coloneqq \{h = (\tau, b): 0 \neq \tau\in \mathbb{R}^{d_\theta}, b\in B_{\gamma}\}$ to the alternative $\mathrm{H}_1: \theta \neq \theta_0$.
As such, $P_{n, \gamma, h}$ for $h\in H_{\gamma, 0}$ will be referred to as \emph{local perturbations consistent with the null hypothesis}, whilst $P_{n, \gamma, h}$ for $h\in H_{\gamma, 1}$ are \emph{local alternatives}.
The subsequent analysis is local with the parameter $\gamma$ being considered fixed at a $\gamma$ consistent with $\mathrm{H}_0$. As such, to lighten the notation, dependence on $\gamma$ will be mostly left implicit: I write  $P_{n, h}$ for $P_{n, \gamma, h}$, $H$ for $H_\gamma$, $H_{i}$ for $H_{\gamma, i}$ ($i=0, 1$) and similarly for other objects. I also use the abbreviation $P_n\coloneqq P_{n, 0}$.

I use the single-index model as a running example throughout the paper.\footnote{Technical details for this example are deferred to Sections \ref{ssec:sim} and \ref{ssec:example-details-sim}.}

\begin{example}[Single-index model]\label{ex:SIM-running-example}
    Suppose that the researcher observes $n$ i.i.d. copies of $W = (Y, X_1, X_2) \in \mathbb{R}^{2 + K}$ where
    \begin{equation}\label{eq:SIM-running-example-mdl}
        Y = f(X_1 + X_2^\prime\theta) + \epsilon, \qquad \mathbb{E}[\epsilon  | X] = 0,
    \end{equation}
    and where $f$ belongs to some set of continuously differentiable functions $\mathscr{F}$. The description of the model is completed by $\zeta \in \mathscr{Z}$, the density function of $(\epsilon, X)$ with respect to some $\sigma$-finite measure.
    The model $\mathcal{P}_n$ consists of the product measures $P_n = P^n$ where $P$ is the probability measure corresponding to the density
    \begin{equation}\label{eq:SIM-running-example-dens}
        p(W) =  p_{\gamma}(W) \coloneqq \zeta(\epsilon_{f, \theta},\, X), \qquad \epsilon_{f, \theta}\coloneqq Y - f(V_\theta),\quad  V_\theta \coloneqq X_1 + X_2^\prime \theta,
    \end{equation}
    for a $\gamma = (\theta, f, \zeta)\in \Theta \times \mathscr{F} \times \mathscr{Z} = \Gamma$.
    A class of local perturbations to this model are the probability measures $P_{n, h} = P_{h}^n$ where $P_h$ has density $p_{\gamma + \varphi_n(h)}$ with
    \begin{equation}\label{eq:SIM-local-alt}
             \varphi_n(h) = \left(\tau,\, b_1,\, b_2\zeta \right) / \sqrt{n},\qquad h = (\tau, (b_1, b_2)) \in H\coloneqq \mathbb{R}^{d_\theta} \times (B_1 \times B_2),
    \end{equation}
    where $B_1$ is a subset of the  bounded, continuously differentiable functions with bounded derivative and $B_2$ is a subset of the bounded functions $b_2:\mathbb{R}^{1+K}\to \mathbb{R}$, continuously differentiable in the first argument with bounded derivative.
\end{example}

\subsection{Local asymptotic normality}

The key technical condition under which the theory is developed is local asymptotic normality (LAN; see e.g. \citealp[Chapter 7]{vdV98} or \citealp[Chapter 6]{LCY00}).
Define the log-likelihood ratios
\begin{equation}\label{eq:LLR}
    L_{n}(h) \coloneqq \log \frac{p_{n, h}}{p_{n, 0}}, \qquad \text{ where } \ p_{n,  h}\coloneqq \frac{\mathrm{d}P_{n, h}}{\mathrm{d}\nu_n}, \text{ for } h\in H.
\end{equation}

\begin{assumption}[LAN]\label{ass:LAN}
    For bounded linear maps $\Delta_{n}:\overline{\operatorname{lin}}\  H_{}\to L_2^0(P_{n})$,
    \begin{equation}\label{eq:LAN}
        L_{n}(h) = \Delta_{n}h - \frac{1}{2} \|\Delta_{n}h\|^2 + R_{n}(h), \qquad h\in H
    \end{equation}
    with  $R_{n}(h) \xrightarrow{P_{n}} 0$ for all $h\in H$. Additionally, for each $h\in H$, the law of
    $\Delta_{n}h$ converges to $\mathcal{N}(0,  \sigma(h))$ in the Mallows-2 metric, $d_2$.
\end{assumption}

The requirement that $\Delta_{n}h$ converges in $d_2$ is equivalent to requiring that it converges weakly and $(\Delta_{n}h)_{n\in \mathbb{N}}$ is uniformly square $P_{n}$-integrable \cite[e.g.][Appendix A.6]{BKRW98}. This implies that $\sigma(h)= \lim_{n\to\infty} \|\Delta_{n}h\|^2$.


\begin{remark}\label{rem:mutual-contiguity}
    Assumption \ref{ass:LAN} ensures that the sequences $(P_{n})_{n\in \mathbb{N}}$ and $(P_{n, h})_{n\in \mathbb{N}}$ are mutually contiguous for any $h\in H$ \cite[see e.g.][Example 6.5]{vdV98}.
\end{remark}

\begin{remark}\label{rem:LAN-vs-ULAN}
    If $H$ is (pseudo-)metrised one may consider a uniform version of Assumption \ref{ass:LAN}, i.e. uniform local asymptotic normality (ULAN). Such a version is given in Assumption \ref{ass:ULAN} and is equivalent to Assumption \ref{ass:LAN} plus asymptotic equicontinuity on compact sets of $h\mapsto \Delta_{n} h$ (in $L_2(P_{n, 0})$) and $h\mapsto P_{n, h}$ (in total variation) (Proposition \ref{prop:ULAN-LAN-equicontinuity}). The latter equicontinuity condition is of interest regarding
      local uniformity of size control; cf. Corollary \ref{cor:psi-locally-uniformly-regular} and Lemma \ref{lem:level-alpha-unif-equicontinuity-TV} below.
\end{remark}

\begin{example}[
    continues=ex:SIM-running-example]
    Under regularity conditions, the single-index model satisfies Assumption \ref{ass:LAN} with
    \begin{equation}\label{eq:SIM-running-example-score-op}
        \Delta_nh \coloneqq \frac{1}{\sqrt{n}}\sum_{i=1}^n \tau^\prime \dot{\ell}_{\gamma}(W_i) + [Db](W_i),
    \end{equation}
    where for $\phi(e, x)\coloneqq \frac{\partial \log \zeta(e, x)}{\partial e}$,
    \begin{equation}\label{eq:SIM-running-example-scores}
        \dot{\ell}_{\gamma}(W) \coloneqq -\phi(\epsilon_{f, \theta}, X) f'(V_\theta)X_2, \quad
        [D b](W) \coloneqq -\phi(\epsilon_{f, \theta}, X)b_{1}(V_{\theta}) + b_2(\epsilon_{f, \theta}, X).
\end{equation}
\end{example}



\subsection{Local regularity for tests}


\begin{definition}\label{defn:locally-regular-test}
    A sequence of tests $\phi_n:\mathcal{W}_n\to [0, 1]$ of the hypothesis $\mathrm{H}_0:\theta = \theta_0$
     against $\mathrm{H}_1: \theta\neq \theta_0$
      is asymptotically of level $\alpha$ and locally regular if
    \begin{equation}\label{eq:locally-regular-test}
        \uppi_{n}(\tau, b) \coloneqq P_{n,  h}\phi_n \to \uppi(\tau), \quad h = (\tau, b)\in H\quad \text{ and }\quad \uppi(0) \le \alpha.
    \end{equation}
\end{definition}

That is, the finite sample (local) power function of the test, $\uppi_n$ converges under each $P_{n, h}$ to a function $\uppi$ which may depend on $\tau$ (and, implicitly, $\gamma$) but not on $b$, the parameter which describes local deviations from the nuisance parameter $\eta$.\footnote{Cf. the definition of a (locally) regular estimator in \citealp[e.g.][p. 365]{vdV98}.}
If a sequence of tests does \emph{not} satisfy \eqref{eq:locally-regular-test} it is \emph{(locally) non-regular}.

Local regularity of test sequences as in \eqref{eq:locally-regular-test} is a pointwise concept. It is also of interest to consider a uniform version.

\begin{definition}\label{defn:locally-uniformly-regular-test}
    A sequence of tests $\phi_n:\mathcal{W}_n\to [0, 1]$ of the hypothesis $\mathrm{H}_0:\theta = \theta_0$
     against $\mathrm{H}_1: \theta\neq \theta_0$
    is asymptotically of level $\alpha$ and locally uniformly regular on $K\subset H$ if \eqref{eq:locally-regular-test} holds uniformly on $K$.
\end{definition}

If $H$ is a (pseudo-)metric space and $K$ is a compact set, for the convergence in \eqref{eq:locally-regular-test} to hold uniformly on $K$ it is necessary and sufficient to show that the sequence of functions $\uppi_n$ is asymptotically equicontinuous on $K$.\footnote{
The same is true if $K$ is totally bounded. See e.g. \cite{D21}, p. 123, for the definition of asymptotic equicontinuity.\label{ftnt:K-compact-totally-bounded}}

Directly working with the power functions $\uppi_n$ to show their asymptotic equicontinuity is complicated in many cases. It is, however, often possible to show results which imply this property. For instance, the functions $h\mapsto P_{n, h}$ being asymptotically equicontinuous in $d_{TV}$ implies the required asymptotic equicontinuity of the power functions. Despite being (much) stronger, this often holds.\footnote{
    See the discussion following Remark \ref{rem:ULAN-compact-equicontinuity-in-TV} below.
}

\paragraph{Weak identification asymptotics and local regularity}

In many models there are parameter values, $\gamma$, at which locally regular estimators do not exist. Points where the parameter of interest, $\theta$, is un- or under-identified provide an important class of examples.
Moreover, as is well known from the literature on weak identification, even if $\theta$ is identified at $\gamma$, finite sample inference may be poor if $\gamma$ is too close to a point of identification failure relative to the amount of information contained in the sample. Such behaviour has been widely studied in models where the part of $\gamma$ causing the identification failure is finite dimensional \citep[e.g.][]{AC12, AM15}.




There are also many examples where weak identification may occur due to the value of \emph{infinite-dimensional} nuisance parameters. \cite{Kaji21} and \cite{AM22} use a differentiability in quadratic mean (DQM) condition to define semiparametric weak identification asymptotics in i.i.d. models. In particular, they consider sequences $P_{n, h}^n$ which satisfy
\begin{equation}\label{eq:dqm-weakid}
    \lim_{n\to\infty} \int \left[\sqrt{n}\left(\sqrt{p_{n, h}} - \sqrt{p_{0}} \right) - \frac{1}{2}f\sqrt{p_{0}} \right]^2\,\mathrm{d}\nu_n  = 0
\end{equation}
for a point $P_0$ where the parameter of interest is unidentified.
In the i.i.d. case, \eqref{eq:dqm-weakid} implies the LAN expansion in Assumption \ref{ass:LAN} with $\Delta_nh = \frac{1}{\sqrt{n}}\sum_{i=1}^n f(W_i)$ \cite[e.g.][Lemma 25.14]{vdV98}.\footnote{
If Assumption \ref{ass:LAN} holds with $\Delta_n$ having this form, the converse is also true.
} Working with Assumption \ref{ass:LAN} in place of \eqref{eq:dqm-weakid} broadens the applicability of this class of semiparametric weak identification asymptotics to non-i.i.d. models. It is clear from Definition \ref{defn:locally-regular-test} that a locally regular test sequence will have asymptotic null rejection probability (NRP) which does not exceed the nominal level under weak identification asymptotics $P_{n, h}$.\footnote{Of course, a (non-regular) test sequence may have asymptotic NRP which depends on $b$ and yet is bounded by the nominal level under $P_{n, h}$ for all $h = (0, b)\in H_0$ and / or have an asymptotic power function which depends on $b$ for $h = (\tau, b)\in H_1$. Restricting attention to locally regular test sequences may be justified by the power optimality results of Section \ref{sec:theory}.}

I now give two examples of semiparametric models where the parameter of interest $\theta$ may be un- or under-identified depending on the value of an infinite dimensional nuisance parameter.\footnote{A further example is the linear simultaneous equations model in \cite{LM21}.}
The first is the running example.

\begin{example}[
    continues=ex:SIM-running-example]
    As is clear from the model equation $ Y = f(X_1 + X_2^\prime\theta) + \epsilon$, if $f$ is flat, i.e. $f' = 0$, then the parameter $\theta$ is unidentified. The sequences given in \eqref{eq:SIM-local-alt} are weak identification asymptotic sequences if $f'=0$.
\end{example}

\begin{example}[IV]\label{ex:IV}
    Suppose the researcher observes $n$ i.i.d. copies of $W=(Y, X, Z)$,
    \begin{equation*}
        Y = X^\prime \theta  +Z_1^\prime\beta + \epsilon, \qquad \mathbb{E}[\epsilon |Z] = 0, \qquad Z = (Z_1^\prime, Z_2^\prime)^\prime.
    \end{equation*}
    If $\pi(Z) \coloneqq \mathbb{E}[X|Z]$ is constant, $\theta$ is unidentified; if some components of $\pi(Z)$ are constant, $\theta$ is underidentified.
\end{example}



In Examples \ref{ex:SIM-running-example} and \ref{ex:IV}, at the points of identification failure, no locally regular estimator exists, however locally regular C($\alpha$) tests are developed in Section \ref{sec:examples}.\footnote{These examples consider i.i.d. data for simplicity. See \cite{HLM22} for an example of a locally regular C($\alpha$) test of the form proposed in this paper for the potentially un- / under-identified parameter in a structural vector autoregressive model.}


\subsection{A class of locally regular tests}

To construct locally regular tests of $\mathrm{H}_0: \theta = \theta_0$ against $\mathrm{H}_1: \theta\neq \theta_0$,
I use a generalisation of the class of C($\alpha$) tests introduced by \cite{N59, N79} to characterise optimal tests in regular parametric models. These tests are a based on a quadratic form of (estimators of) a vector of $d_\theta$ moment conditions $g_n\in L_2(P_n)$ which satisfy the following requirements.

\begin{assumption}[Joint convergence]\label{ass:joint-conv}
    For $g_{n}\in L_2(P_{n})^{d_\theta}$ and each $h= (\tau, b)\in H$,
    \begin{equation*}
        \left(\Delta_{n}h,\; g_{n}^\prime\right)^\prime \overset{P_{n}}{\rightsquigarrow }\mathcal{N}\left(0, \Sigma(h)\right),
    \end{equation*}
    \begin{equation*}
        \Sigma(h) \coloneqq \begin{bmatrix}
            \sigma(h) & \tau^\prime\Sigma_{ 21}^\prime\\
            \Sigma_{ 21}\tau & V
        \end{bmatrix} = \lim_{n\to\infty} \begin{bmatrix}
            \|\Delta_{n}h\|^2 & \left\langle {\Delta_{n}(\tau, 0)}\, ,\, {g_{n}^\prime} \right\rangle\\
            \left\langle {g_{n}}\, ,\, {\Delta_{n}(\tau, 0)} \right\rangle & \left\langle {g_{n}}\, ,\, {g_{n}^\prime} \right\rangle
        \end{bmatrix}.
    \end{equation*}
\end{assumption}

Built-in to Assumption \ref{ass:joint-conv} is a requirement of asymptotic orthogonality of $g_n$ and the scores for the nuisance parameters $\eta$. This generalises the analogous condition in \cite{N59, N79} and is key to the local regularity of C$(\alpha)$ tests.

\begin{remark}\label{rem:orth}
    For Assumption \ref{ass:joint-conv} to hold it is necessary that the $g_{n}$ are approximately zero mean: since $(g_{n})_{n\in \mathbb{N}}$ is uniformly $P_n$-integrable, $P_{n} g_{n} = o(1)$.
    It is also necessary that the $g_{n}$ satisfy an approximate orthogonality property with the scores for nuisance parameters: as $([\Delta_{n}h]g_{n})_{n\in \mathbb{N}}$ is uniformly $P_n$-integrable for each $h = (\tau, b)\in H$,
    $   \lim_{n\to\infty} \left\langle {\Delta_{n}h}\, ,\, {g_{n}^\prime} \right\rangle =  \tau^\prime\Sigma_{21}^\prime = \lim_{n\to\infty} \left\langle {\Delta_{n}(\tau, 0)}\, ,\, {g_{n}^\prime} \right\rangle$,  and so
    \begin{equation}\label{rem:orth:eq:orth}
        \left\langle {\Delta_{n}(0, b)}\, ,\, {g_{n}^\prime} \right\rangle = \left\langle {\Delta_{n}h}\, ,\, {g_{n}^\prime} \right\rangle - \left\langle {\Delta_{n}(\tau, 0)}\, ,\, {g_{n}^\prime} \right\rangle = o(1).
    \end{equation}
\end{remark}

Given any $d_\theta$ moment conditions $f_{n}\in L_2^0(P_{n})$, moment conditions which satisfy an exact version of the orthogonality condition \eqref{rem:orth:eq:orth} may be obtained as
\begin{equation}\label{eq:orth-proj-g}
    g_{n} \coloneqq  \Pi\left[f_{n} \middle| \left\{\Delta_{n}(0, b) : b\in B\right\}^{\perp}\right].
\end{equation}

An important special case of this construction is with $f_{n}$ the score function for $\theta$, i.e. $f_{n} = \dot{\ell}_{n}$ such that $\tau^\prime \dot{\ell}_{n}= \Delta_{n}(\tau, 0)$ for each $\tau\in \mathbb{R}^{d_\theta}$. The function
\begin{equation}\label{eq:effscr}
    g_{n} = \tilde{\ell}_{n} \coloneqq \Pi\left[\dot{\ell}_{n}\middle| \left\{\Delta_{n}(0, b) : b\in B\right\}^{\perp}\right],
\end{equation}
is called the \emph{efficient score function}.
 This yields a power optimal choice of moment conditions satisfying \eqref{rem:orth:eq:orth} as shown in Section \ref{sec:theory} below.

\begin{example}[continues=ex:SIM-running-example,
    ]
    Let $\upomega: \mathbb{R}^{K}\to [\underline{\upomega},
    \overline{\upomega}]\subset (0, \infty)$. Then $g_n\coloneqq \mathbb{G}_n g$,
    \begin{equation}\label{ex:SIM-running-example:eq:g}
        g(W)\coloneqq\upomega(X)(Y - f(V_{\theta}))f'(V_{\theta})\left(X_2 - \frac{\mathbb{E}[\upomega(X) X_2| V_{\theta}]}{\mathbb{E}[\upomega(X) |V_{\theta}]}\right),
    \end{equation}
    has components which belong to $\{\Delta_n(0, b): b\in B\}^\perp$ (where $\Delta_n$ is as in \eqref{eq:SIM-running-example-score-op}).\footnote{
   $g$ coincides with the efficient score function $\tilde{\ell}$ (derived by \citealp{NS93}),
\begin{equation*}
   \tilde{\ell}(W)= \tilde{\upomega}(X)(Y - f(V_{\theta}))f'(V_{\theta})\left(X_2 - \frac{\mathbb{E}[\tilde{\upomega}(X)X_2| V_{\theta}]}{\mathbb{E}[\tilde{\upomega}(X) | V_{\theta}]}\right), \quad \tilde{\upomega}(X)\coloneqq \mathbb{E}[\epsilon^2|X]^{-1},
\end{equation*}
in the (typically infeasible) case with $\upomega = \tilde{\upomega}$.
    } Under regularity conditions, $g_n$ satisfies Assumption \ref{ass:joint-conv} (see Section \ref{ssec:sim} below).

\end{example}

To construct the test statistic, I assume that consistent estimators of $g_{n}$, $V^\dagger$ (the Moore-Penrose pseudo-inverse of $V$) and $r\coloneqq \operatorname{rank}(V)$ are available, given $\theta$.

\begin{assumption}[Consistent estimation]\label{ass:consistent}
    $\hat{g}_{n, \theta}$, $\hat\Lambda_{n, \theta}$, $\hat{r}_{n, \theta}\in \{0, 1, \ldots, d_\theta\}$ satisfy
    \begin{enumerate}
        \item $\hat{g}_{n, \theta} - g_{n} \xrightarrow{P_{n}}0$;\label{ass:consistent:itm:ghat}
        \item $\hat\Lambda_{n, \theta} \xrightarrow{P_{n}}  V^\dagger$;\label{ass:consistent:itm:lambdahat}
        \item If $r \ge1$, then $\hat{r}_{n, \theta} \xrightarrow{P_{n}}r$; if $r = 0$, then $\operatorname{rank}(\hat\Lambda_{n, \theta}) \xrightarrow{P_{n}} 0$.
        \label{ass:consistent:itm:rank}
    \end{enumerate}
\end{assumption}

Verification of Assumption \ref{ass:consistent}\ref{ass:consistent:itm:ghat} typically proceeds by model specific arguments. That $g_n$ is (asymptotically) orthogonal to $\{\Delta_n(0, b):b\in B\}$ often helps in establishing this consistency (cf. \citealp{CEINR22}).
One generally applicable approach to obtain an estimator which satisfies Assumption \ref{ass:consistent}\ref{ass:consistent:itm:lambdahat} is to take an initial estimator which is consistent for $V$, threshold its eigenvalues at an appropriate rate and then take the pseudo-inverse.\footnote{See Section S5 of \cite{LM21-S} for full details of this approach. Other regularisation schemes are also possible (see e.g. \citealp{DV15}).} If one uses the estimator $\hat{\Lambda}_{n, \theta} \coloneqq \hat{V}_{n, \theta}^\dagger$ where $\hat{V}_{n, \theta}\xrightarrow{P_{n}}V$ and $\hat{r}_{n, \theta}\coloneqq \operatorname{rank}(\hat{V}_{n, \theta})$ then condition \ref{ass:consistent:itm:lambdahat} holds if and only if condition \ref{ass:consistent:itm:rank} holds \cite[Theorem 2]{A87}. Nevertheless, as emphasised by the notation, it is not necessary that the estimate $\hat{\Lambda}_{n, \theta}$ be the pseudo-inverse of an initial estimate.

\begin{example}[continues = ex:SIM-running-example]
    Given estimators $\hat{f}_{n, i}$, $\widehat{f'}_{n, i}$ of $f, f'$ and $\hat{Z}_{1, n,i}, \hat{Z}_{2, n, i}$ of $Z_1\coloneqq \mathbb{E}[\upomega(X)X_2|V_{\theta}]$, $Z_2\coloneqq \mathbb{E}[\upomega(X)|V_{\theta}]$, define $ \hat{g}_{n, \theta} \coloneqq \frac{1}{\sqrt{n}}\sum_{i=1}^n \hat{g}_{n, \theta, i}$,
    \begin{equation}\label{ex:SIM-running-example:eq:ghati}
        \hat{g}_{n, \theta, i}\coloneqq  \upomega(X_i)(Y_i - \hat{f}_{n, i}(V_{\theta, i}))\widehat{f'}_{n, i}(V_{\theta, i})\left(X_{2, i} - \hat{Z}_{1, n, i}(V_{\theta, i}) / \hat{Z}_{2, n, i}(V_{\theta, i})\right).
    \end{equation}
    Under regularity conditions (see Section \ref{ssec:sim}), $\hat{g}_{n, \theta}$ satisfies part \ref{ass:consistent:itm:ghat} of Assumption \ref{ass:consistent}, thresholding the eigenvalues of $\frac{1}{n}\sum_{i=1}^n \hat{g}_{n, \theta, i}\hat{g}_{n, \theta, i}^\prime$ at an appropriate rate yields an estimator $\hat{\Lambda}_{n, \theta}$ which satisfies part \ref{ass:consistent:itm:lambdahat} and $\hat{r}_{n, \theta}\coloneqq \operatorname{rank}(\hat\Lambda_{n, \theta})$ satisfies part \ref{ass:consistent:itm:rank}.
\end{example}


Given the estimators of Assumption \ref{ass:consistent}, the C($\alpha$)-style test statistic is
\begin{equation}\label{eq:Shat}
    \hat{S}_{n, \theta} \coloneqq \hat{g}_{n, \theta}^\prime \hat{\Lambda}_{n, \theta} \hat{g}_{n, \theta} .
\end{equation}

The C($\alpha$) -- style test $\psi_{n, \theta_0}$ of $\mathrm{H}_0$ against $\mathrm{H}_1$
 at level $\alpha$ is:
\begin{equation}\label{eq:psi-test}
    \psi_{n, \theta_0} \coloneqq \bm{1}\left\{
            \hat{S}_{n, \theta_0}  > c_n
        \right\},
\end{equation}
where $c_n$ is the $1-\alpha$ quantile of a $\chi^2_{\hat{r}_n}$ random variable.

\paragraph{Local regularity} Assumptions \ref{ass:LAN} -- \ref{ass:consistent} suffice for local regularity of $\psi_{n, \theta_0}$.


\begin{proposition}\label{prop:asymp-dist}
Under Assumptions \ref{ass:LAN} and \ref{ass:joint-conv}, for $h = (\tau, b)\in H$
    \begin{equation*}
        g_{n} \overset{P_{n,  h}}{\rightsquigarrow} \mathcal{N}\left(\Sigma_{ 21}\tau, V\right).
    \end{equation*}
    If Assumption \ref{ass:consistent} also holds, then additionally
     \begin{equation*}
        \hat{g}_{n, \theta_0} \overset{P_{n,  h}}{\rightsquigarrow} \mathcal{N}\left(\Sigma_{ 21}\tau, V\right) \qquad \text{ and } \qquad \hat{S}_{n, \theta_0} \overset{P_{n,  h}}{\rightsquigarrow} \chi^2_r\left(
            \tau^\prime \Sigma_{ 21}^\prime V \Sigma_{ 21}\tau
        \right).
    \end{equation*}
\end{proposition}

\begin{theorem}\label{thm:psi-pwr-local-alt}
    Suppose that Assumptions \ref{ass:LAN}, \ref{ass:joint-conv} and \ref{ass:consistent} hold and $h = (\tau, b) \in H$. Then,
    \begin{equation*}
        \lim_{n\to\infty} P_{n,  h} \psi_{n, \theta_0} = \uppi(\tau)\coloneqq  \begin{cases}
            1- \mathrm{P}\left(\chi_r^2\left(\tau^\prime\Sigma_{ 21}^\prime  V ^\dagger \Sigma_{ 21}\tau\right) \le c_r\right) & \text{ if } r\ge 1\\
            0 &\text{ if }r=0
        \end{cases},
    \end{equation*}
    where $c_r$ is the $1-\alpha$ quantile of the $\chi^2_r$ distribution.
\end{theorem}

Theorem \ref{thm:psi-pwr-local-alt} immediately shows that $\psi_{n, \theta_0}$ is locally regular (cf. \eqref{eq:locally-regular-test}). The asymptotic orthogonality in \eqref{rem:orth:eq:orth} is key to this result.
If, instead, $\lim_{n\to\infty} \left\langle {\Delta_n h}\, ,\, {g_{n}^\prime} \right\rangle = \tau^\prime \Sigma_{21}^\prime + c(b)$ with $c(b)\neq 0$, then (by Le Cam's third Lemma) the limiting distribution of $g_n$ under $P_{n, h}$ would be $\mathcal{N}(\Sigma_{21}\tau + c(b), V)$  and hence the limiting power function of the test sequence would not be free of $b$.

\paragraph{Uniform local regularity}

The local regularity given by \ref{thm:psi-pwr-local-alt} may be ``upgraded'' to local uniform regularity (Definition \ref{defn:locally-uniformly-regular-test}) under various conditions. Here I consider the case where $H$ posseses a (pseudo-)metric structure (e.g. if $H$ is a subset of a (semi-)normed linear space).\footnote{If $H$ posseses a (finite) measure structure and the functions $h =(\tau, b)\mapsto \uppi_{n}(\tau, b)$ are measurable then $\psi_{n, \theta_0}$ is locally uniformly regular except on a ``small'' subset of $H$ by Egorov's Theorem. See Section \ref{ssec:unif-measure-structure} for details.}
In this case, for $\psi_{n, \theta_0}$ to be locally uniformly regular on a compact (or totally bounded) $K\subset H$ it is necessary and sufficient that the functions $h =(\tau, b)\mapsto \uppi_{n}(\tau, b)$ are asymptotically equicontinuous.

\begin{corollary}\label{cor:psi-locally-uniformly-regular}
    Suppose that the conditions of Theorem \ref{thm:psi-pwr-local-alt} hold and that $(H, d)$ is a pseudometric space. If the functions $h = (\tau, b)\mapsto \uppi_n(\tau, b) \coloneqq  P_{n,  h} \psi_{n, \theta_0}$ are asymptotically equicontinuous on a compact (or totally bounded) subset $K \subset H$,
    \begin{equation*}
        \lim_{n\to\infty} \sup_{(\tau, b)\in K}\left|\uppi_n(\tau, b) - \uppi(\tau)
        \right| = 0.
    \end{equation*}
\end{corollary}

I now give a sufficient condition for the asymptotic equicontinuity required by Corollary \ref{cor:psi-locally-uniformly-regular}.

\begin{lemma}\label{lem:level-alpha-unif-equicontinuity-TV}
    If $(H, d)$ is a pseudometric space and $(h\mapsto P_{n, h})_{n\in \mathbb{N}}$ is asymptotically equicontinuous in $d_{TV}$ on $K\subset H$, then $(h\mapsto P_{n, h} \psi_{n, \theta})_{n\in \mathbb{N}}$ is asymptotically equicontinuous on $K$.
\end{lemma}

\begin{remark}\label{rem:ULAN-compact-equicontinuity-in-TV}
    Lemma \ref{lem:level-alpha-unif-equicontinuity-TV} requires asymptotic equicontinuity in total variation of the functions $(h\mapsto P_{n, h})_{n\in \mathbb{N}}$ on subsets $K\subset H$. This holds for any compact $K$ under ULAN (Assumption \ref{ass:ULAN}), as shown in Proposition \ref{prop:ULAN-LAN-equicontinuity}.
\end{remark}



In the parametric i.i.d. case LAN is often verified by establishing a DQM condition, e.g. equation (7.1) in \cite{vdV98}. This is sufficient for the ULAN expansion in Assumption \ref{ass:ULAN} to hold \cite[e.g.][Theorem 7.2]{vdV98}. Semiparametric generalisations of this result are available (e.g. combine Proposition \ref{prop:ULAN-LAN-equicontinuity} and Lemma \ref{lem:iid-ULAN-DQM}).\footnote{\cite{LM21} and \cite{HLM22} verify this asymptotic equicontinuity property in i.i.d. and time series semiparametric examples respectively.}

The condition in Lemma \ref{lem:level-alpha-unif-equicontinuity-TV} is natural given its link with the ULAN condition. Neverthelesss, it is (much) stronger than necessary for the condition required by Corollary \ref{cor:psi-locally-uniformly-regular}; see Lemma \ref{lem:level-alpha-unif-equicontinuity-weak} for  a weaker sufficient condition.




























































































\section{Power optimality}\label{sec:theory}






















































































The preceding section established the local regularity of the tests $\psi_{n, \theta_0}$ based on (estimates of) moment functions $g_{n}$ satisfying certain asymptotic orthogonality conditions. Thus far, nothing has been said about the choice of $g_{n}$ beyond these orthogonality requirements.
The choice of the functions $g_{n}$ determines the power of the corresponding test. As such, they ought to be chosen such that the resulting test has good power against alternatives of interest.

One natural choice is the efficient score function \eqref{eq:effscr}.
It is well known that tests based on the efficient score function have certain optimality properties in regular models when (a) the observations are i.i.d. \cite[cf. Section 25.6][]{vdV98} or (b) when the information operator for $\eta$ is boundedly invertible \citep{CHS96}. I show that this optimality persists in non-regular models and does not require (a) or (b).






The results in this section are derived using the limits of experiments framework of Le Cam \cite[e.g.][]{LC86, vdV98}. In particular, I show that the local experiments consisting of the measures $P_{n, h}$ for $h\in H$ converge weakly to a limit experiment which has a close relationship to a Gaussian shift experiment on the Hilbert space formed by taking the quotient of $H$ under the seminorm induced by the variance function $\sigma(h)$. The connection between these experiments is sufficiently tight that power bounds derived in the latter transfer to the former.\footnote{That the local experiments do not converge to the mentioned Gaussian shift experiment is essentially a purely technical point: the Gaussian shift experiment is defined on a different parameter space to the local experiments, whilst (weak) convergence of experiments (in the sense of \citealp{LC86}) is defined for experiments with the same parameter space.}

\paragraph{The limit experiment}
For this development $H$ is required to be linear and I will therefore assume that $B$ (hence $H$) is a linear space.
Under LAN, there exists a positive semi-definite symmetric bilinear form $\left\langle {\cdot}\, ,\, {\cdot} \right\rangle_{K}$ on $H=  \mathbb{R}^{d_\theta} \times B$ such that $\sigma(h) = \left\langle {h}\, ,\, {h} \right\rangle_{K}$.
This can be seen as a by-product of the following Lemma.
\begin{lemma}\label{lem:IP-GP}
    Suppose Assumption \ref{ass:LAN} holds and $B$ is a linear space. Let $\Delta$ be a square integrable stochastic process defined on $H$ such that $\Delta_{n}h \overset{P_{n}}{\rightsquigarrow} \Delta h$.
    Then $\Delta$ is a mean-zero Gaussian linear process with covariance kernel $K$, where
    \begin{equation*}
        K(h, g) \coloneqq \lim_{n\to\infty} P_{n}\left[\Delta_{n}h\Delta_{n}g\right].
    \end{equation*}
\end{lemma}

For $h, g\in H$, setting $\left\langle {h}\, ,\, {g} \right\rangle_{K} \coloneqq K(h, g)$
gives a positive semi-definite symmetric bilinear form. Let $\|\cdot\|_{K}$ denote the seminorm induced by $\left\langle {\cdot}\, ,\, {\cdot} \right\rangle_{K}$ on $H$.

\begin{remark}\label{rem:IP-linear-operator}
    Suppose that $\left\langle {\cdot}\, ,\, {\cdot} \right\rangle_{H}$ is an inner product on $H = \mathbb{R}^{d_\theta} \times B$.
    The existence of the positive semi-definite symmetric bilinear form $\left\langle {\cdot}\, ,\, {\cdot} \right\rangle_{K}$ is equivalent to the existence of a bounded, self-adjoint, positive semi-definite linear operator $\mathsf{B}$ such that $\left\langle {h}\, ,\, {h} \right\rangle_{K} = \left\langle {h}\, ,\, {\mathsf{B} h} \right\rangle_{H}$ for $h\in H$
    \cite[cf.][p. 845]{CHS96}.
\end{remark}

Define $\mathbb{H}$ as the quotient of $H$ by the subspace on which $\|\cdot\|_{K}$ vanishes:
\begin{equation}\label{eq:mbHgam}
    \mathbb{H} \coloneqq {H} \,/\, {\{h\in H : \|h\|_{K} = 0\}}.
\end{equation}
which is an inner product space when equipped with the natural inner product induced by $\left\langle {\cdot}\, ,\, {\cdot} \right\rangle_{K}$, which I also denote by $\left\langle {\cdot}\, ,\, {\cdot} \right\rangle_{K}$.
An element of $\mathbb{H}$ corresponding to representative element $h\in H$ will be denoted by $[h]$.\footnote{
    Analogous comments apply to the related space $\mathbb{H}_1$, defined below. In both cases, to avoid an excess of parentheses / brackets, if $h = (\tau, b)$ I will write either $[h]$ or $[\tau, b]$, rather than $[(\tau, b)]$.
\label{ftnt:h-vs-[h]}
}

The (weak) limit of the sequence of experiments consisting of the measures $P_{n, h}$ can be obtained by standard results on weak convergence of experiments.

\begin{proposition}\label{prop:conv-exp}
    Suppose that Assumption \ref{ass:LAN} holds, that $B$ is a linear space
    and define the sequence of experiments $ \mathscr{E}_{n}\coloneqq \left(\mathcal{W}_n, \mathcal{B}(\mathcal{W}_n), (P_{n,  h}: h\in H)\right)$.
    Let $\Delta$ be the Gaussian process of Lemma \ref{lem:IP-GP} and let $(\Omega, \mathcal{F}, \mathrm{P})$ be the probability space on which it is defined. Define the experiment $\mathscr{E}\coloneqq (\Omega, \mathcal{F}, (P_{ h}: h\in H))$ according to
    \begin{equation*}
        P_{ 0}\coloneqq \mathrm{P}; \qquad \frac{\mathrm{d}P_{ h}}{\mathrm{d}P_{ 0}} = \exp\left(\Delta h - \frac{1}{2}\|h\|^2\right), \quad h\in H.
    \end{equation*}
    Then $\mathscr{E}_{n}$ converges weakly to $\mathscr{E}$.
\end{proposition}

Under the assumption that $\mathbb{H}$ is separable, the experiment $\mathcal{E}$ is equivalent to a Gaussian shift on $(\mathbb{H}, \left\langle {\cdot}\, ,\, {\cdot} \right\rangle_{K})$, in the sense given by  Proposition \ref{prop:exp-equiv-shift} below.

\begin{assumption}\label{ass:IP}
    $B$ is a linear space and $\mathbb{H}$ as defined in \eqref{eq:mbHgam} is separable.
\end{assumption}





























\begin{proposition}\label{prop:exp-equiv-shift}
    Suppose Assumptions \ref{ass:LAN} and \ref{ass:IP} hold. If $\mathscr{E}$ is as in Proposition \ref{prop:conv-exp}, there is a Gaussian shift experiment $\mathscr{G} \coloneqq (\Omega, \mathcal{F}, (G_{[h]} : [h]\in \mathbb{H}))$
    such that  $d_{TV}(P_{h}, G_{[h]}) = 0$ for each $h\in H$.
\end{proposition}

\paragraph{The efficient information matrix}

Power bounds for tests of $\mathrm{K}_0: h\in H_0$ against $\mathrm{K}_1: h\notin H_1$ can be expressed in terms of the \emph{efficient information matrix}, $\tilde{\mathcal{I}}_{\gamma}$, so named because in the i.i.d. setting it is the covariance matrix of the efficient score function for a single observation. Here I provide an alternative definition of this matrix which applies more generally, and reduces to the classical definition in the i.i.d. case (as shown in Lemma \ref{lem:effscr-coincides-with-usual-defn}).



Let $\|\tau\| \coloneqq  \inf_{b\in B}\|(\tau, b)\|_{K}$, which defines a semi-norm on $\mathbb{R}^{d_\theta}$.
Equipping the quotient $\mathbb{H}_1 \coloneqq  {\mathbb{R}^{d_\theta}} \,/\, {\{\tau \in \mathbb{R}^{d_\theta} :\|\tau\| = 0\}}$ with the natural norm induced by $\|\cdot\|$ (which I also denote by $\|\cdot\|$) turns it into a normed space.
Define the linear map $\pi_1:\mathbb{H} \to\mathbb{H}_1$ as $\pi_1([\tau, b]) \coloneqq [\tau]$.
As $\pi_1$ is continuous it may be uniquely extended to a continuous function defined on $\overline{\mathbb{H}}$, the completion of $\mathbb{H}$; this extension will also be called $\pi_1$.
Since $\pi_1$ is continuous, $\ker \pi_1\subset \overline{\mathbb{H}} $
is closed. Let $\Pi$ be the orthogonal projection onto $\ker \pi_1$ and define $\Pi^\perp\coloneqq I - \Pi$, the orthogonal projection onto $[\ker \pi_1]^\perp$.
Let $e_i$ be the $i$-th canonical basis vector in $\mathbb{R}^{d_\theta}$ and define the \emph{efficient information matrix} $\tilde{\mathcal{I}}_{\gamma}$ as the $d_\theta \times d_\theta$ matrix with $i,j$-th entry $\tilde{\mathcal{I}}_{ij}$ given by\footnote{Lemma \ref{lem:calculation-of-effinfo} gives an alternative expression for $\tilde{\mathcal{I}}_{\gamma}$ based on the Gaussian process $\Delta$ of Lemma \ref{lem:IP-GP}.}
\begin{equation}\label{eq:effinfo}
    \tilde{\mathcal{I}}_{ij} = \left\langle {\Pi^\perp [e_i, 0]}\, ,\, {\Pi^\perp [e_j, 0]} \right\rangle_{K}.
\end{equation}



\begin{lemma}\label{lem:ker-effinfo-is-fzero}
    Under Assumption \ref{ass:IP}, $\|\tau\|^2 = \tau^\prime \tilde{\mathcal{I}}_{\gamma}\tau$ and $\ker \tilde{\mathcal{I}}_{\gamma} = \{\tau\in \mathbb{R}^{d_\theta}: \|\tau\| = 0\}$.
\end{lemma}


\subsection{Tests of a scalar parameter}

The following Theorem records the power bound for (locally asymptotically) unbiased two-sided tests of a scalar $\theta$. In this case
the matrix $\tilde{\mathcal{I}}_{\gamma}$ has rank either 0 or 1.
Theorem \ref{thm:two-sided-pwr-bound} handles both cases simultaneously.
















\begin{theorem}\label{thm:two-sided-pwr-bound}
    Suppose that Assumptions \ref{ass:LAN} and \ref{ass:IP} hold  and $d_{\theta}=1$.
    Let $\phi_n: \mathcal{W}_n\to [0, 1]$ be a sequence of locally asymptotically unbiased level $\alpha$ tests of $\mathrm{K}_0: \tau = 0$ against $\mathrm{K}_1:\tau\neq 0$. That is,
    \begin{equation*}
        \limsup_{n\to\infty} P_{n, h} \phi_n \le \alpha, \quad h\in H_{0},\qquad  \text{ and }\qquad \liminf_{n\to\infty}  P_{n, h} \phi_n \ge \alpha, \quad h\in H_{1}.
    \end{equation*}
    Then, for any
     $h\in H$,
    \begin{equation}\label{thm:two-sided-pwr-bound:eq:power-bound}
        \limsup_{n\to\infty}  P_{n, h} \phi_n \le 1 - \Phi\left(z_{\alpha/2} - \tilde{\mathcal{I}}_{\gamma}^{1/2}\tau\right)  + 1- \Phi\left(z_{\alpha/2} + \tilde{\mathcal{I}}_{\gamma}^{1/2}\tau
\right),
    \end{equation}
    where $z_{\alpha}$ is the $1-\alpha$ quantile and $\Phi$ the CDF of the standard normal distribution.
\end{theorem}



Theorem \ref{thm:psi-pwr-local-alt} implies that the  power bound of Theorem \ref{thm:two-sided-pwr-bound} is achieved by the test $\psi_{n, \theta_0}$ provided $\Sigma_{21}V^\dagger \Sigma_{21}^\prime = \tilde{\mathcal{I}}_{\gamma}$ and $r=1$.

\begin{corollary}\label{cor:psi-two-sided}
    Suppose that Assumptions \ref{ass:LAN}, \ref{ass:joint-conv} and \ref{ass:consistent}
     hold with $\Sigma_{21}V^\dagger\Sigma_{21}^\prime = \tilde{\mathcal{I}}_{\gamma}$ and $r=1$.
    Then, for $h\in H$,
    \begin{equation}\label{eq:two-sided-power-bound}
        \lim_{n\to\infty} P_{n, h}\psi_{n, \theta_0}   =
            1 - \Phi\left(z_{\alpha/2} -\tilde{\mathcal{I}}_{\gamma}^{1/2}\tau\right)  + 1- \Phi\left(z_{\alpha/2} +\tilde{\mathcal{I}}_{\gamma}^{1/2}\tau\right).
    \end{equation}
\end{corollary}


\subsection{Tests of a multivariate parameter}

When $d_\theta >1$ there is an intermediate case where $0<\operatorname{rank}(\tilde{\mathcal{I}}_{\gamma})<d_\theta$. Here I permit $0<\operatorname{rank}(\tilde{\mathcal{I}}_{\gamma})\le d_\theta$ and establish  a maximin power bound for (potentially) non -- regular models, which contains the (regular) full rank case as a special case.\footnote{Section \ref{sm:ssec:stringent} shows that the most stringent test (in the sense of \citealp{W43}) in the limit experiment has the same power function as the maximin test, and no sequence of asymptotically level $\alpha$ tests can correspond to a test in the limit experiment with smaller regret.}





\begin{theorem}\label{thm:maximin-pwr-bound}
    Suppose that Assumptions \ref{ass:LAN} and \ref{ass:IP} hold and $r \coloneqq \operatorname{rank}(\tilde{\mathcal{I}}_{\gamma}) \ge 1$. Let $\phi_n:\mathcal{W}_n \to [0, 1]$ be a sequence of tests
     such that for each $h = (0, b)\in H_{0}$
        \begin{equation*}
            \limsup_{n\to\infty} P_{n, h} \phi_n \le \alpha
        \end{equation*}
    Let $c_r$ the $1-\alpha$ quantile of a $\chi^2_r$ random variable.
    Then, if $a\ge0$,
    \begin{equation}\label{thm:maximin-pwr-bound:eq:bound}
        \limsup_{n\to\infty} \inf \left\{P_{n, h}\phi_n :h = (\tau, b)\in H,\ \tau^\prime\tilde{\mathcal{I}}_{\gamma}\tau \ge a\right\} \le 1 - \mathrm{P}(\chi_r^2(a) \le c_r).
    \end{equation}
\end{theorem}





By Theorem \ref{thm:psi-pwr-local-alt}, the power bound on the right hand side of \eqref{thm:maximin-pwr-bound:eq:bound} is achieved by $\psi_{n, \theta_0}$ provided $\Sigma_{21}V^\dagger \Sigma_{21}^\prime = \tilde{\mathcal{I}}_{\gamma}$ and $\operatorname{rank}(V) = \operatorname{rank}(\tilde{\mathcal{I}}_{\gamma}) = r\ge1$.
In order that the test be asymptotically maximin over a compact subset $K_a$ of $\{h = (\tau, b)\in H: \tau^\prime \tilde{\mathcal{I}}_{\gamma} \tau \ge a\}$, with $a = \inf\{\tau^\prime \tilde{\mathcal{I}}_{\gamma} \tau = a: h\in K_a\}$, some uniformity is required.\footnote{  The pseudometric $d$ in Corollary \ref{cor:psi-maximin}
need not be related to the seminorm $\|\cdot\|_K$.}


\begin{corollary}\label{cor:psi-maximin}
    Suppose that Assumptions \ref{ass:LAN}, \ref{ass:joint-conv} and \ref{ass:consistent}
    hold with $ \Sigma_{ 21} V^{\dagger} \Sigma_{ 21}^\prime = \tilde{\mathcal{I}}_{\gamma}$ and $r = \operatorname{rank}(\tilde{\mathcal{I}}_{\gamma}) = \operatorname{rank}(V)\ge 1$.
    Then for $h = (\tau, b)\in H$
     \begin{equation}\label{cor:psi-maximin:eq-1}
        \lim_{n\to\infty} P_{n,  h}\psi_{n, \theta_0} = 1 - \mathrm{P}\left(\chi^2_r\left(a\right) \le c_r\right), \quad a = \tau^\prime \tilde{\mathcal{I}}_{\gamma} \tau.
     \end{equation}
     Additionally, suppose that $(H, d)$ is a (pseudo-)metric space and let $K_a$ be a compact subset of $\{h = (\tau, b)\in H: \tau^\prime \tilde{\mathcal{I}}_{\gamma} \tau \ge a\}$ such that $a = \inf\{\tau^\prime \tilde{\mathcal{I}}_{\gamma} \tau: h = (\tau, b)\in K_a\}$. If the functions $h\mapsto P_{n,  h}\psi_{n, \theta_0}$ are asymptotically equicontinuous on $K_a$,
     \begin{equation}\label{cor:psi-maximin:eq-2}
        \lim_{n\to\infty} \inf_{h\in K_a} P_{n,  h}\psi_{n, \theta_0} = 1 - \mathrm{P}\left(\chi^2_r\left(a\right) \le c_r\right).
     \end{equation}
\end{corollary}

A sufficient condition for the asymptotic equicontinuity required for the second part of Corollary \ref{cor:psi-maximin} based on an asymptotic equicontinuity in total variation requirement was given as Lemma \ref{lem:level-alpha-unif-equicontinuity-TV} in the previous section.\footnote{As noted there Lemma \ref{lem:level-alpha-unif-equicontinuity-weak} provides some weaker sufficient conditions; in the present context, condition \ref{lem:level-alpha-unif-equicontinuity-weak.itm.Lambda} of Lemma \ref{lem:level-alpha-unif-equicontinuity-weak} is not required (cf. Remark \ref{rem:level-alpha-unif-equicontinuity-weak-condition-split}).}














\subsection{The degenerate case}

If the efficient information matrix $\tilde{\mathcal{I}}_{\gamma}$ is zero, no test with correct asymptotic size has non -- trivial asymptotic power against any sequence of local alternatives.

\begin{proposition}\label{prop:asy-size-alpha-and-r-0-imply-no-power}
    Suppose Assumptions \ref{ass:LAN} and \ref{ass:IP} hold and $r \coloneqq \operatorname{rank}(\tilde{\mathcal{I}}_{}) = 0$. Let $\phi_n:\mathcal{W}_n \to [0, 1]$ be a sequence of tests such that $\limsup_{n\to\infty} P_{n, h} \phi_n \le \alpha$ for each $h = (0, b)\in H_{0}$
    Then, for $h\in H$, $\limsup_{n\to\infty} P _{n, h} \phi_n \le \alpha$.
\end{proposition}

\subsection{Discussion of the power bounds}

There are a number of important aspects to highlight regarding the interpretation of the power bounds obtained in the preceding subsections.

\paragraph{Optimality in multivariate testing problems}


Just as in the classical finite -- dimensional case, the multivariate optimality results just presented
should not be taken in an absolute sense. Nevertheless they seem reasonable if the researcher does not have directions against which they wish to direct power a priori. If there are alternatives of particular interest, one could construct a locally regular test by utilising the same moment conditions $g_{n}$ but weighting them differently  \cite[cf.][]{BRS06}.


\paragraph{The intermediate case with $1\le\operatorname{rank}(\tilde{\mathcal{I}}_{\gamma})<d_\theta$}

A key benefit of the multivariate power results is that they apply equally to non-regular models, i.e. cases where $\tilde{\mathcal{I}}_{\gamma}$ is rank deficient. This scenario can occur for various reasons. Firstly the model may not identify all parameters of interest $\theta$ (i.e. underidentification). Secondly some of the elements of $\theta$ may be weakly identified (i.e. weak underidentification). The power results above apply in either of these cases.

There are a number of other papers which provide inference results in similarly rank deficient settings \citep[e.g.][]{RCBR00, HM19, AG19, ABS23}; none of these papers consider optimal testing.



\paragraph{Alternative approximations}

In the case where $\operatorname{rank}(\tilde{\mathcal{I}}_{\gamma}) = 0$, Proposition \ref{prop:asy-size-alpha-and-r-0-imply-no-power} reveals that the LAN approximation in Assumption \ref{ass:LAN} is, in a certain sense, the wrong approximation: it does not provide any useful way of (asymptotically) comparing tests. Other approximations might provide valuable comparisons. Alternative approximations have been explored in, for example, the IV model \cite[e.g.][]{M09} and semiparametric GMM models by \cite{AM22, AM23}. For example, in the IV case \cite{M09} considers alternatives which are at a fixed distance from the true parameter, rather than in a shrinking $\sqrt{n}$-neighbourhood. Whether such an approach can be developed for the class of models considered here is an interesting question for future work.

\subsection{Attaining the power bounds}\label{ssec:attainment-pwr-bounds}

Provided that the $L_2$ distance between $g_{n}$ and $\tilde{\ell}_{n}$ (as defined in \eqref{eq:effscr}) vanishes, $\psi_{n, \theta_0}$ attains the power bounds established in the preceding subsections. For regular models this result is well known in two special cases: (a) the i.i.d. case (cf. Section 25.6 in \cite{vdV98}; Lemma \ref{lem:effscr-coincides-with-usual-defn}) and (b) when the information operator $\mathsf{B}$ in Remark \ref{rem:IP-linear-operator} is positive-definite with $\mathsf{B}_{22}$, the information operator for $\eta$, boundedly invertible \citep{CHS96}. Here I provide a general version of this result which applies to both regular and non-regular models and does not require (a) or (b). In particular, I show that $\Sigma_{21}V^\dagger \Sigma_{21}^\prime = \tilde{\mathcal{I}}_{\gamma}$, which suffices given Theorem \ref{thm:psi-pwr-local-alt} and the power bounds in Theorems \ref{thm:two-sided-pwr-bound} and \ref{thm:maximin-pwr-bound}.

\begin{theorem}\label{thm:pwr-bounds-effscr}
    Suppose that Assumptions \ref{ass:LAN}, \ref{ass:joint-conv}, \ref{ass:consistent} and \ref{ass:IP} hold and $g_n$ is such that $\lim_{n\to\infty} \int \|g_{n}  - \tilde{\ell}_{n}\|^2\,\mathrm{d}P_n = 0$. Then $\Sigma_{ 21} = V = \tilde{\mathcal{I}}_{\gamma}$, hence $\Sigma_{ 21}V^\dagger \Sigma_{21}^\prime = \tilde{\mathcal{I}}_{\gamma}$.
\end{theorem}































\section{The smooth i.i.d. case}\label{ssec:smooth-iid}

In this section I give conditions which  are sufficient for some of the foregoing Assumptions in the in the benchmark case for semiparametric theory: where the observations are i.i.d. and the model is ``smooth''.


\begin{assumption}[Product measures]\label{ass:iid}
    Suppose $W^{(n)} = (W_1, \ldots, W_n)\in \prod_{i=1}^n\mathcal{W} = \mathcal{W}_n$
     and that each $P_{n, h}$ is a product measure: $P_{n,h} = P_{h}^n$. Each probability measure in $\mathcal{P}_n$ is dominated by the $n$-fold product of a $\sigma$-finite measure $\nu$.
\end{assumption}


In the i.i.d. setting, it is well known that quadratic mean differentiability of the square root of the density $p = \frac{\mathrm{d}P}{\mathrm{d}\nu}$ of $P \coloneqq P_0$ is sufficient for LAN. In particular, if
\begin{equation}\label{eq:dqm}
    \lim_{n\to\infty} \int \left[\sqrt{n}\left(\sqrt{p_{h_n}} - \sqrt{p}\right) - \frac{1}{2}A h\sqrt{p}\right]^2 = 0,
\end{equation}
for a measurable $Ah: \mathcal{W}\to \mathbb{R}$, then with $\Delta_{n}h \coloneqq \frac{1}{\sqrt{n}} \sum_{i=1}^n [Ah](W_i)$ the remainder term $R_{n}$ in the LAN expansion satisfies $R_{n}(h_n)\xrightarrow{P}0$ \cite[e.g.][Lemma 3.10.11]{vdVW96}. This can be used to establish either the LAN expansion required by Assumption \ref{ass:LAN} by taking $h_n = h$ for each $n\in\mathbb{N}$ or the ULAN expansion as in Assumption \ref{ass:ULAN} by considering sequences $h_n\to h$. Sufficient conditions for \eqref{eq:dqm} are well known \cite[e.g.][Lemma 7.6]{vdV98}.

In this case, the scores $Ah$ typically take the form
\begin{equation}\label{eq:Agamh}
    [Ah](W_i) = \tau^\prime \dot{\ell}_{\gamma} (W_i) + [Db](W_i), \quad  h = (\tau, b)\in H,
\end{equation}
where $\dot{\ell}_{\gamma}$ is a $d_\theta$-vector of functions in $L_2^0(P)$ (typically the partial derivatives of $\theta \mapsto \log p_{\gamma}$ at $\gamma$) and $D:\overline{\operatorname{lin}}\  B \to L_2^0(P)$ a bounded linear map. Showing that \eqref{eq:dqm} holds (with $h_n =h$) is typically the most straightforward way to verify the LAN expansion required by Assumption \ref{ass:LAN}. If $A: \overline{\operatorname{lin}}\  H \to L_2(P)$ is a bounded linear map, then the remainder of Assumption \ref{ass:LAN} also follows directly.\footnote{A version of Lemma \ref{lem:iid-LAN-DQM} for ULAN (Assumption \ref{ass:ULAN}) is Lemma \ref{lem:iid-ULAN-DQM} in the supplementary material.}


\begin{lemma}\label{lem:iid-LAN-DQM}
    Suppose that Assumption \ref{ass:iid} holds and for each $h\in H$ equation \eqref{eq:dqm} holds (with $h_n=  h$) with $A: \overline{\operatorname{lin}}\  H \to L_2(P)$ a bounded linear map. Then Assumption \ref{ass:LAN} holds with $P_{n, h} = P_{h/\sqrt{n}}^n$ and
    $[\Delta_{n} h](W^{(n)}) = \mathbb{G}_n A h$.
\end{lemma}








When the data are i.i.d., the the joint convergence of $(\Delta_{n}h, g_{n}^\prime)$ as in Assumption \ref{ass:joint-conv} is particularly straightforward to verify. As noted in the discussion around \eqref{eq:orth-proj-g}, the required orthogonality condition can be ensured by performing an orthogonal projection. Assumption \ref{ass:joint-conv} then follows straightforwardly.
In the i.i.d. setting typically $g_{n}$ will have the form $g_{n}(W^{(n)})
 = \mathbb{G}_n g$.



\begin{lemma}\label{lem:iid-orthocomp-joint-conv}
    Suppose that Assumptions \ref{ass:LAN}  and \ref{ass:iid} hold, with $\Delta_{n} h  = \frac{1}{\sqrt{n}}\sum_{i=1}^n Ah$, where $A h$ is as in equation \eqref{eq:Agamh}. Additionally suppose that
    $g \in \left\{Db : b\in B\right\}^\perp\subset L_2^0(P)$. Then Assumption \ref{ass:joint-conv} holds with $g_{n}(W^{(n)})\coloneqq \mathbb{G}_n g$.
\end{lemma}
\begin{corollary}\label{cor:iid-orth-projection-joint-conv}
    In the setting of Lemma \ref{lem:iid-orthocomp-joint-conv}, if $f\in L_2^0(P)$ and $g$ is the orthogonal projection $g = \Pi[f | \{Db : b\in B\}^\perp ]$
    then Assumption \ref{ass:joint-conv} holds with $g_{n}(W^{(n)})\coloneqq \mathbb{G}_n g$.
\end{corollary}















































































































\section{Examples}\label{sec:examples}

I now illustrate the application of the theoretical results to the single index and IV models and conduct simulation studies to investigate finite sample performance of the proposed approach.
In this section I work under high level conditions to avoid repeating standard regularity conditions; lower level sufficient conditions are given in section \ref{sm:sec:examples-detail} of the supplementary material.


\subsection{Single index model}\label{ssec:sim}

Consider the single index model of Example \ref{ex:SIM-running-example}. I now formalise the development given in Section \ref{sec:heuristic}.
 The model parameters are $\gamma = (\theta, \eta)$ where $\eta= (f, \zeta)$ and the density of one observation with respect to a $\sigma$-finite measure $\tilde{\nu}$ is $p_{\gamma}$ as in \eqref{eq:SIM-running-example-dens}; $P_{\gamma}$ denotes the corresponding probability measure.
The parameters $\gamma$ are restricted by the following Asssumption. Let $\mathscr{X}$ be the support of $X$, $\mathscr{D}$ a convex open set containing $\{x_1 + x_2^\prime \theta: \theta\in \Theta,x\in \mathscr{X}\}$ and $C_b^1(\mathscr{D})$ the class of real functions which are bounded and continuously differentiable with bounded derivative on $\mathscr{D}$.
\begin{assumption}\label{ass:sim-mdl}
    The parameters $\gamma = (\theta, f, \zeta)\in \Gamma = \Theta\times \mathscr{F} \times \mathscr{Z}$ where $\Theta$ is an open subset of $\mathbb{R}^{d_\theta}$, $\mathscr{F} = C_b^1(\mathscr{D})$ and  $\zeta\in \mathscr{Z}$, for
    \begin{equation*}
                \mathscr{Z}\coloneqq \left\{\zeta\in L_1(\mathbb{R}^{1+K}, \nu): \zeta\ge 0,\, \int_{\mathbb{R}\times \mathscr{X}}\zeta\,\mathrm{d}\nu = 1,\,
                  \text{if } (\epsilon, X)\sim \zeta \text { then } \eqref{ass:sim-parameters:eq:sim-moments} \text{ holds}\right\},
        \end{equation*}
        with $L_1(A, \nu)$ is the space of $\nu$ -- integrable functions on $A$ and
        \begin{equation}\label{ass:sim-parameters:eq:sim-moments}
            \mathbb{E}[\epsilon|X] = 0,
            \  \mathbb{E}[\epsilon^2] < \infty,\,
            \  \mathbb{E}[(|\epsilon|^{2+\rho} + |\phi(\epsilon, X)|^{2+\rho} + 1)\|X\|^{2+\rho}] < \infty,\,
            \  \mathbb{E}[XX^\prime] \succ 0,
        \end{equation}
        for $\phi(\epsilon, X)$ the derivative of $e\mapsto \log \zeta(e, X)$.
    Additionally, for each $\gamma\in \Gamma$, $p_{\gamma}$ is a probability density with respect to some $\sigma$-finite measure $\tilde{\nu}$.
\end{assumption}
That $p_{\gamma}$ is a valid probability density holds automatically (with $\tilde{\nu} = \nu$) when $\epsilon|X$ is continuously distributed.

\paragraph*{Local Asymptotic Normality}
Consider local perturbations $P_{\gamma + \varphi_n(h)}$ for
\begin{equation}\label{eq:sim-varphin}
    \varphi_n(h) = \left(\frac{\tau}{\sqrt{n}},\ \varphi_{n, 2}(b_1, b_2)\right), \qquad h = (\tau, b_1, b_2)\in H = \mathbb{R}^{d_\theta}\times B_{1}\times B_{2}.
\end{equation}
$B_{1}$ is the set which indexes the perturbations to $f$ and consists of a subset of the continuously differentiable functions $b_1:\mathscr{D}\to\mathbb{R}$.
$B_{2}$ indexes the perturbations to $\zeta$ and consists of a subset of the functions $b_2:\mathbb{R}^{1+K}\to \mathbb{R}$ which are continuously differentiable in their first argument and satisfy
\begin{equation}
    \label{eq:sim-b2-conditions}
    \mathbb{E}[b_2(\epsilon, X)] = 0,\  \mathbb{E}[\epsilon b_2(\epsilon, X)|X] = 0,\  \mathbb{E}[b_2(\epsilon, X)^2] < \infty \quad \text{ for } (\epsilon, X)\sim \zeta.
\end{equation}
The precise form of $\varphi_{n, 2}$ is left unspecified. It is required only that the local perturbations satisfy the LAN property below.\footnote{Examples of $\varphi_{n, 2}$ and $B$ for which Assumption \ref{ass:sim-LAN} holds are given in Section \ref{app:ssec:sim-LAN}. }

\begin{assumption}\label{ass:sim-LAN}
    Suppose that $\mathcal{W}_n = \prod_{i=1}^n \mathbb{R}^{1+K}$ and
    $P_{n,h}\coloneqq P_{\gamma + \varphi_n(h)}^n \ll \nu_n$ for all $\gamma\in \Gamma$ and $h\in H$ and are such that Assumption \ref{ass:LAN} holds with
    \begin{equation}
        \log \frac{p_{n, h}}{p_{n,  0}} = \frac{1}{\sqrt{n}}\sum_{i=1}^n [Ah](W_i) - \frac{1}{2}\sigma(h) + o_{P_{n, 0}}(1), \quad h\in H_,
    \end{equation}
    where $\sigma(h) = \int [Ah]^2\,\mathrm{d}P$ and $A$ is as in equation \eqref{eq:Agamh} with
    \begin{align*}
        \dot{\ell}(W) &\coloneqq -\phi(Y - f(V_\theta), X) f'(V_\theta)X_2\\
        [D b](W) &\coloneqq -\phi(Y - f(V_\theta), X)b_{1}(V_{\theta}) + b_2(Y - f(V_\theta), X).
    \end{align*}
\end{assumption}









\paragraph*{The moment conditions}\label{para:sim:g}
The test statistic is based on $g_{n}\coloneqq \mathbb{G}_n g$ for $g$ given in \eqref{ex:SIM-running-example:eq:g}. This satisfies Assumption \ref{ass:joint-conv} under  Assumptions \ref{ass:sim-mdl} \& \ref{ass:sim-LAN} and \eqref{eq:sim:eps-phi--1} below.


\begin{proposition}\label{prop:sim-joint-conv}
    Suppose Assumptions \ref{ass:sim-mdl} \& \ref{ass:sim-LAN} hold and under $P$,
    \begin{equation}\label{eq:sim:eps-phi--1}
        \mathbb{E}[\epsilon^2|X] \le C < \infty,\quad \mathbb{E}\left[\epsilon\phi(\epsilon, X)| X\right] = -1, \quad \text{a.s.}~.
    \end{equation}
    Then Assumption \ref{ass:joint-conv} holds with $g_{n}\coloneqq \mathbb{G}_n g$ for $g$ given in \eqref{ex:SIM-running-example:eq:g}.
\end{proposition}

\paragraph*{A feasible test}\label{para:sim:ghat}




To form a feasible test $g_n$ must be replaced by an estimator $\hat{g}_{n, \theta}$. Let this have the form $\hat{g}_{n, \theta}(W^{(n)}) \coloneqq \frac{1}{\sqrt{n}}\sum_{i=1}^n \hat{g}_{n, \theta, i}$, for $\hat{g}_{n, \theta, i}$ defined as in \eqref{ex:SIM-running-example:eq:ghati}.
To keep the notation concise let $Z_{3}\coloneqq f$, $Z_4\coloneqq f'$,
$Z_0 \coloneqq Z_{1} / Z_2$ and correspondingly $\hat{Z}_{0,n, i}\coloneqq \hat{Z}_{1, n, i} / \hat{Z}_{2, n, i}$.
Let $\check{V}_{n, \theta}\coloneqq \frac{1}{n}\sum_{i=1}^n \hat{g}_{n, \theta, i}\hat{g}_{n, \theta, i}^\prime$ and
If $V$ is known to have full rank then let $\hat{V}_{n, \theta} \coloneqq \check{V}_{n, \theta}$, $\hat{\Lambda}_{n, \theta} \coloneqq \hat{V}_{n, \theta}^{-1}$ and $\hat{r}_{n, \theta} = \operatorname{rank}(V)$.
Form the estimator $\hat{V}_{n, \theta}$ according to the construction in Section S5 of \cite{LM21-S} using a truncation rate $\upnu_n$. $\hat{\Lambda}_{n, \theta}$ is then taken to be $\hat{V}_{n, \theta}^\dagger$ and $\hat{r}_{n, \theta}\coloneqq \operatorname{rank}(\hat{V}_{n, \theta})$.
Under the following condition, these estimators satisfy the conditions of Assumption \ref{ass:consistent}.
\begin{assumption}\label{ass:sim-estimation-HL}
    Suppose that equation
    \eqref{eq:sim:eps-phi--1} holds (under $P$), $X$ has compact support,
    $\mathbb{E}[\epsilon^4]<\infty$, and with $P$ probability approaching one $\mathsf{R}_{l, n, i} \le r_n = o(n^{-1/4})$,
    \begin{equation*}
        \mathsf{R}_{l, n, i} \coloneqq \left[\int \left\|\hat{Z}_{l, n, i}(v) - Z_l(v)\right\|^2 \,\mathrm{d}\mathcal{V}(v)\right]^{1/2}, \quad l = 1,\ldots 4,
    \end{equation*}
    where $\mathcal{V}$ is the law of $V_{\theta}$ under $P$ and where $\hat{Z}_{l, n, i}(V_{\theta, i})$ is $\sigma(\{V_{\theta, i}\} \,\cup\, \mathcal{C}_{n, j})$ measurable with $j=1$ if $i >  \lfloor n/2\rfloor$ and 2 otherwise, with $\mathcal{C}_{n, 1}\coloneqq \{W_j : j\in \{1, \ldots, \lfloor n/2\rfloor\}\}$ and $\mathcal{C}_{n, 2}\coloneqq \{W_j : j\in \{\lfloor n/2\rfloor + 1, \ldots, n\}\}$.
\end{assumption}

The rate conditions in Assumption \ref{ass:sim-estimation-HL} can be satisfied by  e.g. (sample -- split) series estimators under standard conditions; see e.g. \cite{BCCK15}.

\begin{proposition}\label{prop:sim-estimation-HL}
    Suppose Assumptions \ref{ass:sim-mdl}, \ref{ass:sim-LAN} and \ref{ass:sim-estimation-HL} hold and $\upnu_n$ is such that $r_n = o(\upnu_n)$.
    Then Assumption \ref{ass:consistent} holds with $V\coloneqq \int gg^\prime\,\mathrm{d}P$.
\end{proposition}

A consequence of Assumption \ref{ass:sim-LAN} and Propositions \ref{prop:sim-joint-conv} and \ref{prop:sim-estimation-HL} is that the test $\psi_{n, \theta_0}$ formed as in \eqref{eq:psi-test} is locally regular by Theorem \ref{thm:psi-pwr-local-alt}.



\paragraph*{Simulation study}

I take $K=1$ and test $H_0: \theta=\theta_0 = 1$ at a nominal level of 5\%. Each study reports the results of 5000 monte carlo replications with a sample size of $n\in \{400, 600, 800\}$.
I report empirical rejection frequencies for the $\psi_{n, \theta_0}$ test along with a Wald test based on an estimator in the style of \cite{I93}.

I consider two different classes of link function. The first sets $f(v) =f_j(v) = 5\exp (-v^2 / 2c_j^2)$ (``exponential''); the second $f(v) =f_j(v) = 25\left(1 + \exp(-v  /c_j)\right)^{-1}$ (``logistic'').  The values of $c_j$ considered are recorded in Table \ref{tbl:SIM-index-fcns}.

\begin{table*}[!htbp]
	\footnotesize
	\begin{center}
		\caption{\label{tbl:SIM-index-fcns} Functions used in the simulation experiments}
		\begin{threeparttable}

\begin{tabular}{llrrr}
\toprule
name & expression & $c_1$ & $c_2$ & $c_3$\\
\midrule
Exponential & $f_j(v) = 5 \exp (-v^2 / 2c_j^2)$ & 1 & 2 & 4\\
Logistic & $f_j(v) = 25(1 + \exp (- v / c_j))^{-1}$ & 1 & 4 & 32\\
\bottomrule
\end{tabular}

		\end{threeparttable}
	\end{center}
\end{table*}

In each case, as $c_j$ increases, the derivative of $f$ flattens out, moving towards a point with $f'=0$, at which $\theta$ is unidentified.\footnote{The functions $f$ and $f'$ are plotted in  Figures \ref{fig:SIM-gaussian-fs} and \ref{fig:SIM-logistic-fs}.}
I draw covariates as $X= (Z_1, 0.2 Z_1 + 0.4 Z_2 + 0.8)$, where each $Z_k\sim U(-1, 1)$ is independent. The error term is drawn either as $\epsilon = \upsilon / \sqrt{3/2}$ with $\upsilon \sim t(6)$ (``homoskedastic'') or $\epsilon \sim \mathcal{N}(0,  1 + \sin(X_1)^2)$ (``heteroskedastic'').

I compute the test $\psi_{n, \theta_0}$ as described on p. \pageref{para:sim:ghat}, with $\upomega(X) = 1$. The functions $f,\, f'$ and $Z_{1}$ are estimated via sample split smoothing splines.\footnote{I use the base \textsf{R} function
\texttt{smooth.spline} with 20 knots. In this setting $Z_2(V_\theta) = 1$ is known.} The truncation parameter $\upnu$ is set to $10^{-3}$. I additionally compute a Wald test in the style of \cite{I93}, using the same non-parametric estimates as for $\hat{g}_{n, \theta}$.\footnote{Given $\hat{f}$, $\hat\theta = \operatorname*{arg\,min}_{\theta\in \Theta_\star} \frac{1}{n}\sum_{i=1}^n (Y_i - \hat{f}(V_{\theta, i}))^2$, for $\Theta_\star = [-10, 10]$. The asymptotic variance is estimated by $\hat{\sigma}^2 / \frac{1}{n}\sum_{i=1}^n \left(\widehat{f'}(V_{\hat\theta, i}) \left[X_2 - \hat{Z}(V_{\hat\theta, i})\right]\right)^2$, for $\hat{\sigma}^2 = \frac{1}{n}\sum_{i=1}^n (Y_i - \hat{f}(V_{\hat\theta, i}))^2$.}

The empirical rejection frequencies of these procedures are recorded in Table \ref{tbl:SIM-size}: $\psi_{n, \theta_0}$ rejects at close to the nominal 5\%  for all simulation designs considered whilst the Wald test over -- rejects in most of the simulation designs considered.
Figure \ref{fig:SIM-power} contains power plots of the $\psi_{n, \theta_0}$ test ($n=800$). For almost flat link functions there is very identifying information and hence very little power available. As the link function moves away from the point of identification failure ($f'=0$), the available power increases and is captured by the $\psi_{n, \theta_0}$ test.


\begin{table*}[!htbp]
	\setlength{\tabcolsep}{5pt}
	\footnotesize
	\begin{center}
		\caption{\label{tbl:SIM-size} \small ERF (\%), Single-index model}
		\begin{threeparttable}

\begin{tabular}{rrrrrrrrrrrrr}
\toprule
\multicolumn{1}{c}{} & \multicolumn{6}{c}{Exponential} & \multicolumn{6}{c}{Logistic} \\
\cmidrule(l{3pt}r{3pt}){2-7} \cmidrule(l{3pt}r{3pt}){8-13}
\multicolumn{1}{c}{} & \multicolumn{3}{c}{Homoskedastic} & \multicolumn{3}{c}{Heteroskedastic} & \multicolumn{3}{c}{Homoskedastic} & \multicolumn{3}{c}{Heteroskedastic} \\
\cmidrule(l{3pt}r{3pt}){2-4} \cmidrule(l{3pt}r{3pt}){5-7} \cmidrule(l{3pt}r{3pt}){8-10} \cmidrule(l{3pt}r{3pt}){11-13}
$n$ & $f_1$ & $f_2$ & $f_3$ & $f_1$ & $f_2$ & $f_3$ & $f_1$ & $f_2$ & $f_3$ & $f_1$ & $f_2$ & $f_3$\\
\midrule\multicolumn{7}{l}{$\psi_{n, \theta_0}$}\\
\midrule
400 & 6.08 & 5.90 & 5.30 & 5.66 & 5.68 & 5.12 & 6.36 & 5.94 & 4.80 & 5.66 & 6.00 & 4.68\\
600 & 6.12 & 5.68 & 5.26 & 5.44 & 4.76 & 4.32 & 6.04 & 5.74 & 4.10 & 5.26 & 5.18 & 4.62\\
800 & 6.20 & 6.04 & 5.28 & 5.46 & 5.72 & 5.22 & 6.00 & 5.90 & 4.26 & 5.62 & 5.16 & 4.46\\
\midrule\multicolumn{7}{l}{Wald}\\\midrule
\addlinespace
400 & 13.06 & 18.94 & 13.54 & 15.20 & 20.32 & 14.38 & 8.12 & 13.60 & 14.74 & 8.24 & 15.52 & 14.92\\
600 & 10.28 & 16.32 & 12.58 & 10.60 & 18.92 & 14.18 & 7.18 & 10.74 & 13.88 & 6.82 & 12.60 & 14.84\\
800 & 10.30 & 16.64 & 12.84 & 9.60 & 19.00 & 13.62 & 6.86 & 10.68 & 12.94 & 6.74 & 11.26 & 14.18\\
\bottomrule
\end{tabular}

		\end{threeparttable}
	\end{center}
\end{table*}


\begin{figure}[!htbp]
		\centering
		\caption{\label{fig:SIM-power} Single-index model, $\psi$ power}
		\begin{subfigure}[b]{0.495\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_exponential.pdf}
			\caption[]
			{{\small Exponential}}
		\end{subfigure}
		\hfill
		\begin{subfigure}[b]{0.495\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_logistic.pdf}
			\caption[]
			{{\small Logistic}}
		\end{subfigure}
\end{figure}














\subsection{IV model}\label{ssec:iv}

In the IV model of Example \ref{ex:IV},  $n$ i.i.d. copies of $W = (Y, X, Z)$ are observed where
\begin{equation}\label{eq:iv-mdl-0}
    Y = X^\prime \theta  +Z_1^\prime\beta + \epsilon, \qquad \mathbb{E}[\epsilon |Z] = 0, \qquad Z = (Z_1^\prime, Z_2^\prime)^\prime.
\end{equation}
Let $d_W\coloneqq d_{\theta} + d_Z + 1$.
With $\pi(Z)\coloneqq \mathbb{E}[X|Z]$ and $\upsilon = X - \pi(Z)$,
\begin{equation}\label{eq:iv-mdl}
    \begin{aligned}
        Y &= X^\prime \theta + Z_1^\prime\beta   + \epsilon\\
        X &= \pi(Z) + \upsilon
    \end{aligned},\qquad   \mathbb{E}[U|Z] = 0, \quad   U \coloneqq (\epsilon, \upsilon^\prime)^\prime.
    \end{equation}
If $\pi(Z)$ is constant the instruments $Z$ provide no information about $\theta$.
Weak identification in this model can be very different from in the IV model with a linear first stage: there are many data configurations in which $\mathrm{Var}(\mathbb{E}[X|Z])$ may be ``large'' whilst $\mathbb{E}[XZ^\prime]\mathbb{E}[ZZ^\prime]^{-1}\approx 0$.
In such situations, tests which can exploit such non-linear identifying information can provide substantially more power than tests which (implicitly) use a linear first stage. In this section I develop a $\psi_{n, \theta_0}$ test which can capture such identifying information whilst remaining robust to weak identification.\footnote{This does not contradict optimality results that are known for, e.g., the AR test \cite[][]{M09, CHJ09} as these results assume a linear first stage.
}$^,$ \footnote{An alternative approach to capturing this non-linear identifying information (whilst remaining robust to weak instruments) is to use a large number of transformations of the instruments, $f_1(Z), \ldots, f_M(Z)$, in a linear first stage, combined with a testing procedure which remains robust in the presence of many weak instruments. In the simulation study below, I compare the $\psi_{n, \theta_0}$ test to this approach, using the test of \cite{MS21}.}



Let $\zeta$ denote the density of $\xi \coloneqq (\epsilon, \upsilon^\prime, Z^\prime)$ with respect to a $\sigma$-finite measure $\nu$.
The parameters of the IV model are $\gamma = (\theta, \eta)$ with the nuisance parameters collected in $\eta = (\beta, \pi, \zeta)$. The density of one observation is
\begin{equation}\label{eq:iv-dens}
    p_{\gamma}(W) = \zeta(Y - X^\prime\theta - Z_1^\prime \beta, X - \pi(Z), Z),
\end{equation}
with respect to a $\sigma$-finite measure $\tilde{\nu}$ and $P_{\gamma}$ denotes the corresponding measure. The model parameters are restricted as follows.

\begin{assumption}\label{ass:iv-mdl}
    The parameters $\gamma = (\theta, \beta, \pi, \zeta)\in \Gamma = \Theta \times \mathcal{B}\times \mathscr{P}\times \mathscr{Z}$ where
    \begin{enumerate}
        \item $\Theta$ is an open subset of $\mathbb{R}^{d_\theta}$ and $\mathcal{B}$ is an open subset of $\mathbb{R}^{d_\beta}$;
        \item $\mathscr{Z}$ is a subset of the set of density functions on $\mathbb{R}^{d_W}$ with respect to $\nu$;
        \item For $(\pi, \zeta)\in \mathscr{P}\times \mathscr{Z}$, if $\xi\coloneqq (U^\prime, Z^\prime)^\prime$, then
        \begin{equation*}
            \mathbb{E}[U | Z] = 0,\qquad \mathbb{E}\|\xi\|^4 < \infty,\qquad  \mathbb{E}\|\pi(Z)\|^4 < \infty,\qquad \mathbb{E}\|\phi(\xi)\|^4 < \infty,
        \end{equation*}
        where $\phi_1 \coloneqq \nabla_{\epsilon}\log \zeta(\epsilon, \upsilon, Z)$, $\phi_2\coloneqq \nabla_{\upsilon}\log \zeta(\epsilon, \upsilon, Z)$ and $\phi \coloneqq(\phi_1, \phi_2^\prime)^\prime$.
    \end{enumerate}
    Additionally,  $p_{\gamma}$ is a probability density for each $\gamma\in \Gamma$ with respect to a $\sigma$-finite measure $\tilde{\nu}$.
\end{assumption}


Assumption \ref{ass:iv-mdl} imposes the existence of certain moments and the (IV) conditional mean restriction. That $p_{\gamma}$ is a valid probability density holds automatically (with $\nu = \tilde{\nu}$) when $U|Z$ is continuously distributed.


\paragraph*{Local Asymptotic Normality}
Consider local perturbations $P_{\gamma + \varphi_n(h)}$ for
\begin{equation}\label{eq:iv-varphi}
    \varphi_n(h)\coloneqq \left(\frac{\tau}{\sqrt{n}},\,  \frac{b_0}{\sqrt{n}}, \varphi_{n, 1}(b_1),\, \varphi_{n, 2}(b_2)\right)~, \quad h = (\tau, b)\in H\coloneqq \mathbb{R}^{d_\theta}\times B,
\end{equation}
with $ B \coloneqq \mathbb{R}^{d_\beta} \times B_{1}\times B_{2}$. $B_{1}$ is a subset of the bounded functions $b_1:\mathbb{R}^{d_Z}\to \mathbb{R}^{d_\theta}$ and $B_{2}$ a subset of the functions $b_2:\mathbb{R}^{d_W}\to \mathbb{R}$ which are bounded and continuously differentiable in its first $1 + d_\theta$ components with bounded derivative and such that
\begin{equation}\label{eq:iv-b2-conditions}
    \mathbb{E}\left[b_2(U, Z)\right] = 0, \quad \mathbb{E}\left[U b_2(U, Z) | Z\right] = 0, \qquad \text{ for } (U^\prime, Z^\prime)^\prime  \sim \zeta.
\end{equation}
The precise forms of $\varphi_{n, 1}, \varphi_{n, 2}$ are left unspecified. It is required only that the local perturbations satisfy LAN.\footnote{Examples of $\varphi_{n, 1}, \varphi_{n, 2}$ and $B$ for which Assumption \ref{ass:iv-LAN} holds are given in Section \ref{app:ssec:iv-LAN}. }

\begin{assumption}\label{ass:iv-LAN}
    Suppose that $\mathcal{W}_n = \prod_{i=1}^n \mathbb{R}^{d_W}$,
    $P_{n, h}\coloneqq P_{\gamma + \varphi_n(h)}^n \ll \nu_n$ for all $\gamma\in \Gamma$ and $h\in H$
    and are such that Assumption \ref{ass:LAN} holds with
    \begin{equation}
        \log \frac{p_{n, h}}{p_{n,  0}} = \frac{1}{\sqrt{n}}\sum_{i=1}^n [Ah](W_i) - \frac{1}{2}\sigma(h) + o_{P_{n, 0}}(1), \quad h\in H,
    \end{equation}
    where $\sigma(h) = \int [Ah]^2\,\mathrm{d}P$ and $A$ is as in equation \eqref{eq:Agamh} with
    \begin{align*}
        \dot{\ell}_{\gamma}(W)&\coloneqq -\phi_1(\epsilon(\theta, \beta), \upsilon(\pi),Z) X_1 \\
        [Db](W) &\coloneqq -\phi(\epsilon(\theta, \beta),\upsilon(\pi),Z)^\prime \left[\begin{smallmatrix}
            b_0^\prime Z_1 & b_1(Z)
        \end{smallmatrix}\right]
     + b_2(\epsilon(\theta, \beta), \upsilon(\pi), Z),
    \end{align*}
    where $\epsilon(\theta, \beta)\coloneqq Y - X^\prime\theta - Z_1^\prime\beta$ and $\upsilon(\pi)\coloneqq X - \pi(Z)$.
\end{assumption}

\paragraph*{The moment conditions}
The test will be based on moment conditions related to the efficient score function for $\theta$, $\tilde{\ell}_{\gamma}$. This is given in the following Lemma.\footnote{The last two conditions in \eqref{eq:iv-EphiU} hold if $\lim_{|u_i|\to\infty} |u_i|\zeta(u, z)=0$ for $i=1,
\ldots, d_\alpha$.
}


\begin{lemma}\label{lem:iv-effscr}
    Suppose Assumptions \ref{ass:iv-mdl}, \ref{ass:iv-LAN} hold, for  $J(Z)\coloneqq \mathbb{E}[UU^\prime|Z]$,
    \begin{equation}\label{eq:iv-EphiU}
        \begin{aligned}
        0<c \le \lambda_{\min}(J(Z))\le
        \lambda_{\max}(J(Z))
        \le C < \infty, &\qquad \lambda_{\min}(\mathbb{E}[Z_1Z_1^\prime] )> 0,\\
        \mathbb{E}\left[\phi(\epsilon,\upsilon, Z)U^\prime \middle| Z\right] = -I,&\qquad  \mathbb{E}[\phi_1(\epsilon, \upsilon, Z)\upsilon U^\prime]= 0,
    \end{aligned}
    \end{equation}
    and that $B_1$ is dense in $L_2$. Define $\upomega(Z)\coloneqq \mathbb{E}[\epsilon^2|Z]^{-1}$. The efficient score for $\theta$ is
    \begin{equation}\label{eq:iv-effscr}
        \tilde{\ell}_{\gamma}(W) = \upomega(Z)(Y - X^\prime\theta - Z_1^\prime\beta)\left[\pi(Z) - \mathbb{E}[\upomega(Z)XZ_1^\prime]\mathbb{E}[\upomega(Z)Z_1Z_1^\prime]^{-1}Z_1\right].
    \end{equation}
\end{lemma}
For simplicity, I will use the moment functions
\begin{equation}\label{eq:iv-bar-effscr}
    g(W)\coloneqq  \mathbb{E}[\epsilon^2]^{-1}(Y - X^\prime\theta - Z_1^\prime\beta)\left[\pi(Z) - \mathbb{E}[XZ_1^\prime]\mathbb{E}[Z_1Z_1^\prime]^{-1}Z_1\right].
\end{equation}
$g$ belongs to the orthocomplement of $\{Db: b\in B\}$ and coincides with the efficient score function when $J(Z) = \mathbb{E}[UU^\prime]$ a.s. (i.e. under homoskedasticity).\footnote{Nevertheless, homoskedasticity is \emph{not} assumed and the results below hold under heteroskedasticity. For full efficiency one could base the test on \eqref{eq:iv-effscr}. This is left for future work.
}

\begin{lemma}\label{lem:iv-bar-effscr}
    Suppose that Assumptions \ref{ass:iv-mdl}, \ref{ass:iv-LAN} and equation \eqref{eq:iv-EphiU} hold. Then, the moment conditions $g \in \{Db: b\in B\}^\perp$. If $\mathbb{E}[\epsilon^2|Z] = \mathbb{E}[\epsilon^2]$ a.s., then $g = \tilde{\ell}_{\gamma}$ a.s..
\end{lemma}


\begin{proposition}\label{prop:iv-joint-conv}
    Suppose that Assumptions \ref{ass:iv-mdl}, \ref{ass:iv-LAN} and equation \eqref{eq:iv-EphiU} hold. Then Assumption \ref{ass:joint-conv} is satisfied with
    $g_{n} = \mathbb{G}_n g$.
\end{proposition}

\paragraph*{A feasible test}


Suppose that $\hat{\beta}_n$ and $\hat{\pi}_{n, i}(Z_i)$ are estimators of $\beta$ and $\pi(Z_i)$ respectively. Let the $i$-th residual in \eqref{eq:iv-mdl-0} based on $\theta = \theta_0$ and $\hat\beta_n$ be $\hat{\epsilon}_{n, i} \coloneqq Y_i - X_i^\prime \theta - Z_{1, i}^\prime\hat{\beta}_n$. Let $\hat{s}_n\coloneqq \frac{1}{n}\sum_{i=1}^n \hat{\epsilon}_{n, i}^2$ and define

\begin{equation}\label{eq:iv-bar-effscr-est}
   \hat{g}_{n, \theta, i}\coloneqq \hat{s}_{n}^{-1}\hat{\epsilon}_{n, i} \left[\hat{\pi}_n(Z_i)  - \left[\frac{1}{n}\sum_{i=1}^n X_i Z_{1, i}^\prime\right] \left[\frac{1}{n}\sum_{i=1}^n Z_{1,i} Z_{1, i}^\prime\right]^{-1} Z_i\right],
\end{equation}
and
\begin{equation}\label{eq:iv-bar-effinfo-est-check}
    \check{V}_{n, \theta} \coloneqq \frac{1}{n}\sum_{i=1}^n \hat{g}_{n, \theta, i}\hat{g}_{n, \theta, i}^\prime.
\end{equation}

Based on $\check{V}_{n, \theta}$, form $\hat{V}_{n, \theta}$ according to the construction in Section S5 of \cite{LM21-S} using a truncation rate $\upnu_n$, set $\hat{\Lambda}_{n, \theta}\coloneqq \hat{V}_{n, \theta}^\dagger$ and $\hat{r}_{n, \theta}\coloneqq \operatorname{rank}(\hat{V}_{n, \theta})$.







The following assumption provides sufficient high-level conditions on the estimators $\hat\beta_n$ and $\hat{\pi}_{n, i}(Z_i)$ such that Assumption \ref{ass:consistent} holds. These conditions are compatible with $\hat{\pi}_{n, i}$ being a leave-one-out series estimator.\footnote{The discretisation of $\hat\beta_n$ is a technical device
which permits the proof to go through under weaker conditions \cite[cf.][Chapter 6]{LCY00}. This
can be arranged given a $\sqrt{n}$ -- consistent initial estimator, by replacing its value with the closest point in the set $\mathscr{S}_n$.}$^,\,$\footnote{See e.g. \cite{BCCK15} for sufficient conditions for \eqref{ass:iv-est2:eq:pi-est-bound} and Section \ref{sm:sec:IV} for a discussion of \eqref{ass:iv-est2:eq:pi-est-bound2}.}

\begin{assumption}\label{ass:iv-est2}
    Suppose that, given $\theta_0$, (i) $\hat{\beta}_n$ is an estimator valued in $\mathscr{S}_n\coloneqq \{CZ / \sqrt{n}: Z\in \mathbb{Z}^{d_{\beta}}\}$ for some $C\in \mathbb{R}^{d_\beta\times d_\beta}$ and satisfying  $\sqrt{n}(\hat{\beta}_n - \beta) = O_{P_{n, 0}}(1)$ and (ii) $\hat{\pi}_{n, i}(Z_i)$ are estimators such that $\hat{\pi}_{n, i}(Z_i)$ is $\sigma(Z_i, \mathcal{C}_{n,-i})$ measurable for $\mathcal{C}_{n,-i}\coloneqq \{W_j: j=1, \ldots, n,\, j\neq i\}$, and on events $F_n$ with $P_{n, 0}(F_n)\to 1$,
    \begin{equation}\label{ass:iv-est2:eq:pi-est-bound}
        \left[\int \left\|\hat{\pi}_{n, i}(z) - \pi(z)\right\|^2 \,\mathrm{d}\zeta_Z(z)\right]^{1/2} \le \delta_n = o(1),
    \end{equation}
    where $\zeta_Z$ is the marginal distribution of $Z$ and for each $k=1, \ldots, d_\theta$, $i\neq j$,
    \begin{equation}\label{ass:iv-est2:eq:pi-est-bound2}
        \mathbb{E}\left[\bm{1}_{F_n}\bm{1}_{G_n}(\hat{\pi}_{n, i, k}(Z_i) - \pi_k(Z_i))(\hat{\pi}_{n, j, k}(Z_j) - \pi_k(Z_j))^\prime\epsilon_i\epsilon_j\right]\lesssim \delta_n^2 / n, \  P_{n, 0}(G_n)\to 1.
    \end{equation}
    Suppose also $\delta_n^2 + n^{-1/2}  = o(\upnu_n)$, \eqref{eq:iv-EphiU} holds and $\mathbb{E}\left[\epsilon^4(\|\pi(Z)\| + \|Z_1\|)^4\right] < \infty$.
\end{assumption}

There is no requirement on the rate $\delta_n$ in \eqref{ass:iv-est2:eq:pi-est-bound}, \eqref{ass:iv-est2:eq:pi-est-bound2} beyond $\delta_n = o(1)$.


\begin{proposition}\label{prop:iv-estimation-bar-effscr2}
    Suppose that Assumptions \ref{ass:iv-mdl}, \ref{ass:iv-LAN}, \& \ref{ass:iv-est2} hold. Then Assumption \ref{ass:consistent} holds with $\hat{g}_{n, \theta} \coloneqq \frac{1}{\sqrt{n}}\sum_{i=1}^n \hat{g}_{n, \theta, i}$, $g_{n} \coloneqq \mathbb{G}_n g$ and $\hat{\Lambda}_{n, \theta}$ defined below equation \eqref{eq:iv-bar-effinfo-est-check}.
\end{proposition}

A consequence of Assumption \ref{ass:iv-LAN} and Propositions \ref{prop:iv-joint-conv} and \ref{prop:iv-estimation-bar-effscr2} is that the test $\psi_{n, \theta_0}$ formed as in \eqref{eq:psi-test} is locally regular by Theorem \ref{thm:psi-pwr-local-alt}.


\paragraph{Simulation study}
I test $H_0: \theta=\theta_0 = 0$ at a nominal level of 5\%. Each study reports the results of 5000 monte carlo replications with a sample size of $n\in \{200, 400, 600\}$.\footnote{The power surfaces in Design 1 are computed with 2500 replications.} Two simulation designs are considered.

Design 1 is a bivariate, just identified design. Here $d_\theta = 2$ and $Z_2$ is drawn from a zero-mean multivariate normal distribution with covariance matrix $\left[\begin{smallmatrix}
    1 & 0.4\\
    0.4 & 1
\end{smallmatrix}\right]$. The error terms $\epsilon, \upsilon$ are drawn from a zero-mean multivariate normal such that each has variance 1 and the covariances are $\mathrm{Cov}(\epsilon, \upsilon_i) = 0.9$ and $\mathrm{Cov}(\upsilon_1,\upsilon_2) = 0.7$. $Z_1=1$ with $\beta = 1$ and $\pi(Z) = \pi(Z_2) = (\pi_1(Z_{2,1}), \pi_2(Z_{2, 2}))^\prime$ with each $\pi_i$ ($i=1, 2$) being one of the exponential or logistic functions $f_j$ in Table \ref{tbl:SIM-index-fcns}. The exponential form is a prototypical function shape for which the linear projection of $X$ on $Z$ will provide essentially no identifying information; for the logistic form this linear projection should perform well.\footnote{These functions are plotted in Figures \ref{fig:SIM-gaussian-fs} and \ref{fig:SIM-logistic-fs}. The separation  $\pi(Z_2) = (\pi_1(Z_{2,1}), \pi_2(Z_{2, 2}))^\prime$ is assumed unknown and is not imposed in the estimation of $\pi$.
}$^,\, $\footnote{Results for the case where $\pi(Z_2)$ is linear are very similar to the ``approximately linear'' logistic case and are therefore unreported.}

I consider the $\psi_{n, \theta_0}$ test developed above, with a leave-one-out series estimator of $\pi$ based on (tensor product) Legendre polynomials. I consider both  fixing the number of polynomials at $k=3$ in each of the univariate series which form the tensor product basis and choosing $k\in \{3, 4, 5, 6, 7\}$ using information criteria. $\upnu$ is set to $10^{-2}$.
I additionally consider the \cite{AR49} (AR) test.\footnote{The AR test is computed with $Z_2$ as instruments, after partialling out $Z_1$.
}$^,$\footnote{I do not consider alternative weak instrument robust tests based on a linear first stage (e.g. LM, CLR) in this design as the AR test is known to be optimal when the model is just-identified.}

The empirical rejection frequencies under the null are shown in Tables \ref{tbl:IV-size1i} and \ref{tbl:IV-size1ii}. The parameter $j$ controls the level of identification: the larger is $j$ the closer $\pi_j$ is to a constant function. In each specification all the considered tests reject at close to the nominal level. Power surfaces for the $\psi_{n, \theta_0}$ and AR tests are shown in figures \ref{fig:IV-power-1-exp-exp-AR} -- \ref{fig:IV-power-1-exp-log-psi}. As can be seen in these figures, the $\psi_{n, \theta_0}$ test is able to detect deviations from the null when $\pi_j$ has the exponential form, unlike the AR test.\footnote{
One could consider an AR test using e.g. some basis functions $f_1(Z_{2}), \ldots, f_K(Z_{2})$ however as noted in \cite[][p. 2669]{MS21}, the AR statistic is not well behaved for large $K$. The jackknife AR test of \cite{MS21} applies only to the case where $d_\theta=1$.
} For the logistic form, the power of the two tests is similar. Unsurprisingly, neither test provides non-trivial power when identification is very weak.


\begin{table*}[!htbp]
	\footnotesize
	\setlength{\tabcolsep}{4.5pt}
	\begin{center}
		\caption{\label{tbl:IV-size1i} Empirical rejection frequencies, IV, Design 1}
		\begin{threeparttable}

\begin{tabular}{rrrrrrrrrrrrrr}
\toprule
\multicolumn{1}{c}{} & \multicolumn{1}{c}{} & \multicolumn{4}{c}{Exponential - Exponential} & \multicolumn{4}{c}{Logistic - Logistic} & \multicolumn{3}{c}{Exponential - Logistic} \\
\cmidrule(l{3pt}r{3pt}){3-6} \cmidrule(l{3pt}r{3pt}){7-10} \cmidrule(l{3pt}r{3pt}){11-13}
\multicolumn{1}{c}{} & \multicolumn{1}{c}{} & \multicolumn{1}{c}{AR} & \multicolumn{3}{c}{$\psi$} & \multicolumn{1}{c}{AR} & \multicolumn{3}{c}{$\psi$} & \multicolumn{1}{c}{AR} & \multicolumn{3}{c}{$\psi$} \\
\cmidrule(l{3pt}r{3pt}){3-3} \cmidrule(l{3pt}r{3pt}){4-6} \cmidrule(l{3pt}r{3pt}){7-7} \cmidrule(l{3pt}r{3pt}){8-10} \cmidrule(l{3pt}r{3pt}){11-11} \cmidrule(l{3pt}r{3pt}){12-14}
$n$ & $j$ &  & $k=3$ & AIC & BIC &  & $k=3$ & AIC & BIC &  & $k=3$ & AIC & BIC\\
\midrule
200 & 1 & 5.52 & 5.10 & 4.00 & 5.18 & 5.52 & 4.88 & 3.40 & 4.48 & 5.52 & 4.56 & 3.76 & 4.68\\
200 & 2 & 5.52 & 6.24 & 6.26 & 6.24 & 5.52 & 5.46 & 5.30 & 5.46 & 5.52 & 5.78 & 5.86 & 5.78\\
200 & 3 & 5.52 & 8.36 & 7.80 & 8.36 & 5.52 & 8.44 & 7.98 & 8.44 & 5.52 & 7.92 & 7.74 & 7.92\\
\addlinespace
400 & 1 & 5.60 & 4.96 & 4.48 & 4.80 & 5.60 & 4.96 & 4.14 & 4.78 & 5.60 & 4.88 & 4.34 & 4.80\\
400 & 2 & 5.60 & 6.12 & 6.20 & 6.12 & 5.60 & 5.68 & 5.52 & 5.68 & 5.60 & 6.08 & 6.02 & 6.08\\
400 & 3 & 5.60 & 6.76 & 6.94 & 6.76 & 5.60 & 7.72 & 7.60 & 7.72 & 5.60 & 6.46 & 6.60 & 6.46\\
\addlinespace
600 & 1 & 5.38 & 5.30 & 4.90 & 5.22 & 5.38 & 5.20 & 3.80 & 5.16 & 5.38 & 5.32 & 4.64 & 5.04\\
600 & 2 & 5.38 & 6.08 & 6.10 & 6.08 & 5.38 & 4.98 & 4.98 & 4.98 & 5.38 & 5.56 & 5.76 & 5.56\\
600 & 3 & 5.38 & 4.36 & 4.78 & 4.36 & 5.38 & 5.02 & 5.50 & 5.02 & 5.38 & 4.40 & 5.06 & 4.40\\
\bottomrule
\end{tabular}

			\begin{tablenotes}
			\footnotesize
			\item \justify{\footnotesize{\textit{Notes:} E.g. ``Exponential - Logistic'' indicates that $\pi_1, \pi_2$ have the exponential and logistic form in Table \ref{tbl:SIM-index-fcns} respectively, with $c_j$ corresponding to column $j$.
			}}
			\end{tablenotes}
		\end{threeparttable}
	\end{center}
\end{table*}

\begin{table*}[!htbp]
	\setlength{\tabcolsep}{4.5pt}
	\footnotesize
	\begin{center}
		\caption{\label{tbl:IV-size1ii} Empirical rejection frequencies, IV, Design 1}
		\begin{threeparttable}

\begin{tabular}{rlrrrrrrrrrrrr}
\toprule
\multicolumn{1}{c}{} & \multicolumn{1}{c}{} & \multicolumn{4}{c}{Exponential - Exponential} & \multicolumn{4}{c}{Logistic - Logistic} & \multicolumn{4}{c}{Exponential - Logistic} \\
\cmidrule(l{3pt}r{3pt}){3-6} \cmidrule(l{3pt}r{3pt}){7-10} \cmidrule(l{3pt}r{3pt}){11-14}
\multicolumn{1}{c}{} & \multicolumn{1}{c}{} & \multicolumn{1}{c}{AR} & \multicolumn{3}{c}{$\psi$} & \multicolumn{1}{c}{AR} & \multicolumn{3}{c}{$\psi$} & \multicolumn{1}{c}{AR} & \multicolumn{3}{c}{$\psi$} \\
\cmidrule(l{3pt}r{3pt}){3-3} \cmidrule(l{3pt}r{3pt}){4-6} \cmidrule(l{3pt}r{3pt}){7-7} \cmidrule(l{3pt}r{3pt}){8-10} \cmidrule(l{3pt}r{3pt}){11-11} \cmidrule(l{3pt}r{3pt}){12-14}
$n$ & $j_1-j_2$ &  & $k=3$ & AIC & BIC &  & $k=3$ & AIC & BIC &  & $k=3$ & AIC & BIC\\
\midrule
200 & 1 - 3 & 5.52 & 6.36 & 5.30 & 6.48 & 5.52 & 7.18 & 6.40 & 7.12 & 5.52 & 6.34 & 5.40 & 6.30\\
200 & 2 - 3 & 5.52 & 6.26 & 6.44 & 6.26 & 5.52 & 6.94 & 6.78 & 6.94 & 5.52 & 6.46 & 6.64 & 6.46\\
200 & 3 - 3 & 5.52 & 8.36 & 7.80 & 8.36 & 5.52 & 8.44 & 7.98 & 8.44 & 5.52 & 7.92 & 7.74 & 7.92\\
\addlinespace
400 & 1 - 3 & 5.60 & 5.70 & 5.38 & 5.68 & 5.60 & 6.24 & 5.82 & 6.20 & 5.60 & 6.14 & 5.54 & 5.98\\
400 & 2 - 3 & 5.60 & 6.22 & 6.48 & 6.22 & 5.60 & 6.24 & 6.40 & 6.24 & 5.60 & 6.74 & 6.90 & 6.74\\
400 & 3 - 3 & 5.60 & 6.76 & 6.94 & 6.76 & 5.60 & 7.72 & 7.60 & 7.72 & 5.60 & 6.46 & 6.60 & 6.46\\
\addlinespace
600 & 1 - 3 & 5.38 & 5.66 & 5.40 & 5.60 & 5.38 & 5.46 & 4.52 & 5.34 & 5.38 & 5.58 & 5.26 & 5.36\\
600 & 2 - 3 & 5.38 & 6.00 & 6.10 & 6.00 & 5.38 & 5.40 & 5.40 & 5.40 & 5.38 & 6.12 & 6.32 & 6.12\\
600 & 3 - 3 & 5.38 & 4.36 & 4.78 & 4.36 & 5.38 & 5.02 & 5.50 & 5.02 & 5.38 & 4.40 & 5.06 & 4.40\\
\bottomrule
\end{tabular}

			\begin{tablenotes}
			\footnotesize
			\item \justify{\footnotesize{\textit{Notes:} E.g. ``Exponential - Logistic'' indicates that $\pi_1, \pi_2$ have the exponential and logistic form in Table \ref{tbl:SIM-index-fcns} respectively, with $c_{j_1}$ and $c_{j_2}$ corresponding to column $j_1$ - $j_2$.
			}}
			\end{tablenotes}
		\end{threeparttable}
	\end{center}
\end{table*}


\begin{figure}[!htbp]
		\centering
		\caption{\label{fig:IV-power-1-exp-exp-AR} IV Design 1, AR power, $\pi_i$ exponential}
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_ar_1_1_exponential_exponential.pdf}
			\caption[]
			{{\small $j_1=1$, $j_2=1$}}
		\end{subfigure}
		\hfill
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_ar_1_3_exponential_exponential.pdf}
			\caption[]
			{{\small $j_1=1$, $j_2=3$}}
		\end{subfigure}
		\hfill
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_ar_3_3_exponential_exponential.pdf}
			\caption[]
			{{\small $j_1=3$, $j_2=3$}}
		\end{subfigure}
\end{figure}

\begin{figure}[!htbp]
		\centering
		\caption{\label{fig:IV-power-1-exp-exp-psi} IV Design 1, $\psi$ $(k=3)$ power, $\pi_i$ exponential}
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_effscr_k_1_1_exponential_exponential.pdf}
			\caption[]
			{{\small $j_1=1$, $j_2=1$}}
		\end{subfigure}
		\hfill
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_effscr_k_1_3_exponential_exponential.pdf}
			\caption[]
			{{\small $j_1=1$, $j_2=3$}}
		\end{subfigure}
		\hfill
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_effscr_k_3_3_exponential_exponential.pdf}
			\caption[]
			{{\small $j_1=3$, $j_2=3$}}
		\end{subfigure}
\end{figure}


\begin{figure}[!htbp]
		\centering
		\caption{\label{fig:IV-power-1-log-log-AR} IV Design 1, AR power, $\pi_i$ logistic}
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_ar_1_1_logistic_logistic.pdf}
			\caption[]
			{{\small $j_1=1$, $j_2=1$}}
		\end{subfigure}
		\hfill
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_ar_1_3_logistic_logistic.pdf}
			\caption[]
			{{\small $j_1=1$, $j_2=3$}}
		\end{subfigure}
		\hfill
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_ar_3_3_logistic_logistic.pdf}
			\caption[]
			{{\small $j_1=3$, $j_2=3$}}
		\end{subfigure}
\end{figure}


\begin{figure}[!htbp]
		\centering
		\caption{\label{fig:IV-power-1-log-log-psi} IV Design 1, $\psi$ $(k=3)$ power, $\pi_i$ logistic}
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_effscr_k_1_1_logistic_logistic.pdf}
			\caption[]
			{{\small $j_1=1$, $j_2=1$}}
		\end{subfigure}
		\hfill
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_effscr_k_1_3_logistic_logistic.pdf}
			\caption[]
			{{\small $j_1=1$, $j_2=3$}}
		\end{subfigure}
		\hfill
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_effscr_k_3_3_logistic_logistic.pdf}
			\caption[]
			{{\small $j_1=3$, $j_2=3$}}
		\end{subfigure}
\end{figure}


\begin{figure}[!htbp]
		\centering
		\caption{\label{fig:IV-power-1-exp-log-AR} IV Design 1, AR power, $\pi_1$ exponential, $\pi_2$ logistic}
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_ar_1_1_exponential_logistic.pdf}
			\caption[]
			{{\small $j_1=1$, $j_2=1$}}
		\end{subfigure}
		\hfill
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_ar_1_3_exponential_logistic.pdf}
			\caption[]
			{{\small $j_1=1$, $j_2=3$}}
		\end{subfigure}
		\hfill
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_ar_3_3_exponential_logistic.pdf}
			\caption[]
			{{\small $j_1=3$, $j_2=3$}}
		\end{subfigure}
\end{figure}


\begin{figure}[!htbp]
		\centering
		\caption{\label{fig:IV-power-1-exp-log-psi} IV Design 1, $\psi$ $(k=3)$ power, $\pi_1$ exponential, $\pi_2$ logistic}
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_effscr_k_1_1_exponential_logistic.pdf}
			\caption[]
			{{\small $j_1=1$, $j_2=1$}}
		\end{subfigure}
		\hfill
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_effscr_k_1_3_exponential_logistic.pdf}
			\caption[]
			{{\small $j_1=1$, $j_2=3$}}
		\end{subfigure}
		\hfill
		\begin{subfigure}[b]{0.3\textwidth}
			\centering
			\includegraphics[width=0.99\textwidth]{power_1_effscr_k_3_3_exponential_logistic.pdf}
			\caption[]
			{{\small $j_1=3$, $j_2=3$}}
		\end{subfigure}
\end{figure}

Design 2 is a univariate, over identified model with heteroskedastic errors. $Z_1$, $\beta$ and $Z_2$ are as in Design 1, and $\pi(Z_2) = (\pi_1(Z_{2, 1}) + \pi_2(Z_{2, 2}))/2$ where the $\pi_i$ have one of the exponential or logistic forms of Table \ref{tbl:SIM-index-fcns}.\footnote{This functional form is treated as unknown and not imposed in the estimation of $\pi$.} I draw $(\tilde{\epsilon}, \tilde{\upsilon})$ from a zero-mean multivariate normal distribution with unit variances,  covariance $0.95$ and set $(\epsilon, \upsilon)^\prime = \left[\begin{smallmatrix}
    \sqrt{1 + \sin(Z_{2, 1})^2} & 0\\
    0 & \sqrt{1 + \cos(Z_{2, 2})^2}
\end{smallmatrix}\right](\tilde{\epsilon}, \tilde{\upsilon})^\prime$.

The $\psi$ tests are computed in the same manner as in Design 1. I also compute the AR, LM and CLR tests based on $Z_2$ (with $Z_1$ partialled out) as well as the many weak instrument robust jackknife AR test of \cite{MS21}. MS$_1$ uses $Z_2$ as instruments; MS$_2$ uses the (tensor product) of Legendre polynomials used to estimate $\pi$ as instruments.

The empirical rejection frequencies under the null are shown in Tables \ref{tbl:IV-size2i} -- \ref{tbl:IV-size2iii}. As in Design 1, the parameter $j$ controls the level of identification: the larger is $j$ the closer $\pi_j$ is to a constant function and hence $\theta$ unidentified. In each specification all the considered tests reject close to the nominal level; the MS$_2$ test is somewhat oversized for smaller $n$.
The power of these tests is plotted in Figures \ref{fig:IV-power-2-exp-exp} -- \ref{fig:IV-power-2-exp-log}; the $\psi_{n, \theta_0}$ tests are  denoted by $k=3$, AIC and BIC, corresponding to how $\pi$ is estimated. For the design with both $\pi_i$ exponential, the $\psi_{n, \theta_0}$ test clearly delivers the highest power whenever there is non-trivial power available; of the other tests, only MS$_2$ delivers non-trivial power in this specification. For the case with both $\pi_i$ logistic, all tests except MS$_2$ perform similarly, with MS$_2$ offering lower power. The same holds for the final specification, where $\pi_1$ is exponential and $\pi_2$ logistic.



\begin{table*}[!htbp]
		\footnotesize
	\begin{center}
		\caption{\label{tbl:IV-size2i} Empirical rejection frequencies, IV, Design 2, Exponential - Exponential}
		\begin{threeparttable}

\begin{tabular}{rrrrrrrrrr}
\toprule
\multicolumn{1}{c}{} & \multicolumn{1}{c}{} & \multicolumn{1}{c}{AR} & \multicolumn{1}{c}{LM} & \multicolumn{1}{c}{CLR} & \multicolumn{1}{c}{MS$_1$} & \multicolumn{1}{c}{MS$_2$} & \multicolumn{3}{c}{$\psi$} \\
\cmidrule(l{3pt}r{3pt}){3-3} \cmidrule(l{3pt}r{3pt}){4-4} \cmidrule(l{3pt}r{3pt}){5-5} \cmidrule(l{3pt}r{3pt}){6-6} \cmidrule(l{3pt}r{3pt}){7-7} \cmidrule(l{3pt}r{3pt}){8-10}
$n$ & $j$ &  &  &  &  &  & $k=3$ & AIC & BIC\\
\midrule
200 & 1 & 5.68 & 5.52 & 5.94 & 7.56 & 9.24 & 6.36 & 5.74 & 6.36\\
200 & 2 & 5.68 & 6.08 & 5.88 & 7.56 & 9.24 & 7.62 & 7.62 & 7.62\\
200 & 3 & 5.68 & 6.46 & 6.16 & 7.56 & 9.24 & 7.36 & 7.44 & 7.36\\
\addlinespace
400 & 1 & 5.28 & 5.48 & 5.26 & 7.18 & 8.78 & 6.32 & 5.38 & 6.36\\
400 & 2 & 5.28 & 5.72 & 5.48 & 7.18 & 8.78 & 7.98 & 7.98 & 7.98\\
400 & 3 & 5.28 & 6.06 & 5.84 & 7.18 & 8.78 & 3.42 & 3.68 & 3.42\\
\addlinespace
600 & 1 & 5.86 & 5.48 & 6.04 & 7.58 & 7.92 & 5.16 & 5.16 & 5.14\\
600 & 2 & 5.86 & 5.50 & 5.86 & 7.58 & 7.92 & 6.28 & 6.20 & 6.28\\
600 & 3 & 5.86 & 5.76 & 6.30 & 7.58 & 7.92 & 1.32 & 1.74 & 1.32\\
\bottomrule
\end{tabular}

			\begin{tablenotes}
			\footnotesize
			\item \justify{\footnotesize{\textit{Notes:} The functions $\pi_i$ have the exponential form in Table \ref{tbl:SIM-index-fcns} with $c_j$ corresponding to column $j$.
			}}
			\end{tablenotes}
		\end{threeparttable}
	\end{center}
\end{table*}

\begin{table*}[!htbp]
		\footnotesize
	\begin{center}
		\caption{\label{tbl:IV-size2ii} Empirical rejection frequencies, IV, Design 2, Logistic - Logistic}
		\begin{threeparttable}

\begin{tabular}{rrrrrrrrrr}
\toprule
\multicolumn{1}{c}{} & \multicolumn{1}{c}{} & \multicolumn{1}{c}{AR} & \multicolumn{1}{c}{LM} & \multicolumn{1}{c}{CLR} & \multicolumn{1}{c}{MS$_1$} & \multicolumn{1}{c}{MS$_2$} & \multicolumn{3}{c}{$\psi$} \\
\cmidrule(l{3pt}r{3pt}){3-3} \cmidrule(l{3pt}r{3pt}){4-4} \cmidrule(l{3pt}r{3pt}){5-5} \cmidrule(l{3pt}r{3pt}){6-6} \cmidrule(l{3pt}r{3pt}){7-7} \cmidrule(l{3pt}r{3pt}){8-10}
$n$ & $j$ &  &  &  &  &  & $k=3$ & AIC & BIC\\
\midrule
200 & 1 & 5.68 & 5.78 & 6.94 & 7.56 & 9.24 & 4.82 & 4.46 & 4.80\\
200 & 2 & 5.68 & 5.94 & 7.22 & 7.56 & 9.24 & 5.76 & 5.74 & 5.76\\
200 & 3 & 5.68 & 5.82 & 6.86 & 7.56 & 9.24 & 7.68 & 7.78 & 7.68\\
\addlinespace
400 & 1 & 5.28 & 4.76 & 5.98 & 7.18 & 8.78 & 4.32 & 4.04 & 4.18\\
400 & 2 & 5.28 & 4.80 & 6.26 & 7.18 & 8.78 & 4.90 & 4.98 & 4.90\\
400 & 3 & 5.28 & 4.94 & 5.80 & 7.18 & 8.78 & 3.18 & 3.52 & 3.18\\
\addlinespace
600 & 1 & 5.86 & 5.12 & 6.42 & 7.58 & 7.92 & 4.84 & 4.50 & 4.66\\
600 & 2 & 5.86 & 5.20 & 6.32 & 7.58 & 7.92 & 5.00 & 4.90 & 5.00\\
600 & 3 & 5.86 & 5.14 & 6.28 & 7.58 & 7.92 & 1.56 & 1.84 & 1.56\\
\bottomrule
\end{tabular}

			\begin{tablenotes}
			\footnotesize
			\item \justify{\footnotesize{\textit{Notes:} The functions $\pi_i$ have the logistic form in Table \ref{tbl:SIM-index-fcns} with $c_j$ corresponding to column $j$.
			}}
			\end{tablenotes}
		\end{threeparttable}
	\end{center}
\end{table*}

\begin{table*}[!htbp]
		\footnotesize
	\begin{center}
		\caption{\label{tbl:IV-size2iii} Empirical rejection frequencies, IV, Design 2, Exponential - Logistic}
		\begin{threeparttable}

\begin{tabular}{rrrrrrrrrr}
\toprule
\multicolumn{1}{c}{} & \multicolumn{1}{c}{} & \multicolumn{1}{c}{AR} & \multicolumn{1}{c}{LM} & \multicolumn{1}{c}{CLR} & \multicolumn{1}{c}{MS$_1$} & \multicolumn{1}{c}{MS$_2$} & \multicolumn{3}{c}{$\psi$} \\
\cmidrule(l{3pt}r{3pt}){3-3} \cmidrule(l{3pt}r{3pt}){4-4} \cmidrule(l{3pt}r{3pt}){5-5} \cmidrule(l{3pt}r{3pt}){6-6} \cmidrule(l{3pt}r{3pt}){7-7} \cmidrule(l{3pt}r{3pt}){8-10}
$n$ & $j$ &  &  &  &  &  & $k=3$ & AIC & BIC\\
\midrule
200 & 1 & 5.68 & 6.14 & 7.12 & 7.56 & 9.24 & 5.28 & 4.62 & 5.30\\
200 & 2 & 5.68 & 5.92 & 7.18 & 7.56 & 9.24 & 6.90 & 6.90 & 6.90\\
200 & 3 & 5.68 & 5.90 & 6.60 & 7.56 & 9.24 & 7.16 & 7.26 & 7.16\\
\addlinespace
400 & 1 & 5.28 & 4.94 & 6.44 & 7.18 & 8.78 & 4.74 & 4.28 & 4.70\\
400 & 2 & 5.28 & 5.24 & 6.54 & 7.18 & 8.78 & 5.52 & 5.64 & 5.52\\
400 & 3 & 5.28 & 5.00 & 5.66 & 7.18 & 8.78 & 3.10 & 3.42 & 3.10\\
\addlinespace
600 & 1 & 5.86 & 5.20 & 6.30 & 7.58 & 7.92 & 4.86 & 4.56 & 5.00\\
600 & 2 & 5.86 & 5.14 & 6.48 & 7.58 & 7.92 & 5.62 & 5.50 & 5.62\\
600 & 3 & 5.86 & 5.22 & 6.24 & 7.58 & 7.92 & 1.20 & 1.58 & 1.20\\
\bottomrule
\end{tabular}

			\begin{tablenotes}
			\footnotesize
			\item \justify{\footnotesize{\textit{Notes:} $\pi_1$, $\pi_2$ have the exponential and logistic form in Table \ref{tbl:SIM-index-fcns} respectively with $c_{j_i}$ corresponding to column $j$.
			}}
			\end{tablenotes}
		\end{threeparttable}
	\end{center}
\end{table*}

\begin{figure}[htbp]
	\begin{minipage}{.99\textwidth}
		\centering
		\caption{\label{fig:IV-power-2-exp-exp} \small IV design 2, Power curves, exponential $\pi_i$ }
		\includegraphics[scale = 0.45]{power_2_exp_exp.pdf}
		\end{minipage}
\end{figure}
\begin{figure}[htbp]
	\begin{minipage}{.99\textwidth}
		\centering
		\caption{\label{fig:IV-power-2-log-log} \small IV Design 2, Power curves, logistic $\pi_i$ }
		\includegraphics[scale = 0.45]{power_2_log_log.pdf}
		\end{minipage}
\end{figure}

\begin{figure}[htbp]
	\begin{minipage}{.99\textwidth}
		\centering
		\caption{\label{fig:IV-power-2-exp-log} \small IV Design 2, Power curves, exponential $\pi_1$, logistic $\pi_2$ }
		\includegraphics[scale = 0.45]{power_2_exp_log.pdf}
		\end{minipage}
\end{figure}

































































































\section{Empirical applications}\label{sec:ea}

In this section I re-analyse two IV studies with potentially weak instruments by inverting the $\psi_{n, \theta_0}$ test developed in section \ref{ssec:iv} to construct weak-instrument robust confidence intervals (CIs). In each case the $\psi_{n, \theta_0}$ test is able to exploit non-linearities to yield substantial reductions in CI length relative to AR CIs.

\subsection{The effect of skilled immigration on productivity}
\cite{H14} studies the long term effect of skilled immigration on productivity using a natural experiment in which the skilled but religiously persecuted French protestants (Hugenots) fled and settled in Prussia.\footnote{That the Hugenots were more skilled (on average) than the Prussian population seems to be broadly accepted cf. pp. 85-86, 93-95 in \cite{H14}.}
In the notation of Example \ref{ex:IV}, $Y$ is log output in textile manufacturing, $X$ is the proportion of Hugenots in each town and $Z_1$ contains various control variables, see \cite{H14} for details. \cite{H14} argues that ``by the order of centralized ruling by the king and his agents Huguenots were channeled into Prussian towns in order to compensate for severe population losses during the Thirty Years' War'', motivating the instrument $Z_2$: the percentage population losses during the war. In particular, three different measurements of this population loss are used and refered to as specifications (1), (2) and (3) hereafter.\footnote{Specifications (1) \& (2) are those considered in the left and and right hand parts of Table 4 of \cite{H14}; Specification (3) is that considered in the left hand part of Table 5.}

As noted in \cite{H14}, this instrument may be weak: in each case the first stage F statistic is ``small''.\footnote{It is less than the cutoff of 10 suggested by \cite{SS97} for homoskedastic IV.}
I implement the $\psi_{n, \theta_0}$ test using a leave-one-out series estimator of $\pi$ of the form
\begin{equation}\label{eq:pi-hat-ea}
    \hat{\pi}_{i}(Z_i) = \hat{\pi}_{i}^\prime p_{K}(Z_{2,i}) + \hat{\beta}_i Z_{1, i},
\end{equation}
where the $i$ subscript on the estimated coefficients indicates they have been estimated on all observations except for the $i$-th. $p_{K}$ is a vector of a constant and the first $K$ Legendre polynomials. I choose $K\in \{1, 2, 3, 4\}$ and whether to include $Z_1$ in the model for $\pi$ by using BIC: all specifications include $Z_1$ and $K=4$.

Table \ref{tbl:H14-1} reports 2SLS estimates of $\theta$ along with 2SLS (Wald) CIs, AR CIs and CIs found by inverting the $\psi_{n, \theta_0}$ test.\footnote{The inversion is performed over a grid of 5000 equally spaced points from -1 to 7.}  The resulting CIs provide a similar interpretation as that based on the AR CIs: the effect of (skilled) Hugenot immigration was positive on textile output. However, the $\psi_{n, \theta_0}$ based CIs are smaller than the AR CIs, achieving approximately a 40\% - 50\% reduction in length.

\begin{table*}[!htbp]
    \small
    	\begin{center}
		\caption{\label{tbl:H14-1} Point estimates and confidence intervals}
		\begin{threeparttable}

\begin{tabular}{lccc}
\toprule
 & (1) & (2) & (3)\\
\midrule
n & 150 & 150 & 186\\
F & 3.668 & 4.791 & 5.736\\
\midrule\multicolumn{4}{l}{\textit{Point estimate}}\\
2SLS & 3.475 & 3.38 & 1.671\\
\midrule\multicolumn{4}{l}{\textit{Confidence intervals}}\\
2SLS & {}[1.27, 5.68] & {}[1.294, 5.467] & {}[0.032, 3.31]\\
AR & {}[1.427, 6.303] & {}[1.43, 5.985] & {}[-0.022, 3.379]\\
$\psi$ & {}[1.626, 4.099] & {}[1.637, 4.073] & {}[1.136, 3.228]\\
\midrule\multicolumn{4}{l}{\textit{Relative length of confidence intervals to AR}}\\
2SLS & 0.904 & 0.916 & 0.964\\
$\psi$ & 0.507 & 0.535 & 0.615\\
\bottomrule
\end{tabular}

            \begin{tablenotes}
                \scriptsize
                \item \justify{\footnotesize{\textit{Notes:} F is the first stage $F$ statistic. All confidence intervals have nominal coverage of 95\%.
                }}
                \end{tablenotes}
		\end{threeparttable}
	\end{center}
\end{table*}

\subsection{The effect of racial segregation on inequality}
\cite{A11} estimates the effect of racial segregation ($X$) on poverty and inequality ($Y$, measured respectively by the poverty rate and log gini coefficient for black / white city residents), instrumenting segregation by a ``railroad division index'' (RDI, $Z_2$), ``a variation on a Herfindahl index that measures the dispersion of a city's land into subunits'' via the layout of railroad tracks.\footnote{Section 3 and Appendix A of \cite{A11} provides evidence that the choice of railroad placement was not related to local social or economic concerns.} The first and second stages also include an intercept and control for railroad track length ($Z_1$).

The instrument may be weak: I calculate the first stage F statistic to be 2.307.\footnote{\cite{A11} refers to Column 1 of Table 1 when discussing the first stage F statistic. The values in this table imply a first stage F of $(0.357 / 0.088)^2 \approx 16.458$. This ``discrepancy'' arises from different default choices of robust covariance estimate in \textsf{R}'s \texttt{sandwich} package (HC3; my calculation) and STATA's \texttt{robust} command (HC1; \citealp{A11}).} I implement the $\psi_{n, \theta_0}$ test using a leave-one-out series estimator of $\pi$ of the form \eqref{eq:pi-hat-ea}. I choose $K\in \{1, 2, 3, 4\}$ and whether to include $Z_1$ in the model for $\pi$ by using BIC: this excludes $Z_1$ and chooses $K=2$.

Table \ref{tbl:A11-1} reports 2SLS estimates of $\theta$ along with 2SLS (Wald) CIs, AR CIs and CIs found by inverting the $\psi_{n, \theta_0}$ test.\footnote{The inversion is performed over a grid of 5000 equally spaced points from -1 to 2.}  The resulting CIs provide a similar interpretation as that based on the AR CIs: racial segregation increases poverty and inequality within the Black community and decreases poverty and inequality within the White community. The $\psi_{n, \theta_0}$ based CIs are shorter than the AR CIs, achieving a reduction in length varying from 5\% to around 38\%.

\begin{table*}[!htbp]
    \small
    	\begin{center}
		\caption{\label{tbl:A11-1} Point estimates and confidence intervals}
		\begin{threeparttable}

\begin{tabular}{lcccc}
\toprule
\multicolumn{1}{c}{ } & \multicolumn{2}{c}{Poverty rate} & \multicolumn{2}{c}{Gini coefficient} \\
\cmidrule(l{3pt}r{3pt}){2-3} \cmidrule(l{3pt}r{3pt}){4-5}
  & White & Black & White & Black\\
\midrule\multicolumn{5}{l}{\textit{Point estimate}}\\
\midrule
2SLS & 0.258 & -0.196 & 0.875 & -0.334\\
\midrule\multicolumn{5}{l}{\textit{Confidence intervals}}\\
2SLS & {}[-0.026, 0.543] & {}[-0.334, -0.058] & {}[0.277, 1.474] & {}[-0.558, -0.111]\\
AR & {}[-0.04, 0.598] & {}[-0.394, -0.075] & {}[0.319, 1.684] & {}[-0.674, -0.149]\\
$\psi$ & {}[0.072, 0.568] & {}[-0.376, -0.097] & {}[0.29, 1.138] & {}[-0.608, -0.106]\\
\midrule\multicolumn{5}{l}{\textit{Relative length of confidence intervals to AR}}\\
2SLS & 0.891 & 0.866 & 0.877 & 0.853\\
$\psi$ & 0.775 & 0.877 & 0.622 & 0.955\\
\bottomrule
\end{tabular}

            \begin{tablenotes}
                \scriptsize
                \item \justify{\footnotesize{\textit{Notes:} All confidence intervals have nominal coverage of 95\%.
                }}
                \end{tablenotes}
		\end{threeparttable}
	\end{center}
\end{table*}




















\section{Conclusion}\label{sec:concl}

In this paper I establish that C($\alpha$)-style tests are locally regular under mild conditions, including in non-regular cases where locally regular estimators do not exist. As a consequence, these tests do not overreject under semiparametric weak identification asymptotics. Additionally I generalise the classical local asymptotic power bounds for LAN models to the case where the efficient information matrix has positive, but potentially deficient, rank, such that these results also apply in cases of underidentification (or weak underidentification). I show that, if the C($\alpha$) test is based on the efficient score function, it attains these power bounds. This (attainment) result improves on results known in the literature in two ways: (i) it applies also to non-regular models and (ii) it does not require the data to be i.i.d. nor the information operator to be boundedly invertible. A simulation study based on two examples shows that the asymptotic theory provides an accurate approximation to the finite sample performance of the proposed tests.










\bibliographystyle{asa}
\bibliography{Bib}