EconBase
← Back to paper

Semiparametrics via parametrics and contiguity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

113,121 characters · 13 sections · 68 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametrics via parametrics and contiguity

\thispagestyle{empty}

abstract\onehalfspacing Inference on the parametric part of a semiparametric model is no trivial task. If one approximates the infinite dimensional part of the semiparametric model by a parametric function, one obtains a parametric model that is in some sense close to the semiparametric model and inference may proceed by the method of maximum likelihood. Under regularity conditions, the ensuing maximum likelihood estimator is asymptotically normal and efficient in the approximating parametric model. Thus one obtains a sequence of asymptotically normal and efficient estimators in a sequence of growing parametric models that approximate the semiparametric model and, intuitively, the limiting {`}semiparametric{'} estimator should be asymptotically normal and efficient as well. In this paper we make this intuition rigorous: we move much of the semiparametric analysis back into classical parametric terrain, and then translate our parametric results back to the semiparametric world by way of contiguity. Our approach departs from the conventional sieve literature by being more specific about the approximating parametric models, by working not only {\it with} but also {\it under} these when treating the parametric models, and by taking full advantage of the mutual contiguity that we require between the parametric and semiparametric models. We illustrate our theory with two canonical examples of semiparametric models, namely the partially linear regression model and the Cox regression model. An upshot of our theory is a new, relatively simple, and rather parametric proof of the efficiency of the Cox partial likelihood estimator. MSC2020 subject classifications: Primary 62E20, 62G05; secondary 62G20. Keywords and phrases: Asymptotic equivalence, efficiency, Cox model, Fisher information, maximum likelihood, partially linear regression, profile likelihood, sieves

\begingroup \footnote{Adam Lee would like to thank participants at various conferences for their helpful comments. Emil A. Stoltenberg would like to the Department of Data Science at BI Norwegian Business School where the greater part of this work was carried out. Per A. Mykland would like to thank the (U.S.) National Science Foundation for financial support under grant DMS-2413952. } \addtocounter{footnote}{-1} \endgroup

Introduction

Drawing valid inference about the parametric component of a semiparametric model is often a challenging task. To deal with some of these challenges, we explore a simplifying strategy (inspired by research in high-frequency econometrics, see mykland2009inference) in which the semiparametric problem is made (almost) fully parametric. The idea is to pretend that the data stem from a parametric distribution, allowing one to derive estimators and study their properties using standard parametric techniques. We let the these parametric distributions grow in such a way that they are mutually contiguous with respect to the full semiparametric distribution; and, finally, we use contiguity and Le Cam{'}s third lemma to switch the analysis back to the semiparametric model in the limit.

The idea of approximating semiparametric models with growing parametric models is not new, as there is a well developed literature on sieves (grenander1981abstract; see shen1997methods and chenshen98 for seminal contributions to this literature, and chen07survey for an excellent review). Our framework parallels the conventional sieve literature in its use of growing parametric models, but differs by carrying out the analysis {\it under} these parametric models themselves, i.e., assuming that the data stem from a member of such a parametric model. As we shall see, combining this analysis under parametric models with the use of contiguity to return to the semiparametric setting makes our approach -- which we call contiguous sieves -- markedly different from that of conventional sieves (i.e., growing parametric models, analysed under the same semiparametric distribution at all steps). In the examples we have considered this approach allows us to establish the asymptotic normality and efficiency of semiparametric maximum likelihood estimators under conditions which are relatively user-friendly and straightforward to verify, essentially because they require just a little more than standard parametric likelihood theory. In particular, we are able to show asymptotic efficiency of semiparametric maximum likelihood estimators in a broad class of models without imposing any empirical process type conditions. Working with contiguous parametric submodels has several advantages. There is no ambiguity in defining a likelihood function (assuming the model is dominated), and based on this likelihood function, estimators can be derived or defined without difficulty. Carrying out the analysis {\it as if} these submodels were the models generating the data, enables us to work as though we have a correctly specified parametric model at each step; consequently, the required regularity conditions for these estimators to be asymptotically normal and efficient can be verified similarly as in classical parametric maximum likelihood theory. Finally, since the parametric distributions are constructed such that they are mutually contiguous to the true semiparametric distribution, we can transfer the analysis back to the semiparametric distribution in the limit, thus obtaining asymptotic results under the true distribution.

The article proceeds as follows. In Section (ref) we outline the general setup we work in and motivate our approach. Section (ref) introduces the assumptions we work under. Sections (ref) & (ref) contain our theoretical results. In Section (ref) we provide a detailed analysis of two examples to demonstrate the application of our approach: the partially linear (regression) model and the Cox model. As we exemplify with the Cox model, parametric approximations might also be used as a purely theoretical tool (and not for estimation); this allows for efficiency proofs that we believe to be simpler than those currently available. Section (ref) concludes.

General setup and parametric theory

Let $X_1, \ldots, X_n$ be i.i.d. replicates of $X$, where $X$ has distribution $P_0 \coloneqq P_{\theta_0, \eta_0}$ belonging to a semiparametric family of distributions $\mathcal{P}= \{P_{\theta, \eta}: \theta\in \Theta, \eta\in \mathcal{H}\}$, where $\Theta\subset \mathbb{R}^p$ for some $p \geq 1$, and $\mathcal{H}$ is a space of infinite dimension. We suppose that $\mathcal{P}$ is dominated by some $\sigma$-finite measure $\mu$, and write $p_{\theta,\eta}$ for the densities. The problem is to do inference on $\theta$ in the presence of the infinite dimensional nuisance parameter $\eta$. The simplifying strategy discussed in the introduction involves the construction of certain parametric submodels: For each $m \geq 1$, let $\mathcal{H}_m$ be a family of parametric functions $\eta_{\gamma}$ indexed by a parameter $\gamma \in \Gamma_m \subset \mathbb{R}^{k_m}$. Typically these shall be such that $\mathcal{H}_m \subset \mathcal{H}_{m + 1}$ for all $m$ with $\cup_{m \geq 1} \mathcal{H}_m$ dense in $\mathcal{H}$ in an appropriate topology. Let $T_m \colon \mathcal{H}_m \to \Gamma_m$ be an isomorphism between the function space $\mathcal{H}_m$ and $\Gamma^{m}$ so that $T_m \eta_{\gamma} = \gamma \in \Gamma_m$ and $T_{m}^{-1}\gamma = \eta_\gamma \in \mathcal{H}_m$. Then for each $m \geq 1$,

equation[equation omitted — 117 chars of source]

is a $p + k_m$ dimensional parametric model, and $P_{\theta,T_m^{-1}\gamma}$ has density $p_{\theta, T_m^{-1}\gamma}$ with respect to $\mu$. As discussed above, the idea is to carry out the analysis {\it as if} $X_1,\ldots,X_n$ was an i.i.d. sample from a member of $\mathcal{P}_m$, let $m$ increase with the sample size $n$, and use contiguity to switch the analysis back to $P_{\theta_0,\eta_0}$ in the limit. The natural choice for a parametric approximation is to suppose that the sample stems from $P_{m} \coloneqq P_{\theta_0, T_m^{-1}\gamma_{0}}$, where $\theta_0$ equals the true value in the big model $P_0$, and $(\gamma_{0})_{m \geq 1} = (\gamma_{0,m})_{m \geq 1}$ is a sequence of growing vectors so that $\eta_{\gamma_{0}}$ approaches the true $\eta_0$ as $m$ tends to infinity. To declutter the notation, we avoid indexing ($\gamma$ and) $\gamma_0$ by $m$, as the size of these parameter vectors should be clear from the context. The densities of $P_0$ and $P_m$ with respect to $\mu$ are denoted $p_0$ and $p_m$, respectively, with similar subscripts for the expectations, $\mathbb{E}_0\, g(X) = \int g(x)\,{\rm d} P_0(x)$ and $\mathbb{E}_m\, g(X) = \int g(x)\,{\rm d} P_m(x)$. The product measure arising from an i.i.d. sample of size $n$ is indicated by a superscript $n$, e.g.: $P_{\theta,\eta}^n = P_{\theta,\eta} \times \cdots \times P_{\theta,\eta}$. The score functions with respect to $\theta$ and $\gamma$ under $\mathcal{P}_m$ are

equation[equation omitted — 246 chars of source]

When evaluated in $(\theta_0,\gamma_0)$, i.e., the `true values' under $\mathcal{P}_m$, we write $\dot{\ell}_m = \dot{\ell}_{\theta_0,T_m^{-1}\gamma_0}$ and $\dot{v}_m = \dot{v}_{\theta_0,T_m^{-1}\gamma_0}$. The log-likelihood function under $\mathcal{P}_m$ is $(\theta,\gamma) \mapsto \sum_{i=1}^n \log p_{\theta,T_m^{-1}\gamma}(X_i)$ for $(\theta,\gamma) \in \Theta \times \Gamma_m$, and the corresponding maximum likelihood estimator (MLE) for $\theta$ is $\widehat{\theta}_{m,n}$. Let $i_m$ be the Fisher information matrix under $P_m$ and block partition it as follows

equation[equation omitted — 286 chars of source]

The standard maximum likelihood theory for inference on $\theta$ in the presence of a finite dimensional (fixed $m$) nuisance parameter $\gamma \in \Gamma_m$ goes as follows: Under regularity conditions (see, e.g., vandervaart1998asymptotic), maximum likelihood estimators are asymptotically linear in the parametric efficient influence function, that is, for fixed $m$,

equation[equation omitted — 170 chars of source]

as $n$ tends to infinity, where $\tilde{\ell}_m$ is the efficient score function and $J_m$ its variance:

equation[equation omitted — 195 chars of source]

with $\Pi_m$ the orthogonal projection onto the linear span of $\{\dot{v}_{m, j}: j=1, \ldots, k_m\}$ in $L_2(P_m)$.

Provided the model $\mathcal{P}_m$ is differentiable in quadratic mean at $(\theta_0,\gamma_0)$ and the efficient information matrix $J_m$ is nonsingular, then (ref) is equivalent to $\widehat\theta_{m, n}$ being the best regular estimator vandervaart1998asymptotic. This means that as $n$ tends to infinity, but $m$ remains fixed, $\sqrt{n}(\widehat\theta_{n, m} - \theta_0)$ converges in distribution under $P_{m}^n$ to a mean zero normal distribution with variance matrix $J_m^{-1}$, this being the smallest possible asymptotic variance matrix of any regular estimator (see Section (ref) for a formal definition).

To illustrate our general strategy as well as the definitions and the notation introduced above, consider the partially linear regression model where $X = (W,Y,Z)$ for a real valued outcome $Y$ that given covariates $(W,Z) = (w,z)$ follows the regression model

equation[equation omitted — 69 chars of source]

Here $\varepsilon$ is a noise term independent of $(W,Z)$, $\eta$ is an infinite dimensional nuisance parameter, and $\theta \in \mathbb{R}$ is the parameter on which we seek to make inference. If $\varepsilon$ is assumed to be mean zero normal with variance $\sigma^2$ and the covariates have a density, then the observation $(W,Y,Z)$ has a density. This density, however, cannot be used to define a maximum likelihood estimator for $(\theta,\eta)$, because the maximiser for $\eta$ will just interpolate the data (see andersen1993statistical for a discussion of these difficulties). To overcome these issues we instead pretend that $Y$ given $(W,Z) = (w,z)$ stems from the parametric model

equation[equation omitted — 88 chars of source]

where $\beta_m = (\beta_{m,1},\ldots,\beta_{m,k_m})^{{\rm t}}$ is a collection of orthonormal (or other basis) functions, $\gamma = (\gamma_1,\ldots,\gamma_{k_m})^{{\rm t}}$ is a Euclidean parameter vector, and $\varepsilon$ and $(W,Z)$ have the same distribution as the similarly denoted random variables above. Now, maximum likelihood estimation is a least squares problem, and we readily obtain a maximum likelihood estimator for $(\theta,\gamma)$, say $(\widehat{\theta}_{m,n},\widehat{\gamma}_{m,n})$. Assuming that the data in fact stem from the parametric model in (ref), we get from standard parametric likelihood theory that $\sqrt{n}(\widehat{\theta}_{m,n} - \theta)$ converges in distribution to a mean zero normal with variance $J_m^{-1}$, where $J_m = \sigma^{-2}\big(\mathbb{E}\, W^2 - \sum_{j=1}^{k_m} (\mathbb{E}\, \{\beta_{m,j}(Z)W\})^2 \big)$, and that this is the efficient information under the model in (ref).

For parametric inference ($m$ fixed) the conclusion above is in many ways the end of the maximum likelihood story: $\sqrt{n}(\widehat{\theta}_{m,n} - \theta_0) \rightsquigarrow {\rm N}(0,J_m^{-1})$, and $J_m^{-1}$ is the smallest possible variance (of a regular estimator). Part of the motivation for the present paper, however, is the observation that if we let $m$ tend to infinity, it is often the case that $J_{m} \to J$ for some nonsingular matrix $J$. It is then tempting to conjecture both that (i) $\sqrt{n}(\widehat\theta_{m_n, n} - \theta_0) \rightsquigarrow {\rm N}(0,J^{-1})$ under $P_{m_n}^n$ for some subsequence $(m_n)_{n \geq 1}$ tending to infinity with $n$; and (ii) that $J$, being the limit of a sequence of efficient information matrices, must be semiparametrically efficient under $\mathcal{P}$. Of course, we desire $\sqrt{n}(\widehat\theta_{m_n, n} - \theta_0) \rightsquigarrow {\rm N}(0,J^{-1})$ under $P_0^n$ rather than $P_{m_n}^n$; we shall subsequently restrict the class of approximating models such that we can make this change of measure.

A case in point is the partially linear model where it follows from Parseval{'}s identity that with $J_m$ the efficient information under (ref)

equation[equation omitted — 93 chars of source]

where the limit is positive (provided $W$ is not a.s. equal to $\mathbb{E}\,(W \,|\, Z)$). Here we also recognise the limit as the efficient information under the semiparametric model in (ref) (see, e.g., BKRW98), demonstrating one case in which the limit of (parametric) efficient information matrices is efficient.

remarkA well known {`}two-steps weak convergence{'} lemma (see, e.g., billingsley1968 or kallenberg2002foundations) says that if $Z_{m,n}\rightsquigarrow Z_m$ for each $m$ and $Z_{m} \rightsquigarrow Z$ and there is a subsequence $(m_n)_{n \geq 1}$ such that $\lim_m\limsup_{n}\Pr(\|Z_{m,n} - Z_{m_n,n}\| \geq \varepsilon) \to 0$ for any $\varepsilon > 0$, then $Z_{m_n,n}\rightsquigarrow Z$. With $Z_{m,n} = \sqrt{n}(\widehat{\theta}_{m,n} - \theta_0)$, as in the setting outlined in the above paragraph, it is tempting to attempt to couple this two-steps theorem with the mutual contiguity $P_0^n \triangleleft \triangleright\, P_{m_n}^n$ in order to conclude that $Z_{m_n,n}$ converges weakly to ${\rm N}(0,J^{-1})$ under $P_0^n$. A closer look at the proof of the two-steps lemma, however, reveals that this conclusion would require $P_0^n$ to be contiguous with respect to $P_m^n$ {\it for any fixed $m$}. But since both $P_0^n = P_0 \times \cdots \times P_0$ and $P_m^n = P_m \times \cdots \times P_m$ with $P_0 \neq P_m$, $P_0^n$ cannot be contiguous with respect to $P_m^n$ (for this impossibility, see oosterhoff1979note or jacod2003limit). Whether the two-steps lemma can be coupled with contiguity appears to be an open problem.

Assumptions for contiguity and efficiency

For our subsequent efficiency results to make sense, we need to impose some structure on the semiparametric model. In Section (ref) we outline this structure, introduce some notation, and define what we mean by asymptotic efficiency. In Section (ref) the conditions we impose on the parametric approximations are presented, along with a few lemmata easing the verification of these.

The true semiparametric model and efficiency

We concentrate on smooth models for i.i.d. data, as in the classical parametric theory. This is made precise by imposing a differentiability in quadratic mean (DQM) condition on the semiparametric model $\mathcal{P}$ lecam1986, BKRW98. Let $B$ be a linear space. We will consider measures $P_{\theta_0 + \tau/\sqrt{n}, \eta_n(b)}$ for $h = (\tau, b)\in \mathbb{R}^p \times B$ where $\eta_n(b)\to \eta_0$ as $n$ tends to infinity, and $\eta_n(0) = \eta_0$.

assumption[DQM] For each $h\in \mathbb{R}^p \times B$, $P_{\theta + \tau/\sqrt{n}, \eta_n(b)}\in \mathcal{P}$ for all large enough $n$, and $\lim_{n\to\infty} \int \{\sqrt{n}(p_{\theta_0 + \tau/\sqrt{n}, \eta_n(b)}^{1/2}-p_0^{1/2}) - \tfrac{1}{2} Ah \, p_0^{1/2}\}^2 \,{\rm d}\mu = 0$, where $A$ is a bounded linear map between $\mathbb{R}^p \times B$ and $L_2(P_0)$.

It follows from the convergence in Assumption (ref) that $Ah\in L_2(P_0)$ and $\int Ah \,{\rm d} P_0 = 0$ vandervaart1998asymptotic. Since $A$ is assumed to be linear we can split out the contributions of the parametric parameter of interest from the infinite dimensional nuisance as $Ah = \tau^{\rm t} \dot{\ell} + Db$, where $\dot{\ell}$ is the ordinary score function for $\theta$ in a model where the nuisance $\eta$ is fixed; while $D \colon B \to \mathbb{R}$ is a linear operator, and $Db$ has the interpretation of a score function for $\eta$ with $\theta$ fixed.

An estimator $\widehat{\theta}_n$ of $\theta_0$ is said to be regular if $\sqrt{n}(\widehat{\theta}_n - \theta_0 - \tau/\sqrt{n}) \rightsquigarrow L$ under $P_{\theta_0+\tau/\sqrt{n}, \eta_n(b)}$, for some law $L$ and each $(\tau, b)\in \mathbb{R}^p\times B$. Requiring regularity excludes superefficient estimators such as the Hodges-Le Cam estimator (see e.g. vandervaart1998asymptotic). The efficiency bound for regular estimators for estimation of $\theta_0$ is determined by the efficient score,

equation[equation omitted — 69 chars of source]

where $\Pi$ denotes the orthogonal projection onto the closure of $\{Db: b\in B\}$ in $L_2(P_0)$. The efficiency bound is the inverse (provided it exists) of the variance of $\tilde{\ell}$,

equation[equation omitted — 81 chars of source]

More precisely, provided $J$ is nonsingular, by the H{\'a}jek--Le Cam convolution theorem, the limiting distribution of any regular sequence of estimators can be represented by the convolution of a ${\rm N}(0, J^{-1})$ with some probability distribution (see e.g. BKRW98; vandervaart1998asymptotic). As such, any regular estimator whose limiting distribution is $L = {\rm N}(0, J^{-1})$ is called best regular; this is what we refer to as asymptotic efficiency.

The discussion immediately above required $J$ to be nonsingular. In fact, this is necessary for the existence of regular estimators chamberlain86asymptotic and therefore we assume this throughout the rest of the paper.

assumption[Nonsingularity] $J$ is nonsingular.

The parametric approximations

We now introduce our assumption on the parametric approximations $P_m = P_{\theta_0,T_m^{-1}\gamma_0}$ to $P_0 = P_{\theta_0,\eta_0}$. In particular, we assume that $P_0$ can be approximated, in an appropriate sense, by a sequence of contiguous alternatives, similar to those in Assumption (ref). Since $P_0$ is unknown, checking Assumptions (ref) and (ref) below entails in practice checking it for any member of $\mathcal{P}$. Formally, however, these assumptions are required to hold only for the specific member $P_0$ of $\mathcal{P}$ that generated the data.

assumption[Contiguity] There is a subsequence $(m_n)$ and a function $g$ such that $\mathbb{E}_0\,g(X) = 0$, $\mathbb{E}_0\,g(X)^2$ is finite, and $\mathbb{E}_0\,g(X) \tilde{\ell}(X) = 0$, and the log-likelihood ratio satisfies \begin{equation} \log \frac{{\rm d} P_{m_n}^n }{{\rm d} P_{0}^n} = \frac{1}{\sqrt{n}}\sum_{i=1}^n g(X_i) - \tfrac{1}{2}\mathbb{E}_0\,g(X_1)^2 + o_{P_0^n}(1). \end{equation}

In our applications, we will seek that $g=0$, but the orthogonality $\mathbb{E}_0\,g(X) \tilde{\ell}(X) = 0$ is all that is actually needed for our results. The log-likelihood ratio expansion in (ref) is equivalent to the densities $p_{m_n} = {\rm d} P_{m_n}/{\rm d} \mu$ satisfying the DQM type condition

equation[equation omitted — 165 chars of source]

for $g$ such that $\mathbb{E}_0\,g(X) \tilde{\ell}(X) = 0$, that is, for the function $g$ appearing in (ref). See for example strasser1985mathematical or lecam1986 for the equivalence of (ref) and (ref). Thus, to check Assumption (ref) it suffices to show either (ref) or (ref).

The log-likelihood ratio expansion in (ref) is key to our results. In particular, since the data are assumed i.i.d., Assumption (ref) and the central limit theorem yield

equation[equation omitted — 240 chars of source]

Since $\exp(Z)$ is positive and $\mathbb{E}\exp(Z) = 1$, we get from Le Cam{'}s first lemma that $P_0^n$ and $P_{m_n}^n$ are mutually contiguous (see vandervaart1988large). Provided $\sqrt{n}(\widehat{\theta}_{m_n,n} - \theta_0)$ converges jointly with $\log ({\rm d} P_0^n /{\rm d} P_{m_n}^n)$ to a Gaussian limit under $P_{m_n}^n$, the asymptotic distribution of $\sqrt{n}(\widehat{\theta}_{m_n,n} - \theta_0)$ under $P_{0}^n$ can be recovered by Le Cam{'}s third lemma lecam1986. As such, Assumption (ref) restricts the possible change in the limiting distribution of $\sqrt{n}(\widehat{\theta}_{m_n,n} - \theta_0)$ resulting from a change of measure from the parametric $P_{m_n}$ back to the semiparametric $P_0$. In particular, provided the limiting distribution of $\sqrt{n}(\widehat{\theta}_{m_n,n} - \theta_0)$ under $P_{m_n}^n$ is orthogonal to $g$ (i.e. $\lim_{n\to\infty} {\rm Cov}(\sqrt{n}(\widehat{\theta}_{m_n,n} - \theta_0),n^{-1/2}\sum_{i=1}^n g(X_i) ) = 0$), changing the measure from $P_{m_n}$ to $P_0$ does not affect the limiting distribution.

In some examples, approximating models satisfying Assumption (ref) can be derived directly from the submodels used to establish Assumption (ref), given the similarity between the DQM required in Assumption (ref) and equation (ref). We view this as a virtue of our contiguous sieve framework which closely links the submodels used to define semiparametric efficiency with those used to estimate $\theta$.

remarkAs discussed above, the weak convergence in (ref), hence Assumption (ref), implies mutual contiguity of $ P_{m_n}^n$ and $P_{0}^n$. Assumption (ref) also imposes additional structure in that it requires the log-likelihood ratios to admit a local asymptotic normality (LAN) type expansion. This allows us ensure that no bias is incurred when switching back to the semiparametric law $P_0$ by imposing only the orthogonality condition $\mathbb{E}_0\,g(X) \tilde{\ell}(X) = 0$. In principle, this LAN-type expansion is not required for our overall strategy: one requires only the contiguity of $P_0^n$ to $P_{m_n}^n$ and the joint convergence of a sequence of statistics and the log-likelihood ratio under $P_{m_n}^n$ to apply the general form of Le Cam's third lemma (e.g., lecam1986 or vandervaart1998asymptotic). With such a weaker requirement (as compared to Assumption (ref)), however, providing conditions under which no bias obtains when switching back from $P_{m_n}$ to $P_0$ becomes more complex.

In addition to the contiguity condition in Assumption (ref), we require that the efficient scores in the parametric submodels approximate the efficient score of the full semiparametric model in a statistically relevant sense. As we will take our limits along the subsequence $(m_n)_{n\geq 1}$ of Assumption (ref), it is sufficient that this approximation holds along this subsequence. This is convenient as Assumption (ref) implies that $P_{m_n} \to P_0$ in total variation (via (ref)), which can help to simplify the demonstration of (ref) (see Lemma (ref) below).

assumption[Efficient score approximation] The efficient scores $\tilde{\ell}_{m_n}$ exist and \begin{equation} \lim_{n\to\infty} \int \|\tilde{\ell}_{m_n} p_{m_n}^{1/2} - \tilde{\ell}p_0^{1/2}\|^2 \,{\rm d}\mu = 0, \end{equation} where $\|x\| = (\sum_{j=1}^p x_j^2)^{1/2}$ is the Euclidean distance.

Explicitly performing the orthogonal projection to compute $\tilde{\ell}$ can, in many models, be quite challenging. Fortunately, one may verify Assumption (ref) without explicitly performing this projection, as the following lemma demonstrates. Here the space $B$ and the linear operator $D$ are as defined in connection with Assumption (ref).

lemmaSuppose that the scores $\dot{\ell}_{m_n}$ and $\dot{v}_{m_n}$ exist in the DQM sense. If \begin{equation} \lim_{n\to\infty} \int \|\dot{\ell}_{m_n}\,p_{m_n}^{1/2} - \dot{\ell}p_0^{1/2}\|^2 \,{\rm d}\mu = 0, \end{equation} and for any $b\in B$ there are vectors $a_{m_n} \in \mathbb{R}^{k_{m_n}}$, such that \begin{equation} \lim_{n\to\infty} \int (a_{m_n}^{\rm t}\dot{v}_{m_n}\,p_{m_n}^{1/2} - Db p_0^{1/2})^2 \,{\rm d}\mu = 0, \end{equation} then Assumption (ref) holds.
proofUnder the assumption of the lemma, for large enough $n$, $\tilde{\ell}_{m_n}$ exists as soon as $\dot{\ell}_{m_n}$ and $\dot{v}_{m_n}$ do. Given (ref) and (ref), apply Theorem (ref) in the appendix to obtain (ref).

We now present two lemmas which can help with the verification of (ref), or of (ref) and (ref). Their straightforward proofs are deferred to Appendix (ref).

lemmaSuppose Assumption (ref) holds. If $f_{m_n} = f_0 + o_{P_0}(1)$ and $f_{m_n}^2$ is uniformly $P_{m_n}$-integrable then $\lim_{n\to\infty} \int \|f_{m_n}p_{m_n}^{1/2} - f_0p_0^{1/2}\|^2\,{\rm d}\mu = 0$.

That $f_{m_n}^2$ is uniformly $P_{m_n}$-integrable means that $\lim_{K\to\infty}\sup_n \int_{|f_{m_n}| \geq K} f_{m_n}^2\,{\rm d} P_{m_n} = 0$. An alternative approach which circumvents the requirement to establish uniform square integrability directly can be based on a lemma originally due to riesz1928convergence (cf. vandervaart1998asymptotic). This lemma also connects Assumption (ref) with the observation that in certain models the sequence of parametric efficient information matrices has the semiparametric efficient information matrix as a limit, as discussed in Section (ref).

lemmaSuppose that {\rm(i)} $p_{m_n}\to p_0$ in $\mu$-measure or {\rm(ii)} for any measurable set $A$, $P_0(A)\le \liminf_{n\to\infty} P_{m_n}(A)$. If $f_{m_n}\to f_0$ in $\mu$-measure and $\limsup_{n\to\infty} \int f_{m_n}^2\,{\rm d} P_{m_n} \le \int f_0^2\,{\rm d} P_0<\infty$, then $\lim_{n\to\infty} \int \|f_{m_n}p_{m_n}^{1/2} - f_0p_0^{1/2}\|^2\,{\rm d}\mu = 0$.

If we use either of these lemmas to verify (ref) directly, we are required to first find the efficient score under the semiparametric model $\mathcal{P}$, a task which involves performing sometimes complicated projections. Lemma (ref), on the other hand, does not involve any projections. The following corollary permits us to find the efficient semiparametric score $\tilde{\ell}$ as a limit.

corollarySuppose that (ref) and (ref) hold. If there is a vector $f_0$ of functions such that the conditions of either Lemma (ref) or (ref) hold for $f_{m_n}$ with components $P_{m_n}$-a.s. equal to each of the components of $\tilde{\ell}_{m_n}$, then $f_0 = \tilde{\ell}$ $P_0$-a.s..
proofSince Lemma (ref) holds, we have that $\int \|\tilde{\ell}_m p_m^{1/2} - \tilde{\ell}p_0^{1/2}\|^2\,{\rm d}\mu \to 0$, which is Assumption (ref). By either Lemma (ref) or (ref) we have $\int \|\tilde{\ell}_m p_m^{1/2} - f_0p_0^{1/2}\|^2\,{\rm d}\mu \to 0$. Since $L_2$ limits are unique up to sets of measure zero, $f_0p_{0}^{1/2} = \tilde{\ell}p_{0}^{1/2}$ $\mu$-almost surely, hence $f_0 = \tilde{\ell}$ $P_0$-almost surely.

Asymptotic efficiency

We now present the main result of the paper. In order to approximate the semiparametric model, we let $m$ increase with $n$. This means that, for our purposes, the property corresponding to the asymptotic linearity in the parametric efficient score function exhibited in (ref), is

equation[equation omitted — 176 chars of source]

Note that property (ref) is not implied by (ref), as (ref) only requires that the remainder $\sqrt{n}(\widehat\theta_{m,n} - \theta_0) -n^{-1/2}\sum_{i=1}^n J_{m}^{-1}\tilde{\ell}_{m}(X_i)$ is $o_{P_{m}^n}(1)$ as $n \to \infty$, for fixed $m$. Verification of (ref) depends on the definition of the estimator $\widehat\theta_{m_n,n}$. In Section (ref) we use a profile likelihood technique to provide sufficient conditions for (ref) to hold for a sequence of maximum likelihood estimators $\widehat{\theta}_{m_n,n}$ in growing contiguous parametric submodels.

Combined with Assumptions (ref)--(ref), the linearity of $\sqrt{n}(\widehat\theta_{m_n,n} - \theta_0)$ in the influence function displayed in (ref) implies the asymptotic efficiency of the estimator $\widehat\theta_{m_n, n}$. The essential idea for proving this is outlined in the following heuristic argument: If Assumption (ref) holds, then $J_{m_n}^{-1}\tilde{\ell}_{m_n}$ in (ref) may be replaced by $J^{-1}\tilde{\ell}$, thus (ref) becomes $\sqrt{n}(\widehat\theta_{m_n,n} - \theta_0) = n^{-1/2}\sum_{i=1}^n J^{-1}\tilde{\ell}(X_i)+o_{P_{m_n}^n}(1)$. Combined with Assumption (ref), this asymptotic linearity ensures that $\sqrt{n}(\widehat\theta_{m_n,n} - \theta_0)$ converges jointly with $ \log ({\rm d} P_0^n /{\rm d} P_{m_n}^n)$ under $P_{m_n}^n$, and hence, by Le Cam's third lemma, one can change the measure from $P_{m_n}$ back to $P_0$ at the cost of adding a bias term of $J^{-1}\mathbb{E}_0\,\tilde{\ell}(X)g(X)$ to the limiting distribution of $\sqrt{n}(\widehat\theta_{m_n,n} - \theta_0)$ under $P_0$. But by the condition on $g$ in Assumption (ref), this bias term is zero.

theoremIf Assumptions (ref) {&} (ref) hold, and $(m_n)$ is a subsequence such that Assumptions (ref) {&} (ref) hold, and (ref) is satisfied, then $\widehat{\theta}_{m_n, n}$ is best regular in $\mathcal{P}$.
proofAssumptions (ref) and (ref) along with the i.i.d. assumption on the data, verify the conditions of Proposition A.10 in vandervaart1988large. Applied to our setting, this proposition gives that $n^{-1/2}\sum_{i=1}^n (\tilde{\ell}_{m_n}(X_i) - \tilde{\ell}(X_i)) = \sqrt{n}(\mathbb{E}_{m_n}\tilde{\ell}_{m_n}(X) - \mathbb{E}_0\,\tilde{\ell}(X)\,) - \mathbb{E}_0\, \tilde{\ell}(X) g(X) + o_{P_0^n}(1) = o_{P_0^n}(1)$. The first equality follows from the cited proposition. The second equality ensues because $\mathbb{E}_{m_n}\tilde{\ell}_{m_n}(X) = 0$, $\mathbb{E}_0\,\tilde{\ell}(X) = 0$, and $\mathbb{E}_0\, \tilde{\ell}(X) g(X) = 0$ by Assumption (ref). Since Assumption (ref) implies that $P_{m_n}^n$ and $P_0^n$ are mutually contiguous, we can swap the $o_{P_{m_n}^n}(1)$ in (ref) with $o_{P_0^n}(1)$, so that (ref) reads $\sqrt{n}(\widehat\theta_{m_n,n} - \theta_0) = n^{-1/2}\sum_{i=1}^n J_{m_n}^{-1}\tilde{\ell}_{m_n}(X_i)+o_{P_{0}^n}(1)$. Moreover, for any $a \in \mathbb{R}^p$ the reverse triangle inequality and then Cauchy--Schwarz yield \begin{align*} |(a^{{\rm t}}J_{m_n}a)^{1/2} - (a^{{\rm t}}Ja)^{1/2}|^2 & = |\|a^{{\rm t}}\tilde{\ell}_{m_n}p_{m_n}^{1/2} \|_{\mu} - \|a^{{\rm t}}\tilde{\ell}p_0^{1/2} \|_{\mu} |^2\\ & \leq \|a^{{\rm t}}(\tilde{\ell}_{m_n}p_{m_n}^{1/2} -\tilde{\ell}p_0^{1/2} ) \|_{\mu}^2 \leq \|a\|^2 \int \|\tilde{\ell}_{m_n} p_{m_n}^{1/2} - \tilde{\ell}p_0^{1/2}\|^2 \,{\rm d}\mu, \end{align*} where the right hand side tends to zero by Assumption (ref). By continuity of the square function, this entails that $a^{{\rm t}}J_{m_n}a \to a^{{\rm t}}Ja$, and since the above is true for any $a \in \mathbb{R}^p$, $J_{m_n} \to J$ and therefore $J_{m_n}^{-1} \to J^{-1}$ since the inverse is a continuous operation when $J$ is nonsingular. Combining this with $n^{-1/2}\sum_{i=1}^n (\tilde{\ell}_{m_n}(X_i) - \tilde{\ell}(X_i)) = o_{P_0^n}(1)$ and (ref), we conclude that \begin{equation} \sqrt{n}(\widehat{\theta}_{m_n,n} - \theta_0) = \frac{1}{\sqrt{n}}\sum_{i=1}^n J^{-1}\tilde{\ell}(X_i) + o_{P_0^n}(1). \end{equation} Given Assumptions (ref) and (ref), the result then follows from Lemma 25.23 and Lemma 25.25 in vandervaart1998asymptotic.

Contiguous sieve MLE

In addition to Assumptions (ref)--(ref), Theorem (ref) requires that the linear expansion (ref) holds. In this section we provide two sets of conditions under which this linear expansion is satisfied for MLEs in growing contiguous parametric models. Both sets of conditions are growing parametric versions of a profile likelihood theorem due to murphy2000profile. These authors provide conditions under which the semiparametric profile likelihood admits a quadratic expansion which, in turn, implies a condition like (ref). In this section we argue similarly, but replace the {\it semiparametric} profile likelihood with a {\it contiguous sieve} profile likelihood which permits us to conclude that (ref) holds. There is a key technical advantage of working with contiguous sieve profile likelihoods in place of the semiparametric profile likelihood. The latter requires the careful construction of {`}approximately least favourable submodels{'}, which can be quite complicated to construct, as can be seen from the examples in the cited article. In contrast, as the likelihoods we work with are parametric, exact least favourable submodels can be constructed following a clear recipe as they always take the same form.

To introduce the profile likelihood, let $(\theta,\gamma) \mapsto L_m(\theta,\gamma)(x) = p_{\theta,T_m^{-1}\gamma}(x)$ be the likelihood function under $\mathcal{P}_m$, and $L_{m,n}(\theta,\gamma) = \prod_{i=1}^n L_m(\theta,\gamma)(X_i)$ be the likelihood based on an i.i.d. sample $X_1,\ldots,X_n$. Denote by ${\rm pl}_{m,n}(\theta)$ the profile likelihood based on $\mathcal{P}_m$

equation[equation omitted — 109 chars of source]

For each $\theta$, let $\widehat{\gamma}(\theta)$ be the value achieving this supremum, that is ${\rm pl}_{m,n}(\theta) = L_{m,n}(\theta,\widehat{\gamma}(\theta))$. We assume throughout that for large enough $n$, such a value exists.

Quadratic expansion & log concavity

A straightforward set of sufficient conditions for (ref) (or (ref)) can be obtained using the results of HjortPollard93. Specifically, let $A_{m,n}(h) = \log {\rm pl}_{m,n}(\theta_0 + h/\sqrt{n}) - \log{\rm pl}_{m,n}(\theta_0)$, then, if the functions $h \mapsto A_{m,n}(h)$ are concave and one manages to find a subsequence $(m_n)$ such that

equation[equation omitted — 170 chars of source]

for each $h$, then the {`}Basic Corollary{'} in HjortPollard93 immediately delivers (ref). This setting covers a large class of semiparametric models of practical interest, including the examples we study in detail in Section (ref) below. We emphasise that the concavity requirement is local, being imposed only on $A_{m, n}$ (actually only the $A_{m_n, n}$ being concave suffices).

For cases in which $A_{m, n}$ is concave and (ref) can be shown to hold, this provides a complete proof of asymptotic normality and efficiency of the contiguous sieve MLE without requiring any empirical process type arguments. We summarise this in a proposition.

propSuppose that Assumptions (ref) & (ref) hold and that $(m_n)$ is a subsequence such that Assumptions (ref) {&} (ref) and the quadratic expansion in (ref) hold. If $h \mapsto A_{m_n,n}(h)$ is concave, then (ref) holds, and $\widehat{\theta}_{m_n, n}$ is best regular in $\mathcal{P}$.
proofThis follows from the Basic Corollary in HjortPollard93. In particular, since $J_{m_n}\to J$ under Assumption (ref) we have $A_{m_n,n}(h) = h^{{\rm t}}n^{-1/2}\sum_{i=1}^n \tilde{\ell}_{m_n}(X_i) - \tfrac{1}{2} h^{{\rm t}}J h + o_{P_{m_n}^n}(1)$. As $A_{m,n}(h)$ is concave, $\sqrt{n}(\widehat{\theta}_{m_n,n} - \theta_0) = n^{-1/2} \sum_{i=1}^n J^{-1}\tilde{\ell}_{m_n}(X_i) + o_{P_{m_n}^n}(1)$ by the Basic Corollary. $\widehat{\theta}_{m_n,n}$ is then best regular in $\mathcal{P}$ by Theorem (ref).

Compared to Theorem (ref), the above proposition allows us to replace (ref) with (ref), provided $h \mapsto A_{m,n}(h)$ is concave. One key advantage of this is that (ref) and concavity give us {\it both} consistency and asymptotic normality of $\widehat{\theta}_{m_n,n}$, while to establish (ref) without concavity, one typically first needs to establish consistency.

We now present a theorem giving conditions under which (ref) holds. This theorem is a stripped down and sieved version of Theorem 1 in murphy2000profile, and requires that we introduce some quantities inspired by that paper. It is in the construction of these quantities we gain a lot in simplicity by working with growing parametric models, compared to attacking the semiparametric model directly. This is because under $P_m = P_{\theta_0,T_m^{-1}\gamma_0}$, the least favourable submodel for estimating $\theta_0$ always takes the form $\theta \mapsto P_{\theta,T_{m}^{-1}\gamma(\theta)}$, where $\gamma(\theta) = \gamma_0 + i_{m,11}^{-1}i_{m,10}(\theta_0 - \theta)$. In view of this, define $\gamma_t^{{\rm sub}}(\theta,\gamma) = \gamma + i_{m,11}^{-1}i_{m,10}(\theta - t)$, where we suppress the dependence of $\gamma_t^{{\rm sub}}(\theta,\gamma)$ upon $m$ from the notation. For each $m$ and each $(\theta,\gamma) \in \Theta \times \Gamma_m$ define the mappings $t \mapsto l_m(t,\theta,\gamma) \coloneqq \log L_m(t,\eta_{\gamma_t^{{\rm sub}}(\theta,\gamma)})$. These functions bridge the gap between the log-profile likelihood and the efficient score. To see this, notice that $\eta_{\gamma_{\theta}^{{\rm sub}}}(\theta,\gamma) = \gamma$, and that the derivative of $l_m(t,\theta,\gamma)$ with respect to $t$ is

equation[equation omitted — 255 chars of source]

in particular, $\dot{l}_m(\theta_0,\theta_0,\gamma_0) = \tilde{\ell}_{m}$. At the same time, mimicking the sandwiching technique of murphy2000profile, we have, writing $\mathbb{P}_n f = n^{-1}\sum_{i=1}^n f(X_i)$ for the integral with respect to the empirical measure, that for arbitrary $\tilde{\theta}_n$

equation[equation omitted — 282 chars of source]

and

equation[equation omitted — 314 chars of source]

In particular, the process $h \mapsto A_{m,n}(h) = \log {\rm pl}_{m,n}(\theta_0 + h/\sqrt{n}) - \log {\rm pl}_{m,n}(\theta_0)$ can be squeezed between two quantities approximating the {`}efficient LAN expansion{'} of (ref).

theoremSet $\tilde{\theta}_{n} = \theta_0 + h/\sqrt{n}$ for some fixed $h$. Assume that $t \mapsto l_m(t,\theta,\gamma)(x)$ is twice continuously differentiable for each $m$, $(\theta,\gamma)$ and $x$, with derivatives $\dot{l}_m$ and $\ddot{l}_m$, and that there is a subsequence $(m_n)$ such that for $\tilde{\psi}$ equal to either $(\tilde\theta_n, \widehat{\gamma}(\tilde{\theta}_n))$ or $(\theta_0, \widehat{\gamma}(\theta_0))$, \begin{equation} \sqrt{n}\mathbb{P}_n \dot{l}_{m_n}(\theta_0,\tilde{\psi}) = \sqrt{n}\mathbb{P}_n \tilde{\ell}_{m_n} + o_{P_{m_n}^n}(1), \quad and\quad \mathbb{P}_n \ddot{l}_{m_n}(s_n, \tilde{\psi}) = - J_{m_n} + o_{P_{m_n}^n}(1), \notag \end{equation} for any deterministic sequence $s_n \to \theta_0$. Then (ref) holds.
proofA Taylor expansion keeping $\tilde{\psi}$ fixed: $n\mathbb{P}_n l_{m_n}(\theta_0 + h/\sqrt{n}, \tilde{\psi} ) - n\mathbb{P}_n l_{m_n}(\theta_0, \tilde{\psi} ) = h^{\rm t} \sqrt{n}\,\mathbb{P}_n \dot{l}_{m_n}(\theta_0,\tilde{\psi}) + \tfrac{1}{2} h^{\rm t}\, \mathbb{P}_n \ddot{l}_{m_n}( s_n, \tilde{\psi}) h$ for an $s_n$ between $\tilde{\theta}_n$ and $\theta_0$. Replace $\mathbb{P}_n \ddot{l}_{m_n}(s_n, \tilde{\psi})$ by $-J_{m_n} + o_{P_{m_n}^n}(1)$, and $\sqrt{n}\,\mathbb{P}_n \dot{l}_{m_n}(\theta_0,\tilde{\psi})$ by $\sqrt{n}\mathbb{P}_n \tilde{\ell}_{m_n} + o_{P_{m_n}^n}(1)$ in both of the sandwiching bounds in (ref) and (ref).

Random quadratic expansions

The concavity assumption that we worked under in the previous section can be substituted with a consistency assumption. Suppose that for any random sequence $\tilde{\theta}_n$ such that $\tilde{\theta}_n = \theta_0 + o_{P_{m_n}^n}(1)$ for some subsequence $(m_n)$, with $\tilde{h}_n = \tilde{\theta}_n - \theta_0$, we have

equation[equation omitted — 256 chars of source]

where $r_{n}(\tilde{\theta}_n) = o_{P_{m_n}^n}(\sqrt{n}\|\tilde{h}_n\| + n\|\tilde{h}_n\|^2 + 1)$. This gives a proposition that is similar to Proposition (ref), replacing concavity with consistency.

propSuppose Assumptions (ref)--(ref) hold, and let $\widehat{\theta}_{m,n}$ be the maximiser of ${\rm pl}_{m,n}(\theta)$. If $\widehat{\theta}_{m_n,n} = \theta_0 + o_{P_{m_n}^n}(1)$ and (ref) holds, then (ref) holds, and consequently $\widehat{\theta}_{m_n,n}$ is best regular in $\mathcal{P}$.

The following theorem, which is analogous to Theorem (ref), provides conditions under which the expansion given in (ref) holds.

theoremAssume that $t \mapsto l_m(t,\theta,\gamma)(x)$ is twice continuously differentiable for each $m$, $(\theta,\gamma)$ and $x$, with derivatives $\dot{l}_m$ and $\ddot{l}_m$; and that there is a subsequence $(m_n)$ such that for any random sequence $\tilde{\theta}_n$ with $\tilde{\theta}_n = \theta_0 + o_{P_{m_n}^n}(1)$, we have \begin{equation} \sqrt{n}\mathbb{P}_n \dot{l}_{m_n}(\theta_0,\tilde{\theta}_n,\widehat{\gamma}(\tilde{\theta}_n)) = \sqrt{n}\mathbb{P}_n \tilde{\ell}_{m_n} + o_{P_{m_n}^n}(\sqrt{n}\|\tilde\theta_n - \theta_0\| + 1), \end{equation} and \begin{equation} \mathbb{P}_n \ddot{l}_{m_n}(s_n,\tilde{\theta}_n,\widehat{\gamma}(\tilde{\theta}_n)) = - J_{m_n} + o_{P_{m_n}^n}(1), \end{equation} for any random sequence $s_n = \theta_0 + o_{P_{m_n}^n}(1)$. Then (ref) holds.

Proposition (ref) and Theorem (ref) are growing parametric versions of Corollary 1 and Theorem 1 (respectively) in murphy2000profile. As they are proven by making what are essentially notational adjustments to the proofs of the cited results, we defer their proofs to Appendix (ref). In Appendix (ref) we also provide sufficient conditions for the assumptions made in Theorem (ref).

Applications

In this section we continue the example of the partially linear model, and also apply our theory to the Cox regression model. Both models satisfy the background assumptions on the semiparametric model (Assumption (ref) and (ref)), consequently we concentrate on the assumptions made on the parametric approximations, that is Assumptions (ref) and (ref), along with (ref). For both models we take, for simplicity, $\theta \in \Theta \subset \mathbb{R}$, and $\eta$ a real valued $s$-times continuously differentiable function on the unit interval. The parametric approximations employed take the form $\beta_m(z)^{{\rm t}}\gamma = (T_m^{-1}\gamma)(z)$, where for each $m$, $\beta_m = (\beta_{m,1},\ldots,\beta_{m,k_m})^{{\rm t}}$ is a collection of orthonormal functions, and $\gamma = (\gamma_1,\ldots,\gamma_{k_m})^{{\rm t}} \in \Gamma_m \subset \mathbb{R}^{k_m}$ are coefficients such that $\beta_m^{{\rm t}}\gamma \to \eta$ in $L_2([0,1],\nu)$, where $\nu$ is an appropriate finite measure (which is always possible, see e.g. Theorem 14.3.1 in szego1975orthogonal). To not overburden the notation, the vectors $\gamma$ are not indexed by $m$. Some of the details of the two following examples are left to appendices (ref) and (ref).

We note that the partially linear model considered immediately below is a special case of the class of {`}partially linear GLMs{'} studied by mammen1997penalized, amongst others. It should be uncomplicated to extend our results in Section (ref) below to a large subclass of such partially linear GLMs. In particular, it is straightforward to establish the concavity of $A_{m, n}$ for any such model with a canonical link function.

The partially linear model

We have $n$ i.i.d. replicates of $X = (W,Y,Z)$, for an outcome $Y$ and covariates $(W,Z)$ with values in $\mathbb{R} \times [0,1]$. The observation $X$ stems from the model $P_0 = P_{\theta_0,\eta_0}$, where each member $P_{\theta,\eta} \in \mathcal{P}$ has density

equation[equation omitted — 156 chars of source]

with respect to Lebesgue measure, and $f_{W,Z}(w,z)$ is the joint density of $(W,Z)$. We assume that $\mathbb{E}\,|W|^{2+\delta}<\infty$ for some $\delta>0$, and denote $P_Z$ the (marginal) law of $Z$. The parametric densities in $\mathcal{P}_m$ are of the same form as (ref), with $\eta$ replaced by $\beta_m^{{\rm t}}\gamma$. The score functions under $\mathcal{P}_m$ are then

equation[equation omitted — 266 chars of source]

Due to the orthonormality of $\beta_m = (\beta_{m,1},\ldots,\beta_{m,k_m})^{{\rm t}}$, the Fisher information matrix under $\mathcal{P}_m$ takes an appealing form, in particular $i_{m,11} = \sigma^{-2}I_{k_m}$ where $I_{k_m}$ is the $k_m$-dimensional identity matrix, and $i_{m,01} = \sigma^{-2}\, \mathbb{E}\, W \beta_m(Z)^{{\rm t}}$. The efficient score under $\mathcal{P}_m$ is therefore

equation[equation omitted — 188 chars of source]

Define $b_0(z) = \mathbb{E}\,(W \,|\, Z = z)$ and $b_m(z) = \beta_m(z)^{{\rm t}}\mathbb{E}\,[W \beta_m(Z)] = \sum_{j=1}^{k_m}\beta_{m,j}(z) \langle \beta_{m,j},b_0\rangle$. As the $\beta_m$ form an orthonormal basis for $L_2([0, 1], P_Z)$, we have $b_m\to b_0$ in $L_2([0, 1], P_Z)$. We first turn to Assumption (ref). The log-likelihood ratio of $P_{m}^n$ with respect to $P_0^n$ can be written

equation[equation omitted — 178 chars of source]

in terms of $h_{m,n}(z) = (\sqrt{n}/\sigma)(\beta_m(z)^{{\rm t}}\gamma_0 - \eta_0(z))$ and $\varepsilon_i = Y_i - \eta_0(Z_i) - \theta_0 W_i$. Assume that there is a subsequence $(m_n)_{n \geq 1}$ and a function $h$ such that $h_{m_n,n}\to h$ in $L_2(P_0)$. (Conditions under which this holds are given as (ref) & (ref) in Appendix section (ref).) Under this assumption it follows from the fact that the data are i.i.d. with $\varepsilon$ independent of $Z$ and the law of large numbers that, with $g(x) = h(z)(y - \eta_0(z) - \theta_0 w)/\sigma$, we have

equation[equation omitted — 173 chars of source]

as $\mathbb{E}_0\, h(Z)^2 = \mathbb{E}_0\, h(Z)^2 (\varepsilon / \sigma)^2$. To conclude that Assumption (ref) holds, we need also show that $\mathbb{E}_0 \, g(X) \tilde{\ell}(X) = 0$, which requires that we determine the form of the efficient score. Working formally, we expect that $\tilde{\ell}_{\theta_0,T_m^{-1}\gamma_0}$ will converge to

equation*[equation* omitted — 102 chars of source]

We verify that this is true in $L_1(P_0)$ in Appendix (ref). As $ \tilde{\ell}_{\theta_0,T_{m}^{-1}\gamma_0}(X) \sim \sigma^{-2}(W - b_m(Z))\varepsilon$ under $P_{m}$, $\mathbb{E}\,(|W|^{2+\delta})\,\mathbb{E}\,(|\varepsilon|^{2+\delta})<\infty$ by assumption, and

equation[equation omitted — 144 chars of source]

it follows that $\tilde{\ell}_{\theta_0,T_{m_n}^{-1}\gamma_0}$ is square uniformly $P_{m_n}$-integrable. Since also $P_{m_n} \to P_0$ in total variation by (ref) (via (ref)), Lemma (ref) in the appendix yields

equation*[equation* omitted — 137 chars of source]

Since (ref) and (ref) also hold (as shown in Appendix (ref)), it follows from Corollary (ref) that $u = \tilde{\ell}$ $P_0$-almost surely. We now have an expression for the efficient score, and can verify that Assumption (ref) holds as

equation[equation omitted — 191 chars of source]

In Appendix (ref) we show that equations (ref) and (ref) hold, entailing that Assumption (ref) also holds by Lemma (ref). From the developments so far, we see that the semiparametric efficient information is $J = \sigma^{-2}(\mathbb{E}\,W^2 - \|b_0\|^2) = \sigma^{-2}\mathbb{E}\,\big(W^2 - (\mathbb{E}\,[W \,|\, Z])^2\big)$, and also note that

equation[equation omitted — 275 chars of source]

which may also be verified directly using Parseval{'}s identity.

Let $(\widehat{\theta}_{m,n},\widehat{\gamma}_{m,n})$ be the maximum likelihood estimator under $\mathcal{P}_m$. To establish that $\widehat{\theta}_{m_n,n}$ is best regular, we show that $h \mapsto \log {\rm pl}_{m,n}(\theta_0 + h/\sqrt{n}) - \log {\rm pl}_{m,n}(\theta_0)$ is concave and admits a quadratic expansion as in (ref), which by Proposition (ref) will allow us to conclude that $\widehat{\theta}_{m_n,n}$ is best regular. The value achieving the supremum in $\sup_{\gamma \in \Gamma_m} L_{m,n}(\theta,\gamma)$ is the least squares solution $\widehat{\gamma}(\theta) = B_{m,n}^{-1}n^{-1}\sum_{i=1}^n \beta_m(Z_i)(Y_i - \theta W_i)$, where $B_{m,n} = n^{-1}\sum_{i=1}^n \beta_m(Z_i)\beta_m(Z_i)^{{\rm t}}$. The log-profile likelihood can then be expressed as $\log {\rm pl}_{m,n}(\theta) = - 1/(2\sigma^2)\sum_{i=1}^n\big(Y_i - \breve{Y}_{m,n,i} - \theta(W_i - \breve{W}_{m,n,i})\big)^2$, with $\breve{W}_{m,n,i} = \beta_m(Z_i)^{{\rm t}}B_{m,n}^{-1}n^{-1}\sum_{j=1}^n \beta_m(Z_j) W_j$ and $\breve{Y}_{m,n,i} = \beta_m(Z_i)^{{\rm t}}B_{m,n}^{-1}n^{-1}\sum_{j=1}^n \beta_m(Z_j) Y_j$. We see that $\log {\rm pl}_{m,n}(\theta)$ is concave in $\theta$. The expansion (ref) is verified Appendix (ref) under conditions (ref) and (ref). In summary, under the basic conditions given here along with (ref)--(ref), the conditions of Proposition (ref) are satisfied and consequently the estimator $\widehat{\theta}_{m_n,n}$ is best regular in $\mathcal{P}$.

Efficiency of the Cox partial likelihood estimator

The Cox regression model differs from the partially linear model in an important way. Whilst in the partially linear model a maximum likelihood estimator for $\theta$ cannot be defined without recourse to sieves, penalisation, or the like (as discussed in the Section (ref)), no such techniques are called for in the Cox regression model, as the Cox partial likelihood permits us to work directly with the semiparametric model. This entails that an estimator for the parametric part of the Cox regression model can be derived straightforwardly, without (for example) the theory developed in this paper. This theory does, however, lead to a simple proof of the efficiency of the Cox partial likelihood estimator, as we will now show. In other words, in this application the theory developed in this paper is used solely as a theoretical tool. Consequently, the choice of basis functions is of no practical importance, and one may (and, indeed, one ought to) use basis functions that make the analysis particularly tractable, which is what we do here.

Suppose we have $n$ i.i.d. replicates of $X = (T,\Delta,W)$ observed over $[0,1]$, with $T = \min(T^{\prime},C)$ and $\Delta = I(T^{\prime} \leq C)$, where the lifetime $T^{\prime}$ and the censoring time $C$ are independent given $W$, and $T^{\prime}$ given $W = w$ follows a Cox model, i.e., its hazard rate is of the form $\eta(t) \exp(\theta w)$ where $\theta \in \Theta \subset \mathbb{R}$ for simplicity, and $\eta$ belongs to the space $\mathcal{H}$ of one time continuously differentiable functions $\eta \colon [0,1] \to (0,\infty)$. As above, $P_{0} = P_{\theta_0,\eta_0}$ is the true data generating mechanism, while $P_m = P_{\theta_0,T_m^{-1}\gamma_0}$ are the parametric approximations. Here we take the basis function

equation[equation omitted — 116 chars of source]

where $V_{m,1},\ldots,V_{m,k_m}$ are disjoint intervals whose union make up $[0,1]$, and each interval is of length $1/k_m$. With this basis, a natural choice is to work under the coefficient $\gamma_0 = (\gamma_{0,1},\ldots,\gamma_{0,k_m})^{{\rm t}}$ with $\gamma_{0,j} = \eta_0((j-1)/k_m)/k_m$ for $j = 1,\ldots,k_m$. Introduce the counting process $N(t) = \Delta I(T \leq t)$, the at-risk process $Y(t) = I(T \geq t)$, and let $(\mathcal{F}_{t})_{t \in [0,1]}$ be the filtration generated by these. Then, with respect to this filtration, $M(t) = N(t) - \int_0^t Y(s)\eta_0(s)\exp(\theta_0 W)\,{\rm d} s$ and $M^m(t) = N(t) - \int_0^t Y(s)\beta_m(s)^{{\rm t}} \gamma_0\exp(\theta_0 W)\,{\rm d} s$ are square integrable martingales with respect to $P_0$ and $P_m$, respectively.

Using the Taylor-expansion $\log(1 + a) = a - \tfrac{1}{2} a^2 + a^2R(a)$ where $R(a) \to 0$ as $a \to 0$, the log-likelihood ratio based on the full sample can be written

equation[equation omitted — 242 chars of source]

where $h_{m,n} = \sqrt{n}(\beta_m^{{\rm t}}\gamma_0 - \eta_0)$, and $r_{m,n} = n^{-1}\sum_{i=1}^n\int_0^1 h_{m,n}^2/\eta_0^2R(n^{-1/2}h_{m,n}/\eta_0)\,{\rm d} N_i$ Let $(m_n)$ be so that $\sqrt{n}/k_{m_n} \to 0$. Then we can apply Lenglart{'}s inequality (e.g., andersen1993statistical or jacod2003limit) to show that $\log {\rm d} P_{m_n}^n/{\rm d} P_{0}^n$ tends to zero in probability, and Assumption (ref) is satisfied with $g =0$. In fact, with the above choice of basis functions, this is the only possible limit.

Next, we verify Assumption (ref). Under $\mathcal{P}_m$, the score functions are

equation[equation omitted — 163 chars of source]

when evaluated in $(\theta_0,T_m^{-1}\gamma_0)$. With $s_{m}^{(k)}(t) = \mathbb{E}_m\, Y(t)W^k \exp(W \theta_0)$ for $k = 0,1,2$, this gives

equation[equation omitted — 214 chars of source]

Inserting the locally constant basis functions in (ref), we find the efficient score under $\mathcal{P}_m$,

equation[equation omitted — 197 chars of source]

At this point it is tempting to conjecture that with $s^{(k)}$ the pointwise limits of $s_m^{(k)}$, the efficient score under $\mathcal{P}$ is $\tilde{\ell}(X) = \int_0^1 (W - s^{(1)}(t)/s^{(0)}(t) ) \,{\rm d} M(t)$. This is indeed the case, as can be established using Lemma (ref), the conditions of which we verify in Appendix (ref) by a simple, if tedious, application of Lemma (ref).

The main message is that these lemmata, combined with Corollary (ref), allow us to conclude that the efficient score under $\mathcal{P}$ is as conjectured above, that is, $\tilde{\ell}(X) = \int_0^1 (W - s^{(1)}/s^{(0)} ) \,{\rm d} M$; and consequently, the efficient information is

equation[equation omitted — 182 chars of source]

This method of finding the efficient score and the efficient information, leads to a new, relatively simple, and rather parametric proof of the efficiency of the Cox partial likelihood estimator (see for example BKRW98 or kosorok2008introduction for proofs that differ from ours). Let $L_{n}^{{\rm cox}}(\theta)$ be Cox{'}s partial likelihood function (see gill1984understanding for an excellent introduction). We may then define the process

equation[equation omitted — 274 chars of source]

and note that $h \mapsto A_{n}^{\rm cox}(h)$ is concave. Under standard regularity conditions in the Cox regression setting -- given as (ref)--(ref) in Appendix (ref) -- $A_{n}^{\rm cox}(h)$ admits the expansion $A_{n}^{\rm cox}(h) = hn^{-1/2}\sum_{i=1}^n \tilde{\ell}(X_i) - \tfrac{1}{2} h^2 J + o_{P_0}(1)$, where $\tilde{\ell}$ and $J$ are as defined above (this follows from a Taylor expansion combined with the second part of Theorem 3.2 in andersen1982cox). Since $h \mapsto A_{n}^{\rm cox}(h)$ is concave, the Basic Corollary in HjortPollard93 entails that the maximiser of $L_{n}^{{\rm cox}}(\theta)$, say $\widehat{\theta}_n^{\rm cox}$, satisfies (ref), that is $\sqrt{n}(\widehat{\theta}_n^{\rm cox} - \theta_0) = n^{-1/2} \sum_{i=1}^n J^{-1}\tilde{\ell}(X_i) + o_{P_0^n}(1)$. That the Cox partial likelihood estimator $\widehat{\theta}_n^{\rm cox}$ is best regular under $\mathcal{P}$ now follows from Lemma 25.23 and Lemma 25.25 in vandervaart1998asymptotic.

Conclusion

This paper develops an alternative approach to establishing the asymptotic normality and efficiency of certain approximate maximum likelihood estimators in semiparametric models. These estimators are contiguous sieve estimators, being maximum likelihood estimators in contiguously growing parametric models. The approach detailed in this paper, however, departs substantially from the conventional sieve literature as we work both with and under the approximating contiguous parametric models, switching back to the semiparametric model in the limit via Le Cam's third lemma. Working with (growing) parametric models allows for a straightforward definition of a maximum likelihood estimator; working under these same parametric models ensures that we may work as if our approximating models were correctly specified at each step. In the examples we have considered in detail, this approach leads to substantial simplifications and relatively straightforward proofs of asymptotic efficiency.

appendix\section{Technical results} \begin{theorem} Let $H$ be a Hilbert space, $h_n, h\in H$, and $L_n, L$ closed linear subspaces of $H$. Let $g_n\coloneqq \Pi(h_n | L_n)$ and $g\coloneqq \Pi(h|L)$. If {\rm(i)} $h_n \to h$ and {\rm(ii)} for each $f\in L$, there is a sequence $(f_n)_{n\in \mathbb{N}}$ and a $N\in \mathbb{N}$ such that $f_n \to f$ and $f_n\in L_n$ for $n\ge N$, then $g_n\to g$. \end{theorem} \begin{proof} Let $\Pi_n$ be the orthogonal projection onto $L_n$ and $\Pi$ that onto $L$. First suppose $h_n=h$ ($n\in \mathbb{N}$). As $(g_n)_{n\in \mathbb{N}}$ is bounded, any subsequence contains a weakly convergent subsequence, say $g_{n_k} \rightharpoonup g^\star$. By self-adjointness and idempotency \begin{equation} \left\langle {g_{n_k}}\, ,\, {g_{n_k}} \right\rangle = \left\langle {\Pi_{n_k} h}\, ,\, {\Pi_{n_k} h} \right\rangle = \left\langle {h}\, ,\, {\Pi_{n_k} h} \right\rangle \to \left\langle {h}\, ,\, {g^\star} \right\rangle. \end{equation} Let $f\in L$. By hypothesis there are $(f_n)_{n\in \mathbb{N}}$ with $f_n\to f$ and $f_n\in L_n$ for $n\ge N_1$. So $f_{n_k}\to f$ and $f_{n_k}\in L_{n_k}$ for $k\ge K_1$. Since $h - \Pi_{n_k}h \rightharpoonup h - g^\star$, by Proposition 16.7 in royden2010real and the fact that $h - g_{n_k}\in L_{n_k}^\perp$ for each $k$, $\left\langle {h - g^\star}\, ,\, {f} \right\rangle = \lim_{k\to\infty} \left\langle {h - g_{n_k}}\, ,\, {f_{n_k}} \right\rangle = 0$. Hence $g^\star = \Pi h = g$. By self-adjointness and idempotency of $\Pi$ and (ref), $ \lim_{k\to\infty}\left\langle {g_{n_k}}\, ,\, {g_{n_k}} \right\rangle = \left\langle {h}\, ,\, {\Pi h} \right\rangle = \left\langle {\Pi h}\, ,\, {\Pi h} \right\rangle = \left\langle {g}\, ,\, {g} \right\rangle$, and hence $g_{n_k}\to g$ by the Radon--Riesz Theorem. As the initial subsequence was arbitrary, $g_n\to g$. To complete the proof, for $h_n\to h$ an arbitrary convergent sequence, $\|g_n - g\| \le \|h_n - h\| + \|\Pi_n h - \Pi h\|$. The first right hand side term is $o(1)$ by assumption; the second by the case with $h_n = h$. \end{proof} \begin{lemma} Let $P_n, P_0\ll \mu$ with densities $p_n, p_0$. Suppose that {\rm(i)} $P_n \to P_0$ in total variation; {\rm(ii)} $f_{n}$ converges to $f_0$ in $P_0$-probability; and {\rm(iii)} $f_n$ is uniformly square $P_n$-integrable. Then $\lim_{n\to\infty} \int \big(f_n p_{n}^{1/2} - f_0p_0^{1/2}\big)^2\,\mathrm{d}\mu =0$. \end{lemma} \begin{proof} $\int f_0^2\,\mathrm{d}P_0 < \infty$ by a version of Fatou's Lemma (e.g. Lemma 2.2 in serfozo1982). Expansion of the square yields \begin{equation*} \int \big(f_n p_{n}^{1/2} - f_0p_0^{1/2}\big)^2\,\mathrm{d}\mu = \int f_n^2 \,\mathrm{d}P_n + \int f_0^2 \,\mathrm{d}P_0 -2 \int f_nf_0 p_n^{1/2}p_0^{1/2} \,\mathrm{d}\mu. \end{equation*} Combining (i), (ii), (iii) and a version of Vitali's convergence theorem (Corollary 2.9 in feinberg2016uniform) gives $\lim_{n\to\infty} \int f_n^2 \,\mathrm{d}P_n = \int f_0^2\,\mathrm{d}P_0$. Hence the proof will be complete if we show that $\lim_{n\to\infty} \int f_nf_0 p_n^{1/2}p_0^{1/2}\,\mathrm{d}\mu = \int f_0^2\,\mathrm{d}P_0$. Let $Q_n$ be the probability measure with $\mu$-density $q_n\coloneqq c_n p_{n}^{1/2}p_0^{1/2}$ where $c_n$ is the normalising constant. We have $ c_n^{-1} \coloneqq \int p_{n}^{1/2}p_0^{1/2}\,\mathrm{d}\mu = 1 -\frac{1}{2}\int (p_{n}^{1/2} - p_0^{1/2})^2\,\mathrm{d}\mu \to 1$ as $\int (p_{n}^{1/2} - p_0^{1/2})^2\,\mathrm{d}\mu\le 2d_{{\rm TV}}(P_n, P_0)$, with $d_{{\rm TV}}$ the total variation distance strasser1985mathematical. Similarly, \begin{align*} \int |q_n - p_0|\,\mathrm{d}\mu &\le \int p_0^{1/2} |c_n| |p_n^{1/2} - p_0^{1/2}|\,\mathrm{d}\mu + \int p_0 |c_n-1|\,\mathrm{d}\mu\\ &\le |c_n| \bigg(\int p_0 \,\mathrm{d}\mu\bigg)^{1/2} \bigg(\int (p_n^{1/2} - p_0^{1/2})^2\,\mathrm{d}\mu\bigg)^{1/2} + |c_n-1| \to 0, \end{align*} implying that $d_{{\rm TV}}(Q_n, P_0)\to 0$. Now, let $g_n \coloneqq f_nf_0$. Note that $h_n \coloneqq g_n + |g_n| \ge 0$. Since $|g_n| \to f_0^2$ in $P_0$-probability, we have $\liminf_{n\to\infty} \int |g_n| \,\mathrm{d}Q_n \ge \int f_0^2 \,\mathrm{d}P_0$ by a version of Fatou's lemma (Corollary 2.3 in feinberg2016uniform). Additionally, by the Cauchy--Schwarz inequality \begin{equation*} \limsup_{n\to\infty}\int |g_n|\,\mathrm{d}Q_n \le \limsup_{n\to\infty} c_n \bigg(\int f_n^2 \,\mathrm{d}P_n \bigg)^{1/2}\bigg(\int f_0^2 \,\mathrm{d}P_0 \bigg)^{1/2} = \int f_0^2\,\mathrm{d}P_0. \end{equation*} In consequence, $\lim_{n\to\infty} \int |g_n|\,\mathrm{d}Q_n = \int f_0^2\,\mathrm{d}P_0$. An entirely analogous argument applied to $h_n\ge 0$ shows that $\lim_{n\to\infty} \int h_n\,\mathrm{d}Q_n = 2\int f_0^2\,\mathrm{d}P_0$. Combining these two limit results yields \begin{equation*} \lim_{n\to\infty} \int f_nf_0 \,\mathrm{d}Q_n = \lim_{n\to\infty}\int h_n - |g_n| \,\mathrm{d}Q_n = \int f_0^2\,\mathrm{d}P_0. \end{equation*} Since $\int f_nf_n p_n^{1/2}p_0^{1/2}\,\mathrm{d}\mu = c_n^{-1} \int f_nf_0\,\mathrm{d}Q_n$ and $c_n\to 1$, this completes the proof. \end{proof} We next provide proofs of Lemmas (ref) and (ref). \begin{proof}[Proof of Lemma (ref)] Since Assumption (ref) holds, so does (ref) and hence $P_{m_n} \to P_0$ in total variation. Apply Lemma (ref). \end{proof} \begin{proof}[Proof of Lemma (ref)] In case (i) we have $f_{m_n}p_{m_n}^{1/2} \to f_0p_0^{1/2}$ in $\mu$-measure. The conclusion follows from Proposition 2.29 in vandervaart1998asymptotic as $\limsup_{n}\int (f_{m_n}p_{m_n}^{1/2})^2\,\mathrm{d}\mu \le \int (f_0p_0^{1/2})^2\,\mathrm{d}\mu <\infty$. In case (ii) the result follows directly from Proposition S3.1 in the supplementary material to hoesch24locally. \end{proof} \section{Additional details for Section (ref)} \subsection{Proofs of Proposition (ref) and Theorem (ref)} \begin{proof}[Proof of Proposition (ref)] Let $\Delta_{n}\coloneqq n^{-1/2}\sum_{i=1}^n \tilde{\ell}_{m_n}(X_i)$ and $\widehat{h}_{n}\coloneqq \sqrt{n}(\hat{\theta}_{n, m_n} - \theta)$. Applying (ref) with $\tilde\theta_n = \widehat{\theta}_{m_n,n}$ gives \begin{equation*} \log {\rm pl}_{n, m_n}(\widehat{\theta}_{m_n,n}) = \log {\rm pl}_{n, m_n}(\theta_0) + \widehat{h}_n^{\rm t} \Delta_n - \frac{1}{2}\widehat{h}_n^{\rm t} J_{m_n}\widehat{h}_n + r_{n}(\hat{\theta}_{n, m_n}), \end{equation*} where $r_{n}(\widehat{\theta}_{m_n,n}) = o_{P_{m_n}^n}(\|\widehat{h}_n\| + 1)^2$. Since $J_{m_n} = \mathbb{E}_{m_n}\tilde{\ell}_{m_n}(X)\tilde{\ell}_{m_n}(X)^{\rm t}$, Assumptions (ref) & (ref) ensure that $J_{m_n}^{-1}\Delta_n = O_{P_{m_n}^n}(1)$. By (ref) with $\tilde\theta_n = \theta + n^{-1/2}J_{m_n}^{-1}\Delta_n$ and $r_{n}(\tilde\theta_n) = o_{P_{m_n}^n}(1)$, we have $\log {\rm pl}_{n, m_n}(\tilde\theta_n)= \log {\rm pl}_{n, m_n}(\theta_0) + \Delta_n^{\rm t} J_{m_n}^{-1}\Delta_n - \frac{1}{2}\Delta_n^{\rm t} J_{m_n}^{-1}\Delta_n + r_{n}(\tilde{\theta}_{n})$. By definition, $\log {\rm pl}_{n, m_n}(\widehat{\theta}_{m_n,n})$ is larger than $\log {\rm pl}_{n, m_n}(\tilde\theta_{n})$. Hence \begin{equation*} \widehat{h}_n^{\rm t} \Delta_n - \frac{1}{2}\widehat{h}_n^{\rm t} J_{m_n}\widehat{h}_n - \frac{1}{2}\Delta_n^{\rm t} J_{m_n}^{-1}\Delta_n \ge -o_{P_{m_n}^n}(\|\widehat{h}_n\| + 1)^2. \end{equation*} The left hand side of the preceding display is equal to the left hand side of \begin{equation*} -\frac{1}{2}\left(\widehat{h}_n - J_{m_n}^{-1}\Delta_n\right)^{\rm t} J_{m_n}\left(\widehat{h}_n - J_{m_n}^{-1}\Delta_n\right) \le -\frac{c}{2} \|\widehat{h}_n - J_{m_n}^{-1}\Delta_n\|^2, \end{equation*} where $0 < c \le \lambda_{\min}(J_{m_n})$ for all sufficiently large $n$, where $\lambda_{\min}$ is the smallest eigenvalue of $J_{m_n}$, and we use that these are bounded below for $n$ sufficiently large. Combination of the preceding two displays yields \begin{equation*} \|\widehat{h}_n -J_{m_n}^{-1}\Delta_n\| = o_{P_{m_n}^n}(\|\widehat{h}_n\| + 1). \end{equation*} Since $\|J_{m_n}^{-1}\Delta_n\| = O_{P_{m_n}^n}(1)$, the triangle inequality implies that \begin{equation*} \|\widehat{h}_n\| \le \|\widehat{h}_n - J_{m_n}^{-1}\Delta_n\| + \|J_{m_n}^{-1}\Delta_n\| = o_{P_{m_n}^n}(\|\widehat{h}\|_n + 1) + O_{P_{m_n}^n}(1) =O_{P_{m_n}^n}(1). \end{equation*} Using this in the penultimate display yields $\|\widehat{h}_n - J_{m_n}^{-1}\Delta_n\| = o_{P_{m_n}^n}(1)$, implying (ref). \end{proof} \begin{remark} In the proof of Proposition (ref) the expansion (ref) is used for two different $\tilde\theta_n$, namely $\tilde\theta_n = \widehat{\theta}_{m_n,n}$ and $\tilde\theta_n \coloneqq \breve{\theta}_{m_n, n} \coloneqq n^{-1/2}J_{m_n}^{-1}\Delta_n$, where $\Delta_n = n^{-1/2}\sum_{i=1}^n\tilde{\ell}_{m_n}(X_i)$. Under Assumptions (ref)-(ref), $\sqrt{n}(\breve{\theta}_{m_n, n} - \theta_0) = J_{m_n}^{-1}\Delta_n = O_{P_{m_n}^n}(1)$, as noted in the proof. Therefore, if one establishes that also $\sqrt{n}(\widehat\theta_{m_n,n} - \theta) =O_{P_{m_n}}(1)$, then it suffices to show that $r_{n}(\tilde\theta_n) = o_{P_{m_n}}(1)$ for all $\tilde{\theta}_n$ such that $\sqrt{n}(\tilde\theta_n - \theta) = O_{P_{m_n}}(1)$. \end{remark} \begin{proof}[Proof of Theorem (ref)] Let $\tilde{h}_n = \tilde{\theta}_n - \theta_0$. For fixed $\tilde{\psi}$, a Taylor expansion yields $\mathbb{P}_n l_{m_n}(\tilde{\theta}_n,\tilde{\psi}) - \mathbb{P}_n l_{m_n}(\theta_0,\tilde{\psi}) = \tilde{h}_n^{{\rm t}}\mathbb{P}_n \dot{l}_{m_n}(\theta_0 ,\tilde{\psi}) + \tfrac{1}{2} \tilde{h}_n^{{\rm t}} \mathbb{P}_n \ddot{l}_{m_n}(s_n,\tilde{\psi})\tilde{h}_n$, for a $s_n$ between $\tilde\theta_n$ and $\theta_0$. Multiplying through by $n$ and replacing $\mathbb{P}_n \ddot{l}_{m_n}(s_n, \tilde{\psi})$ by $-J_{m_n} + o_{P_{m_n}^n}(1)$ and $\sqrt{n}\,\mathbb{P}_n \dot{l}_{m_n}(\theta_0,\tilde{\psi})$ by $\sqrt{n}\mathbb{P}_n \tilde{\ell}_{m_n} + o_{P_{m_n}^n}(\sqrt{n}\|\tilde{h}_n\|+1)$ on the right hand side gives \begin{equation} n\mathbb{P}_n l_{m_n}(\tilde{\theta}_n,\tilde{\psi}) - n\mathbb{P}_n l_{m_n}(\theta_0,\tilde{\psi}) = \tilde{h}_n^{{\rm t}}\sum_{i=1}^n \tilde{\ell}_{m_n}(X_i) - \tfrac{1}{2} n\tilde{h}_n^{{\rm t}} J_{m_n}\tilde{h}_n + o_{P_{m_n}^n}((\sqrt{n}\|\tilde{h}_n\| + 1)^2), \notag \end{equation} for $\tilde{\psi}$ equal to either $(\tilde\theta_n, \widehat{\gamma}(\tilde\theta_n)$ or $(\theta_0, \widehat{\gamma}(\theta_0))$. Applying the sandwiching bounds in (ref) and (ref) gives (ref). \end{proof} \subsection{Sufficient conditions for Theorem (ref)} As demonstrated in the examples in the main text, in some models, our approach of working under the parametric models $P_{m}$ allows (ref) to be established directly. For cases where this is not possible, Proposition (ref) and Theorem (ref) provide a general result for such estimators. This result is based on murphy2000profile with the key difference being that we consider a sieved profile likelihood in which, at each step $m$, true least-favourable submodels necessarily exist. This avoids the requirement to construct {`}approximate least favourable submodels{'} as in murphy2000profile. Nevertheless, as the theoretical analysis of this contiguous sieve profile likelihood estimator proceeds very similarly to the analysis of the semiparametric profile likelihood estimator considered in murphy2000profile, Theorem (ref) states the result under high level conditions. Here we give lower-level structural conditions which imply the conditions (ref) and (ref) required by Theorem (ref). The conditions are similar to those given by murphy2000profile, but require some adjustment as we deal with a sequence of parametric likelihoods. The following lemma splits condition (ref) into a no-bias condition and a condition relating to the empirical process $\mathbb{G}_n$, here defined by $\mathbb{G}_n f = \sqrt{n}(\mathbb{P}_n f - P_{m_n}f)$. \begin{lemma} Suppose that for any $\tilde\theta_n = \theta_0 + o_{P_{m_n}^n}(1)$, \begin{equation} \mathbb{E}_{m_n}\dot{l}_{m_n}(\theta_0, \tilde\theta_n, \widehat{\gamma}(\tilde\theta_n)) =o_{P_{m_n}^n}(\|\tilde\theta_n - \theta_0\| + n^{-1/2}) , \end{equation} and \begin{equation} \mathbb{G}_n \dot{l}_{m_n}(\theta_0, \tilde\theta_n, \widehat{\gamma}(\tilde\theta_n)) = \mathbb{G}_n\tilde{\ell}_{m_n} + o_{P_{m_n}^n}(1). \end{equation} Then (ref) holds. \end{lemma} \begin{proof}\belowdisplayskip=-12pt By (ref) and (ref), and since $\mathbb{E}_{m_n}\tilde{\ell}_{m_n} = 0$, \begin{align*} \sqrt{n}\mathbb{P}_n\dot{l}_{m_n}(\theta_0, \tilde\theta_n, \widehat{\gamma}(\tilde\theta_n)) &= \mathbb{G}_n\dot{l}_{m_n}(\theta_0, \tilde\theta_n, \widehat{\gamma}(\tilde\theta_n)) + \sqrt{n}\mathbb{E}_{m_n} \dot{l}_{m_n}(\theta_0, \tilde\theta_n, \widehat{\gamma}(\tilde\theta_n))\\ &= \mathbb{G}_n\tilde{\ell}_{m_n} + o_{P_{m_n}^n}(\sqrt{n}\|\tilde\theta_n - \theta_0\| +1)\\ &=\sqrt{n}\mathbb{P}_n \tilde{\ell}_{m_n}+ o_{P_{m_n}^n}(\sqrt{n}\|\tilde\theta_n - \theta_0\| +1). \end{align*} \end{proof} Condition (ref) is a {`}no-bias{'} condition, cf. the discussion in murphy2000profile. Condition (ref) can be shown to hold if the nuisance parameter estimator $\widehat{\gamma}$ satisfies a consistency condition and a stochastic equicontinuity type condition is satisfied. For any $m$, $\gamma = \gamma_m\in \mathbb{R}^{k_m}$ may be viewed as an eventually zero sequence in $\mathbb{R}^\mathbb{N}$. Let $\Gamma_{m}^\mathbb{N}$ denote the subset of $\mathbb{R}^\mathbb{N}$ corresponding to vectors in $\Gamma_m$. We equip $\mathbb{R}^\mathbb{N}$ with a topology, $\tau$, and shall need the following consistency condition to hold: For any $\tilde\theta_n = \theta_0 + o_{P_{m_n}^n}(1)$, \begin{equation} \widehat{\gamma}(\tilde\theta_n) -\gamma_0 = o_{P_{m_n}^n}(1). \end{equation} That is, for any neighbourhood $\mathcal{U}$ of zero in $(\mathbb{R}^\mathbb{N}, \tau)$, $\lim_{n\to\infty} P_{m_n}( \widehat{\gamma}(\tilde\theta_n) -\gamma_0 \in \mathcal{U})=1$, for $\widehat{\gamma}(\tilde\theta_n), \gamma_0$ considered as elements in $\mathbb{R}^\mathbb{N}$. The topology $\tau$ is arbitrary. It may, for instance, be that induced by $\mathcal{H}$ (if $\mathcal{H}$ is a topological vector space). However, for this condition to be useful, the topology needs to be strong enough to imply certain continuity conditions. \begin{lemma} Suppose that $\tilde\theta_n = \theta_0 + o_{P_{m_n}^n}(1)$ and that (ref) holds. Let $\mathcal{V}$ be a neighbourhood of $(0, 0)\in \mathbb{R}^{p}\times \mathbb{R}^\mathbb{N}$. Suppose that that for each $\varepsilon, \upsilon >0$ there are random $\Delta_n(\varepsilon, \upsilon)\ge0$ and $N(\varepsilon, \upsilon)$ such that if $n\ge N(\varepsilon, \upsilon)$ then {\rm(i)} $P_{m_n}(\Delta_n(\epsilon, \upsilon) > \upsilon)<\varepsilon$ and {\rm(ii)} \begin{equation*} \sup_{v\in \mathcal{V}_n}\| \mathbb{G}_n \dot{l}_{m_n}(\theta_0, \psi_0 + v ) - \mathbb{G}_n \dot{l}_{m_n}(\theta_0, \psi_0) \|\le \Delta_n(\varepsilon, \upsilon), \qquad \psi_0 \coloneqq (\theta_0, \gamma_0), \end{equation*} where $\mathcal{V}_n\coloneqq \{v\in \mathcal{V} : \psi_0 + v \in \Theta\times \Gamma_{m_n}^\mathbb{N}\}$. Then (ref) holds. \end{lemma} \begin{proof} Let $X_n(v) \coloneqq \mathbb{G}_n \dot{l}_{m_n}(\theta_0, \psi_0 + v)$ and note that $X_n(0) = \mathbb{G}_n \dot{l}_{m_n}(\theta_0, \psi_0 ) = \mathbb{G}_n \tilde{\ell}_{m_n}$. Fix $\varepsilon, \upsilon>0$. By (ref) there is a $N_1$ such that $n\ge N_1$ implies $P_{m_n}\left(\hat{v}_n\in \mathcal{V}\right) \ge 1 -\varepsilon/2$, where $\hat{v}_n \coloneqq (\tilde\theta_n - \theta_0, \widehat{\gamma}(\tilde\theta_n) - \gamma_0)$. Note that, by definition, if $\hat{v}_n\in \mathcal{V}$ then $\hat{v}_n\in \mathcal{V}_n$. Using this, (i) and (ii) we have for all $n\ge \max\{N_1, N(\varepsilon/2, \upsilon)\}$, \begin{align*} P_{m_n}\left(\|X_n(\hat{v}) - X_n(0)\| >\upsilon\right) &\le P_{m_n}\left(\hat{v}_n\notin \mathcal{V}\right) + P_{m_n}\left( \|X_n(\hat{v}) - X_n(0)\| >\upsilon, \hat{v}_n \in \mathcal{V}_n \right)\\ &< \varepsilon/2 + P_{m_n}\big(\sup_{v\in \mathcal{V}_n} \|X_n(v) - X_n(0)\|>\upsilon\big)\\ &\le \varepsilon/2 + P_{m_n}\left(\Delta_n(\varepsilon/2, \upsilon) > \upsilon\right)<\varepsilon, \end{align*} as required. \end{proof} Finally, condition (ref) is an approximate information equality. Provided the information equality approximately holds in the least-favourable parametric submodels, this will hold under continuity and moment conditions. \begin{lemma} Suppose that $\tilde\vartheta_n = \theta_0 + o_{P_{m_n}^n}(1)$, $\tilde\theta_n = \theta_0 + o_{P_{m_n}^n}(1)$, that (ref) holds and \begin{enumerate}[label = {\rm(\roman*)}] • $P_m [\dot{l}_{m}(\theta_0, \theta_0, \gamma_0)\dot{l}_{m}(\theta_0, \theta_0,\gamma_0)^{\rm t}] = -P_{m}[\ddot{l}_{m}(\theta_0, \theta_0, \gamma_0)] + o(1)$ as $m\to\infty$; • $(\tilde\vartheta_n, \tilde\theta_n, \widehat{\gamma}(\tilde\theta_n)) - (\theta_0, \theta_0, \gamma_0)= o_{P_{m_n}^n}(1)$ implies that \begin{equation*} \lim_{n\to\infty} \mathbb{E}_{m_n}\|\ddot{l}_{m_n}(\tilde\vartheta_n, \tilde\theta_n, \widehat{\gamma}(\tilde\theta_n))- \ddot{l}_{m_n}(\theta_0, \theta_0, \gamma_0)\|= 0; \end{equation*} • $\|\ddot{l}_{m_n}(\theta_0, \theta_0, \gamma_0)\|$ is uniformly $P_{m_n}$-integrable. \end{enumerate} Then (ref) holds. \end{lemma} \begin{proof} By the construction of the least favourable submodels and (ref), \begin{equation*} J_{m} = \mathbb{E}_m\,\tilde{\ell}_{m}\tilde{\ell}_{m}^{\rm t} = \mathbb{E}_m\,\dot{l}_{m}(\theta_0, \theta_0, \gamma_0)\dot{l}_{m}(\theta_0, \theta_0,\gamma_0)^{\rm t} = - \mathbb{E}_m\,\ddot{l}_{m}(\theta_0, \theta_0, \gamma_0) + o(1). \end{equation*} Write $Y_{n, i}\coloneqq \ddot{l}_{m_n}(\theta_0, \theta_0, \gamma_0)(X_i)$ and $\tilde{Y}_{n, i}\coloneqq \ddot{l}_{m_n}(\tilde\vartheta_n, \tilde\theta_n, \widehat{\gamma}(\tilde\theta_n))(X_i)$. Then (ref) holds if both (a) $n^{-1}\sum_{i=1}^n (\tilde{Y}_{n,i} - Y_{n, i})$ and (b) $n^{-1}\sum_{i=1}^n( Y_{n,i} - \mathbb{E}_{m_n}Y_{n, i})$ are $o_{P_{m_n}^n}(1)$. Requirement (a) follows as condition (ref) implies $\mathbb{E}_{m_n}\|\tilde{Y}_{n,i} - Y_{n, i}\| \to 0$ and hence $n^{-1}\sum_{i=1}^n( \tilde{Y}_{n,i} - Y_{n, i})$ converges to zero in probability by Markov's inequality. (b) holds as under condition (ref) $n^{-1}\sum_{i=1}^n( Y_{n,i} - \mathbb{E}_{m_n}Y_{n, i})$ converges to zero by the weak law of large numbers gut1992weaklaw. \end{proof} \section{The applications} \subsection{The partially linear model} In this section we first verify that (ref) and (ref) hold in the setup of Section (ref). This entails, by way of Lemma (ref), that the partially linear model satisfies Assumption (ref). We then establish that $\tilde{\ell}_{\theta_0,T_m^{-1}\gamma_0}$ converges to the semiparametric efficient score. Consider submodels of the form $\tau \mapsto p_{\theta_0 + a \tau, \eta_0 + b \tau}$ where $a \in \mathbb{R}$ and $b \in B \subset \mathcal{H}$. Differentiating with respect to $\tau$, and evaluating in $\tau = 0$ \begin{equation} \frac{{\rm d}}{{\rm d} \tau} \log p_{\theta + a \tau, \eta + b \tau} \big|_{\tau = 0} = a \dot{\ell}_{\theta,\eta} + \frac{b(z)}{\sigma^2}( y - \eta(z) - \theta w), \notag \end{equation} where $\dot{\ell}_{\theta,\eta}(x) = \sigma^{-2}w(y - \eta(z) - \theta w)$. We now show that the two terms on the right are the limits, in the sense of (ref) and (ref), of their parametric counterparts $\dot{\ell}_{\theta,T_m^{-1}\gamma}$ and $\dot{v}_{\theta,T_m^{-1}\gamma}$ given in (ref). To this end, we use Lemma (ref), and note for future reference that $d_{\rm TV}(P_{m_n}, P_0)\to 0$ by Assumption (ref) and (ref), and that for any $b \in B$, the sequence defined by $\tilde{b}_m(z) = \beta_m^{\rm t}(z)\mathbb{E}\,\{b(Z) \beta_{m}(Z)\}$ is such that $\tilde{b}_m \to b$ in $L_2(P_Z)$. Note also that under $P_m$, $\dot{\ell}_{m}(X) \sim \sigma^{-2}W\varepsilon$ and with $\tilde{\gamma} = \mathbb{E}\,\{b(Z)\beta_m(Z)\}$, $\tilde{\gamma}^{\rm t} \dot{v}_{m}(X) \sim \sigma^{-2}\tilde{b}_m(Z)\varepsilon$. The uniform square $P_{m_n}$-integrability required by Lemma (ref) then follows from the integrability condition on $W$ and (ref). Additionally, it is straightforward to check that $\dot{\ell}_{\theta, \eta}(X)$ and $b(Z)(Y - \eta(Z) - \theta W)$ are square integrable under $P_0$. Finally we show $L_1(P_0)$ convergence of the parametric scores to the semiparametric scores. In particular, \begin{equation*} \int \|\dot{\ell}_m -\dot{\ell} \|\,\mathrm{d}P_0 =\sigma^{-2} \,\mathbb{E}\, \|W (\eta_0(Z) - \eta_m(Z))\| \to 0, \end{equation*} by the Cauchy--Schwarz inequality as $\eta_m(z) \coloneqq \beta_m(z)^{\rm t}\gamma_0 \to \eta_0(z)$ in $L_2(P_Z)$. Similarly with $\tilde{\gamma}= \mathbb{E}\,\{b(Z)\beta_m(Z)\}$ as above, \begin{equation*} \int \|\tilde{\gamma}^{\rm t}\dot{v}_{m} -Db \|\,\mathrm{d}P_0= \sigma^{-2} \mathbb{E} \,\|\epsilon(\tilde{b}_m(Z) - b(Z)) + \tilde{b}_m(Z)(\eta_0(Z) - \eta_m(Z))\| \to 0, \end{equation*} as $\tilde{b}_m(z)\to b(z)$ in $L_2(P_Z)$. Applying Lemma (ref) then verifies (ref) and (ref) of Lemma (ref), meaning that Assumption (ref) holds. To verify that $\tilde{\ell}_{m}$ converges to $\tilde{\ell}$ in $L_1(P_0)$ as claimed in Section (ref) we note that similarly, \begin{equation*} \int \|\tilde{\ell}_{m} - u\|\,\mathrm{d}P_0 = \sigma^{-2} \mathbb{E}\,\{\| \varepsilon (b_0(Z) - b_m(Z)) + (W-b_m(Z))(\eta_0(Z) - \eta_m(Z))\} \to 0. \end{equation*} To establish (ref), we use Theorem (ref). We have \begin{equation*} l_m(t, \theta, \gamma)(x)= - \frac{1}{2\sigma^2}\left(y - w t - \beta_m(z)^{\rm t} ( \gamma + \mathbb{E}\,\{\beta_m(Z)W\}(\theta - t))\right)^2 + C(w, z), \end{equation*} for $C(w, z)$ a term which does not depend on $(t, \theta, \gamma)$. Thus \begin{align*} & \dot{l}_m(t, \theta, \gamma)(x) = \frac{1}{\sigma^2}\left(y - w t - \beta_m(z)^{\rm t} ( \gamma + \mathbb{E}\,\{\beta_m(Z)W\}(\theta - t))\right)\\ & \qquad \qquad\qquad \qquad \times \left( w - \beta_m(z)^{\rm t} \mathbb{E}\,\{\beta_m(Z)W\}\right), \end{align*} and \begin{equation*} \ddot{l}_{m}(t, \theta, \gamma)(x) = -\frac{1}{\sigma^2}\left(w - \beta_m(z)^{\rm t} \mathbb{E}\,\{\beta_m(Z)^{\rm t} W\}\right)^2 = -\sigma^{-2}\left(w - b_m(z)\right)^2. \end{equation*} We first note that $\mathbb{E}_m \ddot{l}_m(t, \theta_0, \gamma_0) = -J_{m}$. Thus for the second part of the condition in Theorem (ref) it suffices to show that $n^{-1}\sum_{i=1}^n \sigma^{-2}(W_i - b_{m_n}(Z_i))^2 - J_{m_n} =o_{P_{m_n}^n}(1)$. This follows from the weak law of large numbers as the $\sigma^{-2}(W_i - b_{m_n}(Z_i))^2$ are uniformly $P_{m_n}$-integrable (see gut1992weaklaw). To establish the first part of Theorem (ref), let $h\in \mathbb{R}$ and set $\tilde\theta_n = \theta_0 + h / \sqrt{n}$. We then have \begin{align*} & \dot{l}_{m_n}(\theta_0, \tilde\theta_n, \hat{\gamma}(\tilde\theta_n))(X_i)\\ & \qquad = \frac{1}{\sigma^2} \bigg(Y_i - W_i \theta_0 - \beta_{m_n}(Z_i)^{\rm t} \bigg[\widehat{\gamma}(\tilde\theta_n)+ \mathbb{E}\,\{\beta_m(Z)W\}\frac{h}{\sqrt{n}}\bigg]\bigg)\left(W_i - b_m(Z_i)\right), \end{align*} where $\widehat{\gamma}(\theta)$ and $B_{m,n}$ are as defined in Section (ref). Therefore, the difference \begin{align*} & \dot{l}_{m_n}(\theta_0, \tilde\theta_n, \widehat{\gamma}(\tilde\theta_n))(X_i) - \tilde{\ell}_{m_n}(X_i) \\ &\; = \frac{1}{\sigma^2}[W_i - b_{m_n}(Z_i)]\beta_{m_n}(Z_i)^{\rm t}\\ & \qquad \times \left[ (\widehat{\gamma}(\theta_0) - \gamma) + \frac{h}{\sqrt{n}}\left(\mathbb{E}\,\{\beta_{m_n}(Z)W\} - B_{m_n, n}^{-1}\frac{1}{n}\sum_{i=1}^n \beta_{m_n}(Z_i) W_i\right) \right]. \end{align*} In consequence, to verify the remaining condition of Theorem (ref), it suffices to show that \begin{equation} \begin{split} &\frac{1}{\sqrt{n}}\sum_{i=1}^n (W_i - b_{0}(Z_i))\beta_{m_n}(Z_i)^{\rm t} (\widehat{\gamma}(\theta_0) - \gamma); \\ &\frac{1}{\sqrt{n}}\sum_{i=1}^n (b_{0}(Z_i)- b_{m_n}(Z_i))\beta_{m_n}(Z_i)^{\rm t} (\widehat{\gamma}(\theta_0) - \gamma);\\ &\frac{1}{n}\sum_{i=1}^n (W_i - b_{m_n}(Z_i))\beta_{m_n}(Z_i)^{\rm t} \left[\mathbb{E}\,\{\beta_{m_n}(Z)W\} - B_{m_n, n}^{-1}\frac{1}{n}\sum_{i=1}^n \beta_{m_n}(Z_i)W_i\right], \end{split} \end{equation} are all $o_{P_{m_n}}(1)$. In order to verify this we impose the following assumptions (which may be relaxed if more conditions are imposed upon $W$, see below). Let $\xi_{m}\coloneqq \sup_{z\in [0, 1]}\|\beta_{m_n}(z)\|$ and let $a_n$ be such that $\|b_0(Z) - b_{m_n}(Z)\|_{L_2(\nu)} \le a_{n}$. (Upper bounds on $a_n$ are available from approximation theory under, for example, smoothness conditions on $b_0$.) \begin{enumerate}[label=(pl\arabic*), series=plmconditions]\itemsep-0.2em • $\mathbb{E}\,(W^2\,|\, Z) \le C$, $P_Z$-almost surely; • $k_{m_n}^2 / \sqrt{n}\to 0$, $\xi_{m_n}k_{m_n} / \sqrt{n}\to 0$ and $a_{n}\sqrt{k_{m_n}}\xi_{m_n} \to 0$. \end{enumerate} For the first sum in (ref), note that as for $k\neq i$ \begin{equation*} \mathbb{E}\,\left\{ (W_i - b_0(Z_i))\beta_{m_n, j}(Z_i) (W_k - b_0(Z_k))\beta_{m_n, j}(Z_k) \right\} = 0 \end{equation*} then for some positive constant $C_0$ \begin{align*} P_{m_n}\big( \sum_{i=1}^n (W_i - b_{0}(Z_i))\beta_{m_n, j}(Z_i) \ge \sqrt{n}M \big) &\le \frac{2\mathbb{E}\,(\mathbb{E}\,(W^2 \,|\, Z) + b_0(Z)^2) \beta_{m_n, j}(Z)^2}{M}\le \frac{C_0}{M}, \end{align*} which can be made arbitrary small by taking $M$ large enough. We also have \begin{equation} \sum_{j=1}^{k_{m_n}} |\widehat{\gamma}_j(\theta_0) - \gamma_j| \le \sqrt{k_{m_n}} \big(\sum_{j=1}^{k_{m_n}}|\widehat{\gamma}_j(\theta_0) - \gamma_j|^2\big)^{1/2} =O_{P_{m_n}}(k_{m_n} / \sqrt{n}) \end{equation} by the Cauchy--Schwarz inequality and Theorem 4.1 in belloni2015some. It follows that by taking $M = M_n =k_{m_n}$ above and using the union bound that \begin{equation*} \sum_{j=1}^{k_{m_n}} (\hat{\gamma}_j(\theta_0) - \gamma_j) \frac{1}{\sqrt{n}}\sum_{i=1}^{n} \beta_{m_n, j}(Z_i) (W_i - b_0(Z_i)) = O_{P_{m_n}}(k_{m_n}^2 n^{-1/2})= o_{P_{m_n}}(1), \end{equation*} which verifies the first term. The second holds as by Theorem 4.1 in belloni2015some again, \begin{align*} &\big|\frac{1}{\sqrt{n}}\sum_{i=1}^n (b_{0}(Z_i)- b_{m_n}(Z_i))\beta_{m_n}(Z_i)^{\rm t} (\hat{\gamma}(\theta_0) - \gamma)\big|\\ & \qquad \le \sqrt{n}\xi_{m_n} \|\hat{\gamma}(\theta_0) - \gamma\| \frac{1}{n}\sum_{i=1}^n |b_{0}(Z_i)- b_{m_n}(Z_i)|\\ &\qquad = O_{P_{m_n}}\left(\sqrt{n}\xi_{m_n}\sqrt{k_{m_n}}n^{-1/2}a_n \right) = O_{P_{m_n}}\left(\xi_{m_n}\sqrt{k_{m_n}}a_n \right) = o_{P_{m_n}}(1), \end{align*} as $P_{m_n}\big( n^{-1}\sum_{i=1}^n |b_0(Z_i) - b_{m_n}(Z_i)| \ge Ma_n \big) \le \|b_0 - b_{m_n}\|_{L_1(P_Z)}/M a_n \le 1/M$. Finally, we can rewrite the third sum in (ref) as $ n^{-1}\sum_{i=1}^n (W_i - b_{m_n}(Z_i)) \beta_{m_n}(Z_i)^{\rm t} (\pi - \widehat{\pi})$, where $\pi = \mathbb{E}\,\{\beta_{m_n}(Z)W\}$ and $\widehat{\pi}$ is its empirical counterpart. Using, once more, Theorem 4.1 in belloni2015some we have \begin{align*} \big|\frac{1}{n}\sum_{i=1}^n (W_i - b_{m_n}(Z_i)) \beta_{m_n}(Z_i)^{\rm t} (\pi - \widehat{\pi})\big| &\le \xi_{m_n} \|\pi - \widehat{\pi}\|_{2} \frac{1}{n}\sum_{i=1}^n |W_i - b_{m_n}(Z_i) |\\ &= O_{P_{m_n}}(\xi_{m_n}\sqrt{k_{m_n} / n}), \end{align*} as $P_{m_n}\big( n^{-1}\sum_{i=1}^n |W_i - b_{m_n}(Z_i) | \ge M \big) \le M^{-1} (\mathbb{E}\,|W| + \mathbb{E}\, |b_0(Z) | + a_n)$. Here we remark that the rate conditions in (ref) can be relaxed if we impose additional conditions on $W$. In particular, if we replace (ref) with \begin{enumerate} • $W$ is sub-exponential with parameters $(\nu, \alpha)$, $\xi_{m_n}k_{m_n} / \sqrt{n}\to 0$ & $a_{n}\sqrt{k_{m_n}}\xi_{m_n} \to 0$ \end{enumerate} then the second and third terms can be shown to be $o_{P_{m_n}}(1)$ exactly as before but we can refine the argument relating to the first term. In particular, (ref) holds as above, but the probabilistic bound on $n^{-1/2}\sum_{i=1}^n (W_i - b_{0}(Z_i))\beta_{m_n, j}(Z_i)$ can be improved. Since $b_0(Z_i)$ is bounded by, say $C_0$ and $\beta_{m_n, j}(Z_i)$ is bounded by $\xi_{m_n}$, $(W - b_{0}(Z))\beta_{m_n, j}(Z)$ is also subexponential with parameters $(\xi_{m_n}(\nu + C_0), \alpha)$. Hence for all $t\ge 0$ the Bernstein-type bound below holds (e.g., wainwright2019high) \begin{equation*} P_{m_n}\left(\left| \frac{1}{n}\sum_{i=1}^n (W_i - b_{0}(Z_i))\beta_{m_n, j}(Z_i)\right| \ge t \right) \le 2\exp\left( -\min\left\{\frac{nt}{2\alpha}, \frac{nt^2}{2\xi_{m_n}^2(\nu + C_0)^2}\right\} \right). \end{equation*} Hence taking $t = t_n = 2(\nu + C_0)\xi_{m_n} \sqrt{\log k_{m_n}} / \sqrt{n}$ we have that \begin{align*} &k_{m_n} P_{m_n}\left(\left| \frac{1}{n}\sum_{i=1}^n (W_i - b_{0}(Z_i))\beta_{m_n, j}(Z_i)\right| \ge t \right) \\ &\qquad\qquad\qquad \le k_{m_n}2\exp\left(-\frac{n n^{-1} 4\xi_{m_n}^2(\nu + C_0)^2 \log k_{m_n}}{2\xi_{m_n}^2(\nu + C_0)^2}\right)\\ &\qquad\qquad\qquad= 2k_{m_n} \exp(-2 \log k_{m_n}) = 2/k_{m_n}\to0. \end{align*} Thus by (ref) and the union bound \begin{equation*} \sum_{j=1}^{k_{m_n}} (\widehat{\gamma}_j(\theta_0) - \gamma_j) \frac{1}{\sqrt{n}}\sum_{i=1}^{n} \beta_{m_n, j}(Z_i) (W_i - b_0(Z_i)) = O_{P_{m_n}}\big(\frac{\xi_{m_n} \sqrt{\log k_{m_n}} k_{m_n}}{n}\big)= o_{P_{m_n}}(1). \end{equation*} Finally we provide an example of $\beta_{m_n}$ for which $h_{m_n, n} \to h$ in $L_2(P_0)$. In particular, let $\beta_{m_n}$ be (orthonormalised) B-splines of degree $l$ with equally spaced knots. Then if \begin{enumerate}[resume*=plmconditions, label = (pl\arabic*)] • $\eta$ is $s$-times continuously differentiable and $l+1\ge s$$k_{m_n}^{-s}\sqrt{n}\to 0$ \end{enumerate} it follows from Theorem 20.3 in powell1981approximation that if $\gamma$ is chosen optimally then $\|\beta_{m_n}(z)^{\rm t} \gamma - \eta_0(z)\|_{\infty} =O(k_{m_n}^{-s})$ un turn implying that $\sqrt{n} \|\beta_{m_n}^{\rm t} \gamma - \eta_0\|_{L_2(P_Z)} = O(\sqrt{n}k_{m_n}^{-s}) = o(1)$, and thus $h_{m_n, n} \to 0$ in $L_2(P_0)$. \begin{remark} Locally constant basis functions are easy to work with, and provide very intelligible results. Consider therefore the sequence of orthonormal basis function $\beta_{m,j}$ that take the form $\beta_{m,j}(z) = I_{V_{m,j}}(z)/(\mathbb{E}\, I_{V_{m,j}}(Z) )^{1/2}$ for $j = 1,\ldots,k_m, \, k_m \geq 1$, where $V_{m,j} = \{z \colon (j-1)/k_m \leq z < j/k_m\}$ for $j = 1,\ldots,k_m$, and $V_{m,1}\cup \cdots \cup V_{m,k_m} = [0,1]$ for all $m$. Consider the following conditions: \begin{enumerate}[label=(lc\arabic*)]\itemsep-0.2em • $Z$ has density $f_Z$ that is continuously differentiable and positive on $[0,1]$; • The function $z \mapsto \mathbb{E}\,(W^k \,|\, Z = z)$ is continuously differentiable for $k=1,2$; • $\eta_0 \in \mathcal{H}$, where $\mathcal{H}$ consists of all continuously differentiable functions $\eta \colon [0,1] \to \mathbb{R}$. \end{enumerate} With these basis function, one may use Theorem (ref) to show that the quadratic expansion in (ref) holds provided the subsequence $(m_n)$ is chosen so that $n/k_{m_n} \to \infty$; and that Assumption (ref), i.e., contiguity, is satisfied provided $\sqrt{n}/k_{m_n} \to 0$. \end{remark} \subsection{The Cox model} In this appendix we provide some details left out of Section (ref). We need to show that $\log {\rm d} P_{m_n}^n/{\rm d} P_0^n = o_{P_0}(1)$; that the scores are indeed uniformly $P_{m_n}$-integrable; and that $A_{n}^{\rm cox}(h)$ admits a quadratic expansion, as claimed. Introduce \begin{equation} S_n^{(k)}(t,\theta) = \frac{1}{n}\sum_{i=1}^n Y_i(t) W_i^k \exp(\theta W_i) ,\quad for $k = 0,1,2$, \notag \end{equation} and define $s_{\theta,\eta}^{(k)}(t) = \mathbb{E}_{\theta,\eta}\,Y(t)W^k \exp(\theta W)$ for $k = 0,1,2,3$, so that due to the data being i.i.d., $s_{\theta,\eta}^{(k)}(t) = \mathbb{E}_{\theta,\eta}\,S_n^{(k)}(t,\theta)$. Write $s^{(k)} = s_{\theta_0,\eta_0}^{(k)}$ and $s_m^{(k)} = s_{\theta_0,T_m^{-1}\gamma_0}^{(k)}$, and define \begin{equation} v_{m,j} = \frac{\int_{V_{m,j}}s_m^{(2)}(s)\,{\rm d} s}{\int_{V_{m,j}}s_m^{(0)}(s)\,{\rm d} s} - \bigg(\frac{\int_{V_{m,j}}s_m^{(1)}(s)\,{\rm d} s}{\int_{V_{m,j}}s_m^{(0)}(s)\,{\rm d} s}\bigg)^2,\quad and\quad v(t) = \frac{s^{(2)}(t)}{s^{(0)}(t)} - \bigg(\frac{s^{(1)}(t)}{s^{(0)}(t)}\bigg)^2. \notag \end{equation} We make the following assumptions: \begin{enumerate}[label=(cx\arabic*)]\itemsep-0.2em • $\mathbb{E}\, W^k \exp(\theta W) < \infty$ for $k = 0,1,2,3$, for all $\theta$ in a neighbourhood of $\theta_0$ and $\mathbb{E}[|W|^{2+\delta}]<\infty$; • The baseline hazard $\eta_0$ is continuously differentiable and positive on $[0,1]$; • In a neighbourhood $U_{\theta_0}$ of $\theta_0$, \begin{equation} \sup_{t \in [0,1] ,\theta \in U}|S_n^{(k)}(t,\theta) - s_{\theta,\eta_0}^{(k)}(t)| = o_{P_0}(1), \quadfor $k = 0,1,2$. \notag \end{equation} • For a neighbourhood $U_{\eta_0}$ of $\eta_0$, the functions $(\theta,\eta)\mapsto s_{\theta,\eta}^{(k)}(t)$ for $k = 0,1,2,3$ are continuous on $U_{\eta_0}\times U_{\theta_0}$, uniformly in $t \in [0,1]$; they are bounded on $U_{\eta_0}\times U_{\theta_0}\times[0,1]$, and $s_{\theta,\eta}^{(0)}$ is bounded below on $U_{\eta_0}\times U_{\theta_0}\times[0,1]$, and $J = \int_0^1 v(t) s^{(0)}(t) \eta_0(t)\,{\rm d} t$ is positive. \end{enumerate} Assumption (ref), (ref), and (ref) are similar to Assumption A., B., and D. of the canonical Cox regression paper andersen1982cox. We start with the details involved in showing Assumption (ref). Let $(m_n)_{n \geq 1}$ be such that $\sqrt{n}/k_{m_n} \to 0$, and recall that \begin{equation} \log \frac{{\rm d} P_{m_n}^n}{{\rm d} P_0^n} = \frac{1}{\sqrt{n}}\sum_{i=1}^n \int_0^1 \frac{h_{m_n,n}}{\eta_0}\,{\rm d} M_i - \frac{1}{2n}\sum_{i=1}^n \int_0^1 \frac{h_{m_n,n}^2}{\eta_0^2}\,{\rm d} N_i + r_{m_n,n}, \end{equation} with $h_{m,n}(t) = \sqrt{n}( \beta_m(t)^{{\rm t}}\gamma_0 - \eta_0(t))$; $r_{m,n} = n^{-1}\sum_{i=1}^n\int_0^1 h_{m,n}^2/\eta_0^2R(h_{m,n}/(\sqrt{n}\eta_0))\,{\rm d} N_i$; and $R$ is the function defined via the Taylor expansion $\log( 1 + a) = a - \tfrac{1}{2} a^2 + a^2 R(a)$, in particular $R(a) = \tfrac{1}{2} \big(1 - 1/(1 + \tilde{a})^2 \big)$, where $\tilde{a}$ is some point between $0$ and $a$. With the basis functions in (ref), we have that $h_{m_n,n} \to 0$ in $L_2({\rm d} t)$ if and only if $(m_n)_{n \geq 1}$ is such that $\sqrt{n}/k_{m_n}\to 0$. But then \begin{equation} \mathbb{E}_0\, \bigg( \frac{1}{\sqrt{n}}\sum_{i=1}^n \int_0^1 \frac{h_{m_n,n}(t)}{\eta_0(t)}\,{\rm d} M_i(t) \bigg)^2 = \int_0^1 \frac{h_{m_n,n}(t)^2}{\eta_0(t)^2} s^{(0)}(t) \eta_0(t) \,{\rm d} t = o(1), \notag \end{equation} by the $L_2({\rm d} t)$ convergence $h_{m_n,n} \to 0$, combined with (ref) and (ref). Similarly, the second term on the right in (ref) is a nonnegative random variable with expectation \begin{equation} \mathbb{E}_0\, \frac{1}{n}\sum_{i=1}^n \int_0^1 \frac{h_{m_n,n}^2}{\eta_0^2}\,{\rm d} N_i = \int_0^1 \frac{h_{m_n,n}(t)^2}{\eta_0(t)^2} s^{(0)}(t) \eta_0(t)\,{\rm d} t = o(1). \notag \end{equation} Thus both these terms tend to zero in probability by Markov{'}s inequality. To show that the remainder $r_{m_n,n}$ vanishes, note that since $\eta_0$ is continuous on $[0,1]$, hence uniformly continuous we can given $0 < \varepsilon < 1$ find $n_0$ such that $\sup_{t\in [0,1]}|\beta_{m_n}(t)^{{\rm t}}\gamma_0 - \eta_0(t)| < \varepsilon$ for all $n \geq n_0$. But then, by the triangle inequality $|R(n^{-1/2} h_{m_n,n}(t)^2/\eta_0(t)^2 ) | \leq \tfrac{1}{2} \big(1 + 1/(1 - \varepsilon)\big)$ for all $n \geq n_0$. Consequently, $\mathbb{E}_0\, |r_{m_n,n}| \leq \tfrac{1}{2} (1 + 1/(1 - \varepsilon) ) \int_0^1 \big(h_{m_n,n}(t)^2/\eta_0(t)^2\big) s^{(0)}(t) \eta_0(t)\,{\rm d} t$ for all $n \geq n_0$, and we have convergence to zero, and $r_{m_n,n} = o_{P_0}(1)$ by Markov{'}s inequality. It remains to verify (ref) and (ref) of Lemma (ref), which we do using Lemma (ref). To this end, consider the submodels $\tau \mapsto p_{\theta_0 + a \tau, \eta_0 + b \tau}$ where $a \in \mathbb{R}$ and $b \in B \subset \mathcal{H}$. Differentiating with respect to $\tau$, and evaluating in $\tau = 0$ \begin{equation} \frac{{\rm d}}{{\rm d} \tau} \log p_{\theta_0 + a \tau, \eta_0 + b \tau}(X) \big|_{\tau = 0} = a \dot{\ell}_{\theta_0,\eta_0} + Db, \notag \end{equation} where $\dot{\ell}(X) = W M(1)$ and $Db = \int_0^1 b(s)/\eta_0(s)\,{\rm d} M(s)$. Then \begin{align*} |\dot{\ell}_{m_n}(X) - \dot{\ell}(X) | & \leq |W|\exp(\theta_0W ) \sum_{j=1}^{k_{m_n}}\int_{V_{m_n,j}}| \eta_0(s) - \eta_0((j-1)/k_{m_n}) |\,{\rm d} s\\ & \leq |W|\exp(\theta_0W )\tfrac{1}{2} \sup_{t\in [0,1]}\eta_0^{\prime}(t)/k_{m_n}, \end{align*} which, as $k_{m_n}\to \infty$, tends to zero in probability by (ref). Next, let $(\gamma_m)_{m \geq 1}$ be a sequence of vectors such that $\gamma_m^{{\rm t}}\beta_m \to b$ uniformly on $[0, 1]$. We have the bound \begin{align*} & |\gamma_{m_n}^{{\rm t}}\dot{v}_{m_n}(X) - Db(X) |\\ &\qquad \qquad \leq |\int_0^1 \bigg( \frac{\gamma_{m_n}^{{\rm t}}\beta_{m_n}(s)}{\gamma_0^{{\rm t}}\beta_{m_n}(s)} - \frac{b(s)}{\eta_0(s)} \bigg)\,{\rm d} M^{m_n}(s)| +\sup_{s\in [0,1]}\frac{b(s)}{\eta_0(s)} \exp(\theta_0W)/k_{m_n}. \end{align*} Here the first term on the right tends to zero in probability by an application of the It{\^o} isometry followed by Markov{'}s inequality (or, alternatively, Lenglart{'}s inequality). The second term tends to zero in probability by (ref) and (ref). In order to show that both $\dot{\ell}_{m_n}^2$ and $(\gamma_{m_n}^{{\rm t}}\dot{v}_{m_n})^2$ are uniformly $P_{m_n}$-integrable, we may argue as follows. We start with $\dot{\ell}_{m_n}^2$. Since its compensator is continuous, the quadratic variation of $M^{m_n}$ is $[M^{m_n}, M^{m_n}](t) = N(t)$. Let $Z^{m_n}\coloneqq WM^{m_n}(t)$ and note that $Z^{m_n}(1) = \dot{\ell}_{m_n}(X)$. It is easy to check that $(Z^{m_n}(t))_{t\in[0, 1]}$ is a martingale. We also have that $Z^{m_n}(0) = W M^{m_n}(0) = W N(0) = 0$ a.s. since $T\ge 0$ and $T=0$ has zero probability. By the definition of quadratic variation, it is clear that $[Z^{m_n}, Z^{m_n}](t) = W^2N(t)$. Clearly $\dot{\ell}_{m_n}(X)^{2+\delta} \le \sup_{t\in [0, 1]} |Z^{m_n}(t)|^{2+\delta}$. Thus, using the Burkholder-Davis-Gundy inequality (e.g., cohen2015stochastic), \begin{equation*} \mathbb{E}_{m_n}\left[\sup_{t\in [0, 1]} |Z^{m_n}(t)|^{2+\delta}\right] \le C\, \mathbb{E}[W^{2+\delta}N(1)^{1+\delta/2}]\le C\,\mathbb{E}[W^{2+\delta}], \end{equation*} which is finite by Condition (ref), and hence $\dot{\ell}_{m_n}^2$ is uniformly $P_{m_n}$-integrable. For $(\gamma_{m_n}^{{\rm t}}\dot{v}_{m_n})^2$ we first note that by (ref) and the fact that $\gamma_{0}^{\rm t} \beta_{m_n}\to \eta_0$ uniformly (as noted above), it follows straightforwardly that \begin{equation*} \liminf_{n\to\infty}\inf_{s\in [0, 1]} \gamma_{0}^{\rm t} \beta_{m_n}(s) \ge c > 0, \end{equation*} for some $c$. Combining this with the uniform convergence of $\gamma_{m_n}^{\rm t}\beta_{m_n}$ to $b$, it follows that for all large enough $n$, the ratio $\gamma_{m_n}^{\rm t}\beta_{m_n} /\gamma_{0}^{\rm t}\beta_{m_n}$ is bounded. Hence by Proposition II.4.1. in andersen1993statistical, $K^{m_n}(t)\coloneqq \int_0^t (\gamma_{m_n}^{\rm t}\beta_{m_n})/(\gamma_{0}^{\rm t}\beta_{m_n})\,\mathrm{d}M^{m_n}$ is a local square integrable martingale with $K^{m_n}(0)=0$ a.s., and \begin{align*} \left[K^{m_n},\, K^{m_n}\right](t) \le \int_0^1 \left(\frac{\gamma_{m_n}^{\rm t}\beta_{m_n}(s)}{\gamma_{0}^{\rm t}\beta_{m_n}(s)}\right)^2\,\mathrm{d}N(s) = \left(\frac{\gamma_{m_n}^{\rm t}\beta_{m_n}(T)}{\gamma_{0}^{\rm t}\beta_{m_n}(T)}\right)^2. \end{align*} For all large enough $n$, we have that $|\gamma_{m_n}^{\rm t}\beta_{m_n}(s)| \le (\overline{b} + 1)$ where $\overline{b}\coloneqq \sup_{t\in [0, 1]} |b(t)|$ and also $\inf_{s\in[0, 1]}\gamma_{0}^{\rm t}\beta_{m_n}(s) \ge c > 0$. Thus for such $n$ the right hand side above can be further bounded above by \begin{equation*} \left(\frac{\gamma_{m_n}^{\rm t}\beta_{m_n}(T)}{\gamma_{0}^{\rm t}\beta_{m_n}(T)}\right)^2 \le \frac{\overline{b}+ 1}{c}. \end{equation*} Then by the Burkholder-Davis-Gundy inequality for a $\delta>0$, and all large enough $n$, \begin{equation*} \mathbb{E}[(\gamma_{m_n}^{\rm t} \dot{v}_{m_n})^{2+\delta}]\le \mathbb{E}\left[\sup_{t\in[0, 1]}K^{m_n}(t)^{2+\delta}\right]\le C \left(\frac{\overline{b} + 1}{c}\right)^{1+\delta/2}, \end{equation*} which is finite, hence $(\gamma_{m_n}^{{\rm t}}\dot{v}_{m_n})^2$ un uniformly $P_{m_n}$-integrable.