EconBase
← Back to paper

Inference on Time Series Nonparametric Conditional Moment Restrictions Using General Sieves

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

78,285 characters · 22 sections · 29 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference on Time Series Nonparametric Conditional Moment Restrictions Using General Sieves

abstractGeneral nonlinear sieve learnings are classes of nonlinear sieves that can approximate nonlinear functions of high dimensional variables much more flexibly than various linear sieves (or series). This paper considers general nonlinear sieve quasi-likelihood ratio (GN-QLR) based inference on expectation functionals of time series data, where the functionals of interest are based on some nonparametric function that satisfy conditional moment restrictions and are learned using multilayer neural networks. While the asymptotic normality of the estimated functionals depends on some unknown Riesz representer of the functional space, we show that the optimally weighted GN-QLR statistic is asymptotically Chi-square distributed, regardless whether the expectation functional is regular (root-$n$ estimable) or not. This holds when the data are weakly dependent beta-mixing condition. We apply our method to the off-policy evaluation in reinforcement learning, by formulating the Bellman equation into the conditional moment restriction framework, so that we can make inference about the state-specific value functional using the proposed GN-QLR method with time series data. In addition, estimating the averaged partial means and averaged partial derivatives of nonparametric instrumental variables and quantile IV models are also presented as leading examples. Finally, a Monte Carlo study shows the finite sample performance of the procedure

\onehalfspacing

Introduction

Consider a conditional moment restriction model

equation[equation omitted — 95 chars of source]

where $\rho$ is a scalar residual function; $\alpha_0=(\theta_0,h_0)$ contains a finite dimensional parameter $\theta_0$ and an infinite dimensional parameter $h_0$, which may depend on some endogenous variables $W_t$. The conditioning filtration $\sigma_t(\mathcal{X})$ is the sigma-algebra generated by variables $\{\mathcal{X}_s: s\leq t\}$, where $ \mathcal{X}_s$ is a vector of multivariate (finite dimensional) exogenous variables, including all relevant lagged variables of $Y_t$ and other instrumental variables. The model therefore allows for endogenous variables and weakly dependent data.

This paper considers optimal estimation and inference for linear functionals $\phi(\alpha_0)$ of the infinite dimension. The functional may be either known or not. When it is unknown, it is assumed to take the form $$ \phi(\alpha_0)=\mathbb El(h_0(W_t))\,, $$ where $l$ is a known linear function and $h_0(W_t)$ is the nonparametric function on endogenous variable. We use general nonlinear sieve learning spaces, whose complexity grows with the sample size, to estimate the infinite dimensional parameter, such as multi-layer neural networks and Gaussian radial basis. The motivation of using general nonlinear sieve learning space, besides being adaptive to high dimensional covariates, is that they allow unbounded supports of the covariates. This is particularly desirable for models of dependent time series data, such as nonlinear autoregressive models.

We formally establish inferential theories of these functionals learned using the general nonlinear sieve learning space, and conduct inference using quasi-likelihood ratio (QLR) statistics based on the optimally weighted minimum distance. Of particular interest is the estimation of an expectation functional, such as averaged partial means, weighted average derivatives and averaged squared partial derivatives, of a nonparametric conditional moment restriction via nonlinear sieve learning sieves. An important insight from our main theory is that the asymptotic distribution does not depend on the actual choice of the learning space, but is only determined by the functional and the loss function. Therefore, estimators produced by either deep neural networks, Gaussian radial basis, or other nonlinear sieve learning basis, have the same asymptotic distribution.

In general, machine learning inference often relies on sample splitting/ cross-fitting, which does not work well in the time series setting. We propose a new time series efficient inference based on the optimal quasi-likelihood ratio test, without requiring cross-fitting. It is shown that the optimally weighted QLR statistic, based on the general nonlinear sieve learning of $h_{0}()$, is asymptotically chi-square distributed regardless of whether the information bound for the expectation functional is singular or not, which can be used to construct confidence sets without the need to compute standard errors. We present a Monte Carlo study to illustrate finite sample performance of our inference procedure.

Depending on the specific applications, our model may involve Fredholm integral equation of either the first kind (NPIV and NPQIV) or the second kind (Bellman equations). In the former case, it is well known that estimating $h_0$ is an ill-posed problem and the rate of convergence might be slow. In the latter case, the problem can be well-posed. As one of the leading examples of the Fredholm integral equation of second kind, we show that our framework implies a natural neural network-based inference in the context of Reinforcement Learning (RL), a popular learning device behind many successful applications of artificial intelligence such as AlphaGo, video games, robotics, and autonomous driving sutton2018reinforcement, silver2016mastering, vinyals2019alphastar, shalev2016safe. Due to the dynamics of the RL model, theoretical analysis of reinforcement learning naturally requires to explicitly allow time series dependency among the observed data. Earlier theoretical studies focused on the settings where the value function are approximated by linear functions. More recent developments on nonlinear learning space include farahmand2016regularized, geist2019theory,fan2020theoretical, duan2021optimal, long20212, chen2022well,shi2020statistical, among others. Our innovation lies in making inference about the functionals (such as the value functional for specific states) of the $Q$-function using general nonlinear sieve learning spaces. While the reinforcement learning is based on the well known Bellman equation, it can be formulated as the conditional moment restriction model with time series data. Therefore, one can apply the GN-QLR inference to estimating the state-specific value function in the setting of the off-policy evaluation. These applications are potentially useful for dynamic causal inference.

In the i.i.d. case, existing theoretical works on neural networks have focused on deriving approximation theories and optimal rates of convergence for estimations. Theoretically, deep learning has been shown to be able to approximate a broad class of highly nonlinear functions, see, e.g. mhaskar2016learning, rolnick2017power, lin2017does, shen2021neural,hsu2021approximation, schmidt2020nonparametric. yang1999information obtained the minimax $L_2$- rate of convergence for neural network models. Recently, chen2021efficient considered NN efficient estimation of the (weighted) average derivatives in a NPIV model for i.i.d. data, and presented consistent variance estimation. In contrast, using a general theory of Riesz representations, we derive the asymptotic distribution of the finite dimensional parameter $\theta_0$ and functionals of the infinite dimensional parameter $h_0$ that is learned from the general learning space. The uncertainty of the general nonlinear sieve learning estimator plays an essential role in the asymptotic distributions. chernozhukov2018double,chernozhukov2018generic, chernozhukov2018automatic proposed double machine learning and debias methods to achieve valid inference; dikkala2020minimax studied a minimax criterion function to study the unknown functional approximated by neural networks for NPIV models. In addition, the Riesz representation is playing a central role in our inferential theory. See newey1994asymptotic, shen1997methods,chen1998sieve,chernozhukov2020adversarial for related approaches.

In the time series setting, the neural networks have been applied to economic demand estimations as in chen2009land, and is widely applicable in financial asset pricing such as guijarro2021deep,gu2020empirical,bali2021different. These papers approximate unknown functions by neural networks, but without rigorous theoretical justifications. All these models can be formulated as an inference problem for conditional moments.

The rest of the paper is organized as follows. Section (ref) first introduces the model, the general nonlinear sieve space, the estimation and inference procedures. Section (ref) establishes the convergence rate of the nonlinear sieve estimator for the unknown function satisfying the conditional moment restrictions with weakly dependent data. Section (ref) provides the limiting distribution of the estimator for functionals that can be regular or irregular. Section (ref) shows that the GN-QLR statistics is asymptotically Chi-square distributed for both the regular and irregular functionals for time series. In Section (ref) we apply our approach to the estimation of the value function of RL and the weighted average derivative of NPIV and NPQIV as leading examples. Section (ref) contains simulation studies and Section (ref) briefly concludes.

The model

The General sieve learning space

This paper studies inference with the general nonlinear sieve learning space. The unknown function is estimated on a learning space, denoted by $\mathcal H_n$, is a general approximation space that consists of either linear or nonlinear sieves, provided that the function of interest can be approximated well by the learning space.

The popular feedforward neural network (NN) is one of the leading examples that fits into this context. Many theoretical studies have shown that NN can well approximate a broad class of functions and achieves nice statistical properties. The multilayer feedforward NN composites functions taking the form: $$ h(x)= \theta_{J+1}h_{J}(x), \quad \cdots \quad h_{j}(x)= \sigma(\theta_{j} h_{j-1}(x)), \quad \cdots \quad,h_{0}(x)=x $$ where the parameters $\theta=(\theta_1,\cdots,\theta_J)$ with $\theta_j\in\mathbb R^{d_j\times d_{j-1}}$ , $h_j(x) \in \mathbb R^{d_j}$, and $\sigma: \mathbb R^{d_j}\rightarrow \mathbb R^{d_{j}}$ is a elementwise nonlinear activation function, usually the same across components and layers. One of the popularly used activation functions is known as ReLU, defined as $ \sigma(x)= \max(0, x). $ The number of neurons being used in layer $j$, denoted by $d_j$, is called the width of that layer.

We could also use other nonlinear approximation learning spaces, which uses nonlinear combinations of inputs and neurons. One such example is the space spanned by Gaussian radial bases, which is a multilayer compositions of functions of the form: $$ h(x)=\alpha_0+\sum_{j=1}^J\alpha_jG(\sigma_j^{-1}\|x-\gamma_j\|),\quad \alpha_0,\alpha_j, \gamma_j\in\mathbb R, \sigma_j>0, $$ where $G$ is the standard normal density function. A key feature is that here inputs and neurons (e.g., a vector of $x$) are “nonlinearly combined" as $\|x-\gamma_j\|$, while they are linearly combined as indices $\theta_jx$ in the ordinary neural networks. Additional examples of nonlinear sieves include spline and wavelet sieves. They are very flexible and enjoy better approximation properties than linear sieves.

One of the key motivations of using general nonlinear sieve learning space, besides being adaptive to high dimensional covariates, is that it allows unbounded supports of input covariates. This is particularly desirable for time series models dependent data, such as nonlinear autoregressive models.

Semiparametric learning

We shall assume a finite-order Markov property: for some known and fixed integer $r\geq 1$, let $X_t:=(\mathcal{X}_t,...,\mathcal{X}_{t-r})$ for all $ t=1,...,n$. define

eqnarray*[eqnarray* omitted — 175 chars of source]

where we assume that $\mathbb{E}[\rho(Y_{t+1},\alpha)|\sigma_t(\mathcal{X})]$ and $\operatorname{Var}(\rho(Y_{t+1},\alpha_0)|\sigma_t(\mathcal{X}))$ only depend on $( \mathcal{X}_t,...,\mathcal{X}_{t-r})$ for all $\alpha$. The model is then equivalent to $Q(\alpha_0)=0$ where

equation*[equation* omitted — 75 chars of source]

Here we use the optimal weighting function $\Sigma(X_t)$. Suppose there are nonparametric estimators $\widehat m(X,\alpha)$ and $\widehat\Sigma(X_t)$ for $m(X_t.,\alpha)$ and $\Sigma(X_t)$, we then define the sample criterion function

equation*[equation* omitted — 106 chars of source]

The estimated optimal weighting matrix is needed for the quasi-likelihood inference. In practice, one can start with the identity weighting function to obtain an initial estimator for $\alpha_0$, use it to estimate $\Sigma(X_t)$ , then update the estimator using the estimated optimal weighting matrix.

We focus on the general nonlinear sieve learning approximation to the true nonparametric function, and restrict to the following estimation space:

equation*[equation* omitted — 61 chars of source]

Here $\Theta$ is a compact set as the parameter space for $\theta_0$ but not necessarily for $\mathcal{H}_n$. In addition, let $P_{en}(h)$ denote some functional penalty for the infinite dimensional parameter. We then define the estimator $\widehat\alpha=(\widehat\theta,\widehat h)\in \mathcal{ A}_n$ as an approximate minimizer of the penalized loss function restricted to the general nonlinear sieve learning space:

equation*[equation* omitted — 152 chars of source]

The tuning parameter $\lambda_n $ is chosen to decay relatively fast, so that the penalization $P_{en}(\cdot)$ does not have a first-order impact on the asymptotic theory. Nevertheless, the functional penalization is imposed to overcome undesirable properties associated with estimates based on a large parameter space. Essentially, it plays a role of forcing the optimization to be carried out within a weakly compact set shen1997methods.

The functions $(x,\alpha )\mapsto \widehat{m}(x,\alpha )$ and $x\mapsto \widehat{\Sigma }(x)$ are nonparametric estimators of $(x,\alpha )\mapsto m(x,\alpha )$ and $x\mapsto \Sigma (x)$ (a positive definite weighting matrix) respectively. The projection $m(X_t,\alpha)$ can be also estimated using linear sieves:

equation*[equation* omitted — 116 chars of source]

where we consider linear sieve space: let $\{\Psi_j: j=1,\dots,k_n\}$ denote a set of sieve bases,

equation*[equation* omitted — 138 chars of source]

So we use the general nonlinear sieve learning space $\mathcal{H}_n$ to approximate the function space for $h_0$, and a linear sieve space $\mathcal{D}_n$ to approximate the instrumental space, which is easier to implement computationally than using nonlinear sieve approximations to the instrumental space. A more important motivation of using linear sieve space to estimate the conditional mean function $\mathbb{E}[\rho(Y_{t+1},\alpha)|\sigma_t(\mathcal{X})]$ is that the sample loss function $Q_n(\alpha)$ can be shown to have a local quadratic approximation (LQA): for some $B_n=O_P(1)$ and $Z_n\to^dN(0,1)$,

equation[equation omitted — 121 chars of source]

uniformly for all $\alpha$ in a shrinking neighborhood of $\alpha_0$ and $ |x|\leq Cn^{-1/2}$; here $\langle u_n, \alpha-\alpha_0\rangle$ is some inner product between $\alpha-\alpha_0$ and some function $u_n$, to be defined explicitly later. This LQA plays a fundamental role for the inferential theory of semiparametric inference using general nonlinear sieve learning methods.

Semiparametric efficient estimations

Let the parameter space of the true function be $\mathcal H_0$ and let $\mathcal A_0=\Theta\times \mathcal H_0$. We are interested in the inference of $\phi(\alpha_0)$, where $\phi:\mathcal{ A}_0\to \mathbb{R}$ can be a known functional of $\alpha_0$. We also study the inference problem of unknown functionals, taking the form

equation*[equation* omitted — 60 chars of source]

where $l(\cdot)$ is a known function. While the naive plug-in estimator $ \frac{1}{n}\sum_{t=1}^nl( \widehat h(W_t)) $ is also asymptotically normal, when the model contains endogenous variables, it is not semiparametrically efficient. An important example of $\phi(\alpha_0)$ is the weighted average derivative of nonparametric instrumental variable regression (NPIV), defined as

equation*[equation* omitted — 86 chars of source]

where $\Omega(\cdot)$ is a known positive weight function and $\nabla h_0$ denotes the gradient of the nonparametric regression function $h_0$. As documented by ai2012semiparametric, the simple plug-in estimator is not an efficient estimator. To obtain a more efficient estimator, on the population level consider conditional (given $X_t$) projection of $l(h_0(W_t))$ onto $\rho(Y_{t+1}, \alpha_0)$, and the corresponding functional of interest also can be represented as $\phi(\alpha_0)$ with the functional:

eqnarray[eqnarray omitted — 122 chars of source]

where $\Gamma_0(X_t) = \mathbb{E }[l(h_0(W_t)) \rho(Y_{t+1}, \alpha_0) |\sigma_t(\mathcal{X}) ] \Sigma(X_t)^{-1} $ is the projection coefficient. We shall obtain efficient estimator of $\phi(\alpha_0)$ based on this expectation expression. It is worthy to know that the added term $\Gamma_0(X_t)\rho(Y_{t+1,\alpha_0})$ is in effect only for endogenous regressors. In pure exogeneous models where $W_t=X_t$, we have $\Gamma_0(X_t)=0$. In this case the moment condition ((ref)) reduces to the original one $\phi(\alpha)= \mathbb El(h(W_t))$.

Let

equation[equation omitted — 117 chars of source]

for some estimator $\widehat\Gamma_t$ to be defined later. Then we estimate the functional by $\widehat\phi(\widehat\alpha)$. Asymptotically, we shall show that

equation[equation omitted — 188 chars of source]

where $\mathcal{W}_t= l(h_0(W_t)) - \Gamma_0(X_t) \rho(Y_{t+1}, \alpha_0)$ and $\sigma^2$ is the asymptotic variance. It is clear that the asymptotic distribution arises from two sources of uncertainties, and importantly, the nonparametric learning error $\phi(\widehat \alpha)- \phi(\alpha_0) $ plays a first-order role.

We shall show that in both known and unknown functional case, estimated $ \phi(\alpha_0)$ is asymptotically normal. We then provide quasi-likelihood inference to construct confidence intervals for $\phi(\alpha_0)$.

Rates of Convergence

Weighted function space and sieve learning space

Since the supports of the endogenous variable $W_t$ could be unbounded, we use a weighted sup-norm metric defined as

equation[equation omitted — 120 chars of source]

This is known as “admissible weight" which is often used for $h_0(W_t)$ when $W_t$ has fat tailed distribution (Remark 2.6 of haroske2020nuclear). Smooth functions with unbounded support might still be well approximated under the weighted sup-norm. The $L^2(W)$-norm can be bounded by the weighted sup-norm as: for any function $h(w)$: $$ \|h\|_{L^2(W)}^2=\int h(s)^2 f_W(s)ds\leq \|h\|^2_{\infty,\omega}\int (1+|s|^2)^{\omega} f_W(s)ds, $$ provided the distribution of the endogenous variable $W$ has as density $f_W$ such that $ f_W(s) (1+|s|^2)^{\omega}$ is integrable.

We do not consider the overparametrized regime, but impose restrictions on the complexity of the general nonlinear sieve learning space $\mathcal H_n$, measured by the “number of parameters" of the space, denoted by $p(\mathcal{H}_n)$. More specifically, we impose the following condition.

assum[function and learning space] (i) The function space: The unknown function $h_0\in\mathcal H_0$, which is a weighted H\"{o}lder ball: for some $\gamma>0,g\geq0$, $$ \mathcal H_0=\{h: \|h(\cdot)(1+|\cdot|^2)^{-g/2}\|_{\Lambda^\gamma}\leq c \} $$ where $$ \|f\|_{\Lambda^\gamma}=\sup_w|f(w)|+\max_{|a|=d}\sup_{w_1\neq w_2}\frac{|\nabla^af(w_1)-\nabla^af(w_2)|}{\|w_1-w_2\|^{\gamma-d}}. $$ Also, we require $g<\omega$ for $\omega$ defined in ((ref)). (ii) Approximation rate under the $\|\|_{\infty,\omega}$ norm: $$ \inf_{h\in\mathcal H_0} \| h_0-h\|_{\infty,\omega}\leq cp(\mathcal H_n)^{-m} $$ for some $m>0$, and some sequence $p(\mathcal H_n)\to\infty$, $p(\mathcal{H}_n)\log n=o(n)$. (iii) Complexity: Let $\mathcal N(\delta,\mathcal H_n, \|.\|_{\infty,\omega})$ denote the minimal covering number, that is, the minimal number of closed balls of radius $\delta$ with respect to $\|.\|_{\infty,\omega}$ needed to cover $\mathcal H_n$. We assume, there is a constant $C>0$, so that for any $\delta>0$, $$ \mathcal N(\delta,\mathcal H_n, \|.\|_{\infty,\omega})\leq \left(\frac{Cn}{\delta}\right)^{p(\mathcal H_n)}. $$

We need to assume that $h(w)$ is smooth in some sense with respect to $h(w)$. Condition (i) is a standard weighted smoothness condition for functions with unbounded support. Here two weighted norms are being defined, the weighted sup norm $\|.\|_{\infty,\omega}$ with a weight parameter $\omega$ in ((ref)). The weighted sup norm intead of the usual sup norm is being considered, as discussed above, for the purpose of allowing the nonparametric function $h(\cdot)$ to have possibly unbounded support, which is the typical case for autoregressive models. The other norm is $\|.\|_{\Lambda^\gamma}$ for the H\"{o}lder ball with a weight parameter $g$. Here we require $g<\omega$ so that the closure of the function space $\mathcal H_0$ with respec to the norm $\|.\|_{\infty,\omega}$ is compact, following from gallant1987semi.

In Condition (ii), $p(\mathcal H_n)\to\infty$ measures the dimension of of the learning space. For multilayer neural networks with ReLU activation functions, anthony2009neural showed that the bound holds with $p(\mathcal H_n)$ being the pseudo-dimension of the space and is bounded by $ C J^2 K^2\log(JK^2) $, where $J$ and $K$ respectively denote the width and depth of the network. For finite-dimensional linear sieve, the inequality also holds with $p(\mathcal H_n)$ being bounded by the number of sieve bases.

When the function $h$ has bounded support, Condition (ii) has been verified for numerous learning spaces. For instance, for feed forward multilayer neural networks, bauer2019deep showed that the approximation rate is $n^{-c},$ for $ c= \frac{p}{2p+d^*}$ and $ p=a+\gamma$, with properly chosen depth and width of layers. Importantly, $d^*\leq \dim(W_t)$ is the “intrinsic dimension" of the true function. For instance if $h_0$ has a hierarchical interaction structure or multi-index structure, $d^*$ is the number of index. When the function $h$ has unbounded support, it is known that for linear sieves such as B-splines and wavelets the approximation rate is $m=p(\mathcal H_n)^{-\gamma/\dim(W_t)}$ where $p(\mathcal H_n)$ is the number of basis. The approximation rate is however still an open question for feed forward neural networks in this case.

Ill-posedness

In this section we present the rate of convergence. For simplicity throughout the rest of the paper, we focus on the case $ \dim(\rho(Y_{t+1},\alpha))=1$. By the identification condition, $Q(\alpha)=0$ if and only if $\alpha=\alpha_0.$ So the usual risk consistency refers to $Q(\widehat\alpha)=o_P(1)$. In the presence of endogenous variables, the risk consistency however, is not sufficient to guarantee the estimation consistency. The latter is often defined under a strong norm:

equation*[equation* omitted — 109 chars of source]

We first introduce a pseudometric on $\mathcal{A}_n$ that is weaker than $ \|.\|_{\infty,\omega}$. To do so, recall the general Gateaux derivative. Given generic $ \alpha=(\theta, h) $ and $v=(v_\theta, v_h)$, let $F(x,\alpha)=F(x,\theta, h) $ be a function that is assumed to be differentiable with respect to $ \theta$. Define

eqnarray*[eqnarray* omitted — 184 chars of source]

where we implicitly assume $\frac{dF(x, \theta, h+\tau v_h)}{d\tau}$ exists at $\tau=0.$ Then the weak norm is defined to be

equation*[equation* omitted — 112 chars of source]

Define $\pi_n\alpha_0\in\mathcal{A}_n$ be such that

equation*[equation* omitted — 130 chars of source]

The following assumption imposes conditions on the local curvature of the criterion function.

assum[criterion function] There are $c_1, c_2>0$ so that (i) $\|\alpha-\alpha_0\|^2\leq c_1\mathbb{E }m(X_t,\alpha) ^2 \Sigma(X_t)^{-1} $ for all $\alpha\in\mathcal A_n$. (ii) $\mathbb{E }m(X_t,\pi_n\alpha_0)^2 \Sigma(X_t)^{-1} \leq c_2 \|\alpha_0-\pi_n\alpha_0\|^2$.

We now discuss the ill-posedness which reflects the relation between the risk consistency and estimation consistency. Let the sieve modulus of continuity be

equation*[equation* omitted — 147 chars of source]

We say that the problem is ill-posed if $\delta=o(\omega_n(\delta))$ as $\delta\to0.$ The growth of $\omega_n(\delta)\delta^{-1}$ reflects the difficulty of recovering $\alpha_0$ through minimizing the criterion function.

Rates of convergence

Below we present regularity conditions to achieve the rates of convergence. We allow weakly dependent time series data satisfying $\beta$-mixing conditions. Define the mixing coefficient

equation*[equation* omitted — 125 chars of source]

where $\mathcal{F}_s^t$ denotes the $\sigma$-field generated by $(Y_{s+1}, X_s),...,(Y_{t+1}, X_t)$.

assum[Dependences] (i) $\{(Y_{t+1}, X_t)\}_{t=1}^n$ is a strictly stationary and $\beta$-mixing sequence with $\beta(j)\leq \beta_0\exp(-cj)$ for some $ \beta_0, c>0$. (ii) There is a known and finite integer $r\geq 1$ so that for each $ \alpha\in\mathcal{A}_n$ and $t=1,...,n$, The conditional expectation $ \mathbb{E}[ f(S_t, \alpha)|\sigma_t(\mathcal{X})]$ depend on $\sigma_t( \mathcal{X})$ only through $X_t:=(\mathcal{X}_t,...,\mathcal{X}_{t-r})$, for $S_t=(Y_{t+1},W_t)$ and \begin{equation*} f(S_t,\alpha)\in \{ \rho(Y_{t+1},\alpha),\quad \rho(Y_{t+1},\alpha_0)^2 , \quad l(h_0(W_t)) \rho(Y_{t+1}, \alpha_0) \}. \end{equation*}
assum$Q(\alpha)=0$ if and only if $\alpha=\alpha_0.$ In addition, $Q(\alpha)$ is lower semicontinuous.

The lower semicontinuity of the criteria function is satisfied by the risk function of many interesting models. This condition ensures that it has a minimum on any compact set.

assum[Penalty] (i) There is $M_0>0$, $P_{en}(h) \leq M_0$ for all $h\in \mathcal{H}_n\cup\{h_0\}$.\newline (ii) $P_{en}$ is lower semicompact on $(\mathcal{A}_n, \|.\|_{\infty,\omega})$, i.e. $\{h: P_{en}(h)\leq M\}$ is compact for any $M>0$. \newline (iii) $\frac{k_n}{n} +Q(\pi_n\alpha_0)=O(\lambda_n) $ where recall $k_n$ is the number of linear sieve bases in $\mathcal{D}_n$.

Define $$\epsilon(S_t,\alpha):= \rho(Y_{t+1},\alpha)-m(X_t,\alpha).$$ One of the major technical steps is to establish the stochastic equicontinuity for the function class $\Psi_j(X_t)\epsilon(S_t,\alpha)$ for $ \beta$-mixing observations, where $\alpha$ belongs to the class of deep neural networks. More specifically, we shall derive the bound for, with $\Psi(X_t):=(\Psi_j(X_t): j\leq k_n)$:

equation*[equation* omitted — 188 chars of source]

for a given convergence sequence $r_n\to 0$. This is achieved under the following Assumption.

assumThere is $C>0$, (i) There are $\kappa>0$ and $C>0$ so that for all $\delta>0$ and all $ \alpha_1, \alpha_2\in\mathcal{A}_n$, \begin{equation*} \max_{j\leq k_n} \mathbb{E}[\Psi_j(X_t)^2+1]\sup_{ \|\alpha_1-\alpha_2 \|_{\infty,\omega}< \delta} | \epsilon( S_t, \alpha_1)- \epsilon( S_t, \alpha_2 )| ^2\leq C\delta^{2\kappa}. \end{equation*} (ii) $\mathbb{E}\max_{j\leq k_n}\Psi_j(X_t)^2 \sup_{\alpha\in\mathcal{A}_n} \rho(Y_{t+1}, \alpha )^2\leq C.$ (iii) There is a $\|.\|_{\infty,\omega}$- neighborhood of $\alpha_0$ on which $m(\cdot, \alpha)$ is continuously pathwise differentiable with respect to $\alpha$, and there is a constant $C>0$ such that $\|\alpha-\alpha_0\|\leq C\|\alpha-\alpha_0\|_{\infty,\omega}.$

Next we present regularity conditions on the linear sieve space $\mathcal{D} _n$ used to approximate the conditional mean function $m(X,\alpha)$.

assum[Linear sieve space] (i) There is $\varphi_n\to0$ so that uniformly for $\alpha\in\mathcal{A}_n$, there is $k_n\times 1$ vector $b_{\alpha}$, \begin{equation*} \mathbb{E }[g(X_t, \alpha)- \Psi(X_t)^{\prime }b_\alpha ]^2=O(\varphi_n^2), \end{equation*} for all $g(X_t,\alpha)\in \{m(X_t,\alpha),\mathbb{E }[l(h_0(W_t)) \rho(Y_{t+1}, \alpha_0)|X_t], \frac{dm(X_t,\alpha)}{d\alpha}[u_n], \frac{ dm(X_t,\alpha)}{d\alpha}[u_n]\Sigma(X_t)^{-1}\}$. (ii) Let $\Psi_n $ be the $n\times k_n$ matrix of the linear sieve bases: $\Psi_n=(\Psi(X_t): t=1...n) _{n\times k_n}$: and let $A:=\frac{1}{n}\mathbb{E }\Psi_n^{\prime }\Psi_n$. The linear sieve satisfies: $\lambda_{\min}(A)>c$ and $\| \frac{1}{n}\Psi_n^{\prime }\Psi_n -A\|=o_P(1)$.

Finally, we apply the pseudo dimension to quantify the complexity of the neural network class.

assum(i) $\sup_x[ \Sigma(x)^{-1}+ \Sigma(x)]<C .$ Also, $\sup_x|\widehat \Sigma(x)-\Sigma(x)|=o_P(1)$. (ii) The distribution of the endogenous variable $W_t$ has a density function $f_W$, which satisfies $\int w(x)^{-2} f_W(x)dx<\infty.$

Recall that $k_n$ denotes the number of sieve bases being used to estimate the expectation function $m(X,\alpha)$; $\varphi_n$ is the approximation rate in Assumption (ref). Let

eqnarray[eqnarray omitted — 270 chars of source]
thm[Rate of convergence] Under Assumptions (ref)-(ref), for any $\epsilon>0$ , \begin{equation*} \|\widehat\alpha-\alpha_0\|_{\infty,\omega}= O_P(\delta_n),\quad Q(\widehat\alpha)=O_P(\bar\delta_n^2). \end{equation*}

The derived rate of convergence is comparable with that of chen2012estimation. In $\bar \delta_n$, the term $\|\pi_n\alpha_0-\alpha_0\|$ is the approximation error on the general nonlinear sieve learning space; $\sqrt{\lambda_n}$ is the effect of penalization. In addition, $\varphi_n$ and $\sqrt{k_n}d_n$ respectively arise from the bias and variance of estimating $m(X,\alpha)$. In particular, the variance term $\sqrt{k_n}d_n$ depends on the complexity of the general nonlinear sieve learning space, which arises from the stochastic equicontinuity. In addition, $\omega_n(\bar\delta_n)$ connects the convergence under the weak norm $O_P(\bar\delta_n)$ to the convergence under the strong norm via the sieve modulus of continuity. When there are no endogeneity, $ \bar\delta_n $ and $\omega_n(\bar\delta_n)$ are of the same order. General nonlinear sieve spaces with more complicated structures (with larger “dimension" $p(\mathcal H_n)$) have increased covering numbers on the learning space, and thus lead to slower decays of these two terms.

Asymptotic Distributions for Functionals

We now study estimating linear functionals of $\alpha_0$. We establish the asymptotically normality of the estimated functionals formed via pluging-in the general learning estimators.

Riesz representation

A key ingredient of our analysis, as in chen2015sieve, relies on representing the estimation error $\phi(\widehat\alpha)-\phi(\alpha_0)$ using a linear inner product induced from the loss function via the Riesz representation theorem. We define an inner product space as follows.

For any space $\mathcal{H}$, let span$\{\mathcal{H}\}$ denote the closed linear span of $\mathcal{H}$. For any $v_1, v_2$ in span$(\mathcal{A}_n \cup\{\alpha_0\})$, the linear span of $\mathcal{A}_n \cup\{\alpha_0\}$, define the inner product:

equation*[equation* omitted — 180 chars of source]

Let $\alpha_{0,n}\in $ span$(\mathcal{A}_n )$ be such that

equation*[equation* omitted — 112 chars of source]

We note that it is likely $\alpha_{0,n}\neq\pi_n\alpha_0$ because $ \pi_n\alpha_0\in\mathcal{A}_n$, which is not the same as $\text{span}( \mathcal{A}_n )$, when $\mathcal{A}_n$ is a nonlinear sieve space.

Given Theorem (ref), we can focus on shrinking neighborhoods

eqnarray[eqnarray omitted — 373 chars of source]

for a generic constant $C>0$, where $v_n^*$ is the Riesz representer to be defined below.

Because both $\mathcal{A}_{osn} $ and $\alpha_{0,n}$ are functions inside the general nonlinear sieve learning space, $(\bar V_n, \langle.\rangle)$ is a finite dimensional Hilbert space under the weak-norm $\|v\|=\sqrt{\langle v, v\rangle}$. Suppose $\frac{d\phi(\alpha_0)}{d\alpha}[v]$ is a linear functional. As any linear functional on a finite dimensional Hilbert space is bounded, by the Riesz representation Theorem, there is $v_n^*\in \bar V_n$ so that

equation*[equation* omitted — 107 chars of source]

To appreciate the role of Riesz representation in the semiparametric inference, note that $\widehat\alpha- \alpha_{0,n}\in \bar V_n$, and we have,

eqnarray*[eqnarray* omitted — 406 chars of source]

where the first equality follows from the smoothness condition (Assumption (ref) below) of the functional; the second equality is to the linearity of the functional pathwise derivative. In addition, suppose $\frac{ d\phi(\alpha_0)}{d\alpha}[ \alpha_{0,n}-\alpha_0]$ is negligible, a claim we shall discuss in Remark (ref) later, we can then apply the Riesz representation theorem to reach the last line of the expansion.

In addition, one of the key technical steps in the proof, by locally expanding the risk function, is to prove:

equation*[equation* omitted — 195 chars of source]

where $\mathcal{Z}_t=\rho(Y_{t+1}, \alpha_0)\Sigma(X_t)^{-1}\frac{d m(X_t,\alpha_0)}{d\alpha}[v^*_n] ,$ and $\|v_n^*\| ^2=\operatorname{Var}(\frac{1}{\sqrt{n}} \sum_t\mathcal{Z}_t).$ Then together we have

equation*[equation* omitted — 110 chars of source]

Importantly, our inference procedure does not require estimating the Riesz representer $v_n^*$ or $\|v_n^*\|$. Instead, we propose a quasi-likelihood ratio (QLR) inference. We shall provide regularity conditions in the next section to formalize the above derivations, and subsequently address estimating the known and unknown functionals.

Asymptotic distributions for known functionals

We have the following assumptions.

assum[smoothness] (i) The functional $\phi$ is linear in the sense that the functional $\phi$ is linear in the sense that $ \phi(\alpha)-\phi(\alpha_0)= \frac{d\phi(\alpha_0)}{d\alpha} [\alpha-\alpha_0] $. (ii) $\sqrt{n}\frac{d\phi(\alpha_0)}{d\alpha}[ \alpha_{0,n}-\alpha_0] = o_P(\|v_n^*\|).$
remarkAssumption (ref) (iii) requires that the neural network bias term $\frac{d\phi(\alpha_0)}{d\alpha}[ \alpha_{0,n}-\alpha_0]$ should be negligible. Here we present a sufficient condition following the discussion of chen2015sieve. First, since $ \alpha_{0,n}$ is the projection of $\alpha_0$ on to span$(\mathcal{A}_n)$ and $v_n^*\in \bar V_n\subset$ span$(\mathcal{A}_n)$, we have $\langle v_n^*, \alpha_{0,n}-\alpha_0\rangle=0$. In addition, define an infinite dimensional Hilbert space $\bar V$ as the closure of the linear span of $ \mathcal{A}- \{\alpha_0\}. $ Suppose $\frac{d\phi(\alpha_0)}{d\alpha}[ \cdot ]$ is bounded, then there is a unique Riesz representer $v^*\in\bar V$ so that \begin{equation*} \frac{d\phi(\alpha_0)}{d\alpha}[v]=\langle v^*,v\rangle, \quad \forall v\in \bar V. \end{equation*} As $\alpha_{0,n}-\alpha_0\in\bar V $, we have \begin{equation*} \left|\sqrt{n}\frac{d\phi(\alpha_0)}{d\alpha}[ \alpha_{0,n}-\alpha_0] \right| = \left|\sqrt{n} \langle v^*-v_n^*, \alpha_{0,n}-\alpha_0\rangle \right| \leq \sqrt{n} \| v^*-v_n^*\| \| \alpha_{0,n}-\alpha_0\|. \end{equation*} So condition (iii) holds as long as $\sqrt{n} \| v^*-v_n^*\| \| \alpha_{0,n}-\alpha_0\|=o_P(\|v_n^*\|)$.

To allow quantile applications that involve nonsmooth loss functions, we need to show that the sample criterion function $Q_n(\alpha)$ can be replaced with a smoothed criterion $\widetilde Q_n(\alpha):= \frac{1}{n}\sum_t\ell(X_t, \alpha)^2\widehat\Sigma(X_t)^{-1}$, where $\Psi_n=(\Psi(X_t): t=1...n) _{n\times k_n}$:

equation*[equation* omitted — 189 chars of source]

and $m_n(\alpha)$ denotes the $n\times 1$ vector of $m(X_t,\alpha)$. The replacement error is negligible:

equation*[equation* omitted — 140 chars of source]

Therefore, theoretical analysis of $Q_n(\alpha)$ is asymptotically equivalent to that of $\widetilde Q_n(\alpha)$, while the latter is second-order pathwise differentiable, and admits a local quadratic approximation. Formalizing this argument would require the following conditions.

assum$m(x, t)$ is twice differentiable with respect to $t$, and there is $C>0$, so that, recall that $u_n= v_n^*/\|v_n^*\|$ being the “normalized Riesz representer": (i) $\mathbb{E }| \rho(Y_{t+1}, \alpha_0)|^{2+\zeta}\left| \frac{d m(X_t,\alpha_0)}{d\alpha}[u_n] \right |^{2+\zeta}+\mathbb{E }| \rho(Y_{t+1}, \alpha_0)|^{2+\zeta}<C$ for some $\zeta > 0$; (ii) $\mathbb{E }\sup_{\alpha\in\mathcal{C}_{n}} \sup_{|\tau|\le Cn^{-1/2}} \frac{1}{n}\sum_t\left[\frac{d^2}{d\tau^2} m(X_t, \alpha+\tau u_n)| \right] ^2 <C$; (iii) $\sup_{\tau\in(0,1)}\sup_{ \alpha\in\mathcal{C}_n} \mathbb{E }\left[ \frac{d ^2}{d\tau^2} m(X_t,\alpha_0+\tau(\alpha-\alpha_0))\right]^2 =o(n^{-1})$; (iv) $k_n \sup_{\alpha\in\mathcal{C}_n}\frac{1}{n}\sum_t[ \frac{ dm(X_t,\alpha )}{d\alpha}[u_n]- \frac{dm(X_t,\alpha_0)}{d\alpha} [u_n]]^2=o_P(1)$; (v) $\mathbb{E} \big\{[\max_{j\leq k_n}\Psi_j(X_t)^2+1]\sup_{\alpha\in\mathcal{C} _n} (\rho(Y_{t+1}, \alpha)-\rho(Y_{t+1}, \alpha_0) )^2\big\} < C\delta_n^{2\eta} $ for some $\kappa, \eta >0$.

Finally, we need to strengthen conditions on the penalty and some rates of convergence as follows.

assum(i) Let $\mathcal{C}_h:=\{h: (\theta, h)\in\mathcal{C}_n \text{ for some $ \theta\in \Theta$}\}$, which is the local neighborhood for the estimated $ h(\cdot)$. We assume \begin{equation*} \lambda_n \sup_{h\in\mathcal{C}_h }|P_{en}( h) - P_{en}(h_0 ) | + \lambda_n \sup_{h\in\mathcal{C}_h }|P_{en}( \pi_nh) - P_{en}(h_0 ) | =o(n^{-1}). \end{equation*} (ii) $\sqrt{n}\bar \delta_n\|\widehat\Sigma_n-\Sigma_n\| =o(1)$, where $\widehat\Sigma_n$ and $\Sigma_n$ be the diagonal matrix of $\widehat\Sigma(X_t)$ and $\Sigma(X_t)$ for all $t$, and furthermore $ \varphi_n^2\bar\delta_n^2 + k_nd_n^2\delta_n^{2\eta} + \sqrt{k_n} d_n\delta_n^{\eta}\bar\delta_n =o(n^{-1})$.

The following condition is similar to Condition C in shen1997methods, which is used to control the approximation error of the learning space for locally perturbed elements.

assumThere is $\mu_n\to 0$ so that $\mu_n\bar \delta_n = o(n^{-1}) $, we have \begin{equation*} \sup_{\alpha\in\mathcal{C}_n} \frac{1}{n}\sum_{t=1}^n [m( X_t, \pi_n \alpha)- m( X_t, \alpha)]^2 =O_P(\mu_n^2). \end{equation*}
thm[Limiting distribution] Under Assumptions (ref)-(ref), \begin{eqnarray*} \sqrt{n} \frac{\phi(\widehat\alpha)-\phi(\alpha_0)}{ \|v_n^*\| } \to^d \mathcal{N}(0,1). \end{eqnarray*}

An important insight from this theorem is that the asymptotic distribution does not depend on the actual choice of the learning space. The asymptotic variance $$\|v_n^*\| ^2=\mathbb{E} \Sigma(X_t)^{-1}\left( \frac{dm(X_t, \alpha_0)}{d\alpha}[v_n^*] \right)^2$$ is only determined by the functional forms $\phi$ and $m(X,\alpha)$, and more generally, the loss function. So whether the multilayer neural network, B-spline, Gaussian radial basis, etc, are being used to estimate $\alpha_0$, the asymptotic distribution is the same. What really matters is the loss function.

Estimation for unknown functionals

We now consider estimating unknown (probably not $\sqrt{n}$-estimable) functionals, taking the form

equation*[equation* omitted — 53 chars of source]

where $l(\cdot)$ is a known function. ai2012semiparametric used the following moment condition ((ref)) to construct the optimal criterion function: \\

equation[equation omitted — 144 chars of source]

where $\Gamma(X_t) = \mathbb{E }[l(h_0(W_t)) \rho(Y_{t+1}, \alpha_0) |\sigma_t(\mathcal{X}) ] \Sigma(X_t)^{-1} . $ They showed that estimating $\gamma_0 $ based on this moment condition leads to more efficient estimator than based on the naive plug-in method $\frac{1}{n}\sum_i l(\widehat h(W_t))$, whenever $W_t$ is endogenous. Because the naive plug-in estimator does not take into account the potential correlations between the moment functions $m(X_t, \alpha)$ and $ l(h(W_t))$.

Using the more efficient moment condition of $\gamma_0$, and letting $$\phi(\alpha):= \mathbb{E}l(h(W_t))-\mathbb{E}\Gamma(X_t)\rho(Y_{t+1},\alpha),$$ we note that $ \phi(\alpha_0)=\gamma_0.$ Suppose the functional $\phi(\cdot)$ were known, and Assumption (ref) continues to hold for $\phi(\alpha)$, then we can show

equation*[equation* omitted — 223 chars of source]

where $v^*_n$ is the Riesz representer. But we in fact are facing a problem of estimating an unknown functional $\phi(\cdot)$. To do so, we first estimate $ \Gamma (X_t)$ by

equation*[equation* omitted — 184 chars of source]

Then define the final estimator:

equation[equation omitted — 197 chars of source]

The following asymptotic expansion holds for the estimated functional:

eqnarray*[eqnarray* omitted — 287 chars of source]

where $\mathcal{Z}_t=\rho(Y_{t+1}, \alpha_0)\Sigma(X_t)^{-1}\frac{d m(X_t,\alpha_0)}{d\alpha}[v^*_n] .$ This explicitly presents two leading sources for the asymptotic distribution, where the asymptotic variance is given by

equation[equation omitted — 220 chars of source]

where $\mathcal W_t$ and $\mathcal Z_t$ are uncorrelated.

We impose the following conditions

assum(i) $\sup_x|\Gamma(x)|^2 +\sup_w\sup_{h\in\mathcal{H} _n}l(h(w))^2<C.$ (ii) $l(h)$ is linear in $h$. (iii) $\mathbb{E}\sup_{\alpha\in\mathcal{C}_n} | l(h(W_t))-l(h_0(W_t))|^2\leq C\delta_n^{2\eta}$, where for simplicity we assume the same $\eta$ as in Assumption (ref) (v).

Assumption (ref) regulates the approximation quality of the instrumental space using linear sieves, which is not stringent since $ \mathbb{E}(l(h(W_t))\rho(Y_{t+1}, \alpha)|\sigma_t(\mathcal{X}))$ is a function of the instrumental variable.

The next assumption imposes a condition on the accuracy of estimating the optimal weighting function $\Sigma(X_t)$. For the NPQIV model this assumption is trivially satisfied since $\widehat\Sigma(X_t)=\Sigma(X_t)= \varpi(1-\varpi)$ is known (see Section (ref) for the definition of $\varpi$). We shall verify it for the NPIV model in Section (ref).

assumThere is a sequence $p_n $ so that $p_n \bar\delta_n\sigma=o(n^{-1})$ and \begin{equation*} \frac{1}{n}\sum_t \Gamma(X_t)\Sigma(X_t) (\widehat\Sigma(X_t)^{-1}-\Sigma(X_t)^{-1})\rho(Y_{t+1}, \alpha_0) =O_P( p_n). \end{equation*}

The asymptotic normality requires some rate restrictions, which we impose below.

assum(i) There is $c_0>0$ so that $\sigma^2>c_0.$ (ii) Let $\nu_n:=\delta_n^{\eta} \sup_x|\widehat\Sigma(x)-\Sigma(x)| +\sqrt{ k_n} d_n\delta_n^{\eta} +\varphi_n^2. $ Then $\nu_n \bar\delta_n\sigma=o(n^{-1})$.
thmSuppose Assumptions (ref)-(ref) hold for $ \phi(\alpha)=\mathbb{E}l(h(W_t))-\mathbb{E}\Gamma(X_t)\rho(Y_{t+1},\alpha)$. In addition, Assumptions (ref)-(ref) hold. Then \begin{equation*} \sqrt{n}\sigma^{-1}( \widehat\gamma-\gamma_0 ) \to^d\mathcal{N}(0,1). \end{equation*}

Quasi-Likelihood Ratio Inference for Functionals

As shown by Theorems (ref) and (ref), computing the asymptotic variance requires estimating Riesz representer. While chen2015sieve and chernozhukov2018automatic proposed framework of estimating the Riesz representer, the task is in general quite challenging when its does not have closed-form approximations. In this section we propose to make inference directly using the optimally weighted quas-likelihood ratio statistic (QLR).

QLR Inference for known functionals

Consider testing

equation*[equation* omitted — 46 chars of source]

for some known $\phi_0\in\mathbb{R}.$ Consider the restricted null space $ \mathcal{A}_n^R:=\{\alpha\in\mathcal{A}_n: \phi(\alpha) =\phi_0\}$. The GN-QLR statistic is defined as

equation*[equation* omitted — 90 chars of source]

where $\widehat\alpha^R\in\mathcal{A}_n^R$ approximately minimizes the penalized loss function over the general nonlinear sieve learning restricted on the null space:

equation*[equation* omitted — 162 chars of source]

Define

eqnarray*[eqnarray* omitted — 111 chars of source]
assum(i) Recall $\mu_n$ as defined in Assumption (ref) . It also satisfies: \begin{equation*} \sup_{\alpha\in\mathcal{A}_{osn}, \phi(\alpha)=\phi_0} \frac{1}{n}\sum_{t=1}^n [m( X_t, \pi_n^R (\alpha+xu_n))- m( X_t, \alpha+xu_n)]^2 =O_P(\mu_n^2) \end{equation*} (ii) $(1+\|v_n^*\|) \sup_{\alpha\in\mathcal{C}_n}|\phi(\pi_n\alpha)-\phi( \alpha)| =o (n^{-1/2})$.

The following theorem shows the asymptotic null distribution of $S_n(\phi_0)$ .

thmSuppose conditions of Theorem (ref) and Assumption (ref) hold. Then under $H_0: \phi(\alpha_0)=\phi_0$, \begin{equation*} S_n(\phi_0)\to^d\chi^2_1. \end{equation*}

QLR inference for unknown functionals

We now move on to the inference for the unknown functional $\gamma_0:= \mathbb{E}l(h_0(W_t))$, which is estimated by $\widehat\gamma$ as defined in ((ref)). Consider testing

equation*[equation* omitted — 52 chars of source]

for some known $\phi_0$. Define

equation*[equation* omitted — 109 chars of source]

where $\widehat\Sigma_2$ consistently estimates the long-run variance (e.g. NW87): $$\Sigma_2:=\operatorname{Var} \left( \frac{1}{\sqrt{n}}\sum_{t=1}^{n} \mathcal W_t\right) =\frac{1}{n}\sum_{t=1}^n\operatorname{Var}(\mathcal W_t) +\frac{1}{n}\sum_{t\neq s} \text{cov}(\mathcal W_t, \mathcal W_s) .$$ We recall that $ \mathcal{W}_t= l(h_0(W_t)) - \Gamma(X_t) \rho(Y_{t+1}, \alpha_0)$.

Note that $(\widehat\alpha,\widehat\gamma)$ is numerically equivalent to the solution to the following problem:

equation*[equation* omitted — 183 chars of source]

We define the GN-QLR statistic as

equation*[equation* omitted — 125 chars of source]

where $\widehat\alpha^R\in\mathcal{A}_n^R$ approximately minimizes the penalized loss function in the learning space $\mathcal H_n$, but fixing $\gamma=\phi_0$:

equation*[equation* omitted — 179 chars of source]

The asymptotic analysis of $\widetilde S_n(\phi_0)$ is rather sophisticated, which requires additional rate constraints stated as follows.

thmSuppose $ \widehat\Sigma_2-\Sigma_2=o_P(1)\Sigma_2$ and conditions of Theorem (ref) hold. Then under $H_0: \gamma_0=\phi_0$ \begin{equation*} \widetilde S_n(\phi_0)\to^d\chi^2_1. \end{equation*}

Examples

In this section, we illustrate our main results using three important models: Reinforcement learning, NPIV and NPQIV. We impose premitive conditions to verify the high level Assumptions (ref), (ref) and (ref) respectively in the two models.

Reinforcement learning

Reinforcement learning (RL) has been an important learning device behind many successes in applications of artificial intelligence. Theories of RL have been developed in the literature of statistical learning and computer science. Most of the existing theoretical works formulate the problem as a least-square regression and approximate the value function by a linear function, such as bradtke1996linear, etc. Nonlinear approximations using kernel methods or deep learning appeared in the more recent literature, for example farahmand2016regularized, geist2019theory, fan2020theoretical, duan2021optimal, long20212,chen2022well. shi2020statistical also conducted inference for the optimal policy using linear sieve representations.

We proceed learning using neural networks, and study the inference for a given policy. We follow the recent literature on the off-policy evaluation problem, and formulate the reinforcement learning problem as a conditional moment restriction model. Assume the observed data trajectory $\{(S_t, A_t, R_t)\}_{t \ge 0}$ is obtained from an unknown behavior policy probability $\pi^b(a|s)$, where $(S_t, A_t, R_t)$ denote the state, action and observed reward at time $t$ respectively and $\pi^b(a|s)$ is the distribution to take action $a$ at state $s$. We denote the space of states and actions as $\mathcal S$ and $\mathcal A$. It is assumed that the reward $R_t$ is jointly determined by $(S_t, A_t, S_{t+1})$. Standing at state $S_t$ at period $t$, one takes action $A_t$ and receives reward $R_t$. The state then transits to $S_{t+1}$ at the next period.

The value of a given policy $\pi$ is measured by the so-called $Q$-function. Specifically, for any given $\pi$ and any state-action pair $(s,a)$, $Q$-function is defined as the expected discounted reward:

equation*[equation* omitted — 101 chars of source]

where $\mathbb E^\pi$ or in short $\mathbb E$ is the expectation when we take actions according to $\pi$, $0\le \gamma < 1$ is the discount factor and we consider the discounted infinite-horizon sum of expected rewards. To estimate $Q^\pi$, a classical approach is to solve the Bellman equation below:

equation*[equation* omitted — 165 chars of source]

The goal is to recover $Q^\pi$ of a given target policy $\pi$. In practice, multiple trajectories $\{(S_{i,t}, A_{i,t}, R_{i,t}, S_{i,t+1})\}_{0\le t \le T,1\le i\le N}$ may be observed to help estimate the $Q$-function. But for simplicity we assume $N=1$ and $T=n$. The more general case can be cast by merging the $N$ time series into a single series of size $n=TN. $

The Bellman equation can be formulated as a conditional moment restriction with respect to $Q^{\pi}$ for weakly dependent time series: $$ \mathbb E[\rho(Y_{t+1}, Q^{\pi})| S_t, A_t]=0,\quad Y_{t+1}= (R_t, S_t, A_t, S_{t+1}), \quad X_t=(S_t, A_t), $$ where $$ \rho(Y_{t+1}, h)=R_t- h(S_t,A_t) + \gamma \int_{x \in \mathcal A} \pi(x|S_{t+1}) h(S_{t+1}, x) \mathrm{d} x. $$

In this framework, the estimation of the function $Q^{\pi}(s,a)$ can be conducted on the neural network space, and we assume that computationally the integration in the $\rho$-function can be well approximately by the Monte Carlo method. For off-policy evaluations, the following value function is of the major interest in this section: given state $s\in\mathcal S$,

equation[equation omitted — 117 chars of source]

which is a known functional $\phi_s(\cdot)$ for a single state $s$.

The Bellman equation also admits a Fredholm integral equation of the second kind kress1989linear, which is a well-posed problem. Therefore, estimating the $Q$-function may achieve fast-rate of convergence. That is, the sieve modulus of continuity satisfies:

equation*[equation* omitted — 145 chars of source]

Recently chen2022well showed this result for $\|.\|_s$ to be either the sup-norm or the $\ell_2$-norm. The inner product is defined, in this case, as $ \langle v_1, v_2\rangle=\mathbb{E}\Sigma(X_t)^{-1}\left( \frac{dm}{d h}[v_1] \right)\left( \frac{dm}{dh} [v_2] \right), $ where

equation[equation omitted — 176 chars of source]

and induced a Riesz representer $v^*$ whose closed form is unavailable. Meanwhile, it follows from the Bellman equation that $ m(X_t, h) = \frac{dm}{dh}[h-Q^{\pi}] $ for all $h\in\mathcal H_n$. Therefore, the weak norm $\|.\|$ can be expressed as: $$ \|h-Q^{\pi}\|^2= \mathbb E m(X_t, h)^2\Sigma(X_t)^{-1}, $$ which shows that the employed minimum distance criterion function is directly estimating the squared weak norm.

Let $\widehat Q^{\pi}$ be the estimated $Q^{\pi}$ using the general nonlinear learning space, and the functional is naturally estimated using $$ \phi_s(\widehat Q^{\pi})= \int_{a \in \mathcal A} \pi(a|s) \widehat Q^{\pi}(s, a) \mathrm{d} a $$ As the moment restriction function $ \mathbb E [ \rho(Y_{t+1}, h)| S_t, A_t] $ is linear in $h$ in this case, it is straightforward to verify the high-level conditions as follows.

assum(i) For some $\zeta>4$, the Riesz representer satisfies $ \mathbb E\int \pi(x|S_{t+1})|v^*_n(S_{t+1}, x)|^{\zeta}dx+ \mathbb E |v_n^*(S_t, A_t)|^{\zeta}\leq \|v_n^*\|^{\zeta}$. (ii) $\mathbb ER_t^4<\infty $, $\mathbb E\max_{j\leq k_n} \Psi_j(X_t)^4<\infty$, $\mathbb E (1+|S_t|^2+|A_t|^2)^{2\omega}<\infty $ and $\mathbb E M(S_{t+1})^4<\infty$, where $M(S_{t+1}):=\int\pi(x|S_{t+1}) (1+x^2+S_{t+1}^2)^{\omega/2}dx$, and $\omega$ is the degree of the weighted-sup metric $\|.\|_{\infty,\omega}$.
propFor the Reinforcement Learning model considered here, Assumption (ref) implies Assumptions (ref), (ref) and (ref).

It then follows from Theorem (ref) that $$\|v_n^*\|^{-1}\sqrt{n}\left(\phi_s(\widehat Q^{\pi}) -\phi_s(Q^{\pi})\right)\to^d\mathcal N(0,1) $$ Inference about $\phi_s(Q^{\pi})$ based on pivotal statistics can be conducted using the GN-QLR test.

The NPIV model

In the nonparametric instrumental variable model (NPIV), consider

equation*[equation* omitted — 95 chars of source]

where $\sigma_t(\mathcal X)$ is the filtration generated from instrumental variables $X_t$. Then $m(X_t, \alpha)= \mathbb{E}[(y_{t+1} - h(W_t))|\sigma_t(\mathcal{X})] $ and the Gateaux derivative is defined as $\frac{dm(X_t, \alpha)}{dh}[v] = \mathbb{E}(v(W_t)|\sigma_t(\mathcal{X})), $ implying

equation*[equation* omitted — 173 chars of source]

We estimate the conditional variance $ \Sigma(X_t)$ by $ \widehat\Sigma_t = \widehat A_n^{\prime }\Psi_n(\Psi_n^{\prime }\Psi_n)^{-1}\Psi(X_t) $ where $\widehat A_n$ is a $n\times 1$ vector of $\rho(Y_{t+1}, \widehat \alpha)^2$. Recall that for $\delta_n$ and $\bar\delta_n$ defined in ((ref)),

equation*[equation* omitted — 107 chars of source]

We impose the following low-level conditions to verify Assumptions (ref) and (ref).

assum(i) $\delta_n^2\bar\delta_n\sigma= o(n^{-1})$, $\mathbb{E }\max_{j\leq k_n}|\Psi_j(X_t) |^2 (U_t^2+1)<C $, and $\mathbb E(U_t^2|\sigma_t(\mathcal X))<C$ almost surely. Also, $\mathbb E(1+|W_t|^2)^{\omega}<C$ and $\mathbb E\max_{j\leq k_n}\Psi_j(X_t)^2(1+|W_t|^2)^{\omega}<C.$ (ii) The Riesz representer $v_n^*$ satisfies: there are $C, \zeta>0$, $\mathbb E(\max_{j\leq k_n}\Psi_j(X_t)^2+1)v_n^*(W_t)^2< C \mathbb EK_t^2$ and $ \mathbb E | U_t|^{2+\zeta}| K_t|^{2+\zeta} \leq C (\mathbb EK_t^2)^{1+\zeta/2}$, where $K_t:=\mathbb E (v_n^*(W_t)|\sigma_t(\mathcal X)).$
propFor the NPIV model, (i) Assumption (ref) implies Assumptions (ref), (ref), (ref) and (ref). (ii) For the known functional $\phi(\cdot)$, in addition Assumptions (ref), (ref), (ref), (ref), (ref), (ref), (ref), (ref) hold. Then \begin{eqnarray*} \frac{\sqrt{n}(\phi(\widehat\alpha)-\phi(\alpha_0))}{\sigma_n } \to^d \mathcal{N}(0,1), \end{eqnarray*} where $\sigma_n^2:= \operatorname{Var}\left ( \mathbb{E}(v_n^*(W_t)|\sigma_t(\mathcal{X})) \Sigma(X_t)^{-1} U_t \right) .$ (iii) For the unknown functional $\gamma_0= \mathbb{E }l(h_0(W_t)),$ if additionally Assumptions (ref),(ref) hold, then \begin{equation*} \sqrt{n}v^{-1}( \widehat\gamma-\gamma_0 ) \to^d\mathcal{N}(0,1), \end{equation*} where $v^2:=\frac{1}{n}\operatorname{Var}(\sum_t\mathcal W_t- \mathcal Z_t) $ with $\mathcal W_t=l(h_0(W_t))-\Gamma_0(X_t)U_{t+1}$ and $\mathcal Z_t= U_{t+1}\Sigma(X_t)^{-1}\mathbb E [v_n^*(W_t)|\sigma_t(\mathcal X)]$.

The NPQIV model

Consider the nonparametric quantile instrumental variable (NPQIV) model

equation*[equation* omitted — 94 chars of source]

Then $m(X_t, \alpha)= P(U_{t+1}<h-h_0|\sigma_t(\mathcal{X}))-\varpi$ where $ U_{t+1}=y_{t+1}- h_0(W_t)$ and $\alpha=h$. Within this framework, we now verify the high-level assumptions presented in the previous sections.

Suppose the conditional distribution of $U_t$ given $ (X_t, W_t)$ is absolutely continuous with density function $f_{U_t|\sigma_t( \mathcal{X}), W_t}(u)$. In this context, $\Sigma(X_t)$ is known, given by

equation*[equation* omitted — 118 chars of source]

Then the Gateaux derivative is defined as

equation*[equation* omitted — 145 chars of source]

implying, for $g_1= f_{U_t|\sigma_t(\mathcal{X}),W_t}(0) u_n(W_t) $ and $ g_2=f_{U_t|\sigma_t(\mathcal{X}),W_t}(0) (h(W_t)-h_0(W_t)) $,

equation*[equation* omitted — 172 chars of source]

Also, $\|v_n^*\| ^2=(\varpi-\varpi^2)^{-1}\mathbb{E }g(X_t)^2 $ where $ g(X_t)= \mathbb{E }[ f_{U_t|\sigma_t(\mathcal{X}),W_t}(0) v^*_n(W_t)|\sigma_t(\mathcal{X})]. $

We impose the following low-level conditions to verify Assumptions (ref) and (ref). Let

eqnarray*[eqnarray* omitted — 191 chars of source]
assum(i) There are $c_1, c_2, \epsilon_0>0$ so that for all $\|h-h_0\|_{\infty,\omega}< \epsilon_0$, \begin{equation*} c_2\mathbb{E }B_t(h ,h)^2 \Sigma(X_t)^{-1}\leq \mathbb{E }B_t(h_0,h)^2 \Sigma(X_t)^{-1}\leq c_1 \mathbb{E }B_t(h ,h)^2 \Sigma(X_t)^{-1} . \end{equation*} (ii) Almost surely, $\sup_{u} f^{^{\prime }}_{U_t|\sigma_t(\mathcal{X}),W_t} (u) <C$ and $\sup_{u, x,w} f_{U_t|\sigma_t(\mathcal{X}),W_t} (u) <C.$ Also and there is $L>0$, for all $u$, almost surely, $\sup_{x, w}|f_{U_t|\sigma_t( \mathcal{X}),W_t}(u)-f_{U_t|\sigma_t(\mathcal{X}),W_t}( 0)|\leq L |u|. $ (iii) $\mathbb{E}[\max_{j\leq k_n}\Psi_j(X_t)^2+A_t(h_0)^2](1+|W_t|^2)^{\omega}<C $ and $ \mathbb{E } [u_n(W_t)^4|\sigma_t(\mathcal{X}) ]<C$. (iv) $\delta_n^2k_n=o(1)$ and $\delta_n^4=o(n^{-1})$.

The following proposition, proved in the appendix, is the main result in this subsection, which verifies the high-level conditions in the NPQIV context.

propFor the NPQIV model, (i) Assumption (ref) implies Assumptions (ref), (ref), (ref) and (ref). (ii) For the known functional $\phi(\cdot)$, in addition Assumptions (ref), (ref), (ref), (ref), (ref), (ref), (ref), (ref) hold. Then \begin{eqnarray*} \frac{\sqrt{n}(\phi(\widehat\alpha)-\phi(\alpha_0))}{\sigma_n } \to^d \mathcal{N}(0,1), \end{eqnarray*} where $\sigma_n^2:=(\varpi-\varpi^2)^{-1}\mathbb{E}\left([\mathbb{E } f_{U_t|\sigma_t(\mathcal{X}),W_t}(0) v^*_n(W_t)|\sigma_t(\mathcal{X} )]^2\right) .$ (iii) For the unknown functional $\gamma_0= \mathbb{E }l(h_0(W_t)),$ if additionally Assumptions (ref),(ref) hold, then \begin{equation*} \sqrt{n}v^{-1}( \widehat\gamma-\gamma_0 ) \to^d\mathcal{N}(0,1), \end{equation*} where $v^2:=\frac{1}{n}\operatorname{Var}(\sum_t\mathcal W_t - \mathcal Z_t) $ with $\mathcal W_t=l(h_0(W_t))-\Gamma_0(X_t)U_{t+1}$ and \\$\mathcal Z_t= (\varpi-\varpi^2)^{-1}1\{U_{t+1}\leq0\} \mathbb{E } f_{U_t|\sigma_t(\mathcal{X}),W_t}(0) v^*_n(W_t)|\sigma_t(\mathcal{X} )$.

Simulation Studies

In this section, we set up nonparametric endogenous models to illustrate the performance of our proposed estimators and testing statistics using some synthetic data. Consider the following data generating process

equation*[equation* omitted — 66 chars of source]

where

equation*[equation* omitted — 104 chars of source]

and $\phi(\alpha) = \mathbb E[\partial h/\partial Z_t]=\vartheta_0 =1$ is the quantity to be estimated. We choose $L=3, b_l = 0.4^l$ and consider the nonlinear mapping $f(x) = \frac{1-\exp(-x)}{1+\exp(-x)}$. The endogenous $Z_t$ is generated using the following auto-regressive model:

equation*[equation* omitted — 171 chars of source]

And $e_t$ is generated with the following ARCH model using $\varepsilon_t$ as the innovation:

equation*[equation* omitted — 101 chars of source]

We set $\rho = 0.5$ to make $Z_t$ endogenous. We also make $e_t$ heterogeneous. Note that $\mathbb E[e_t^2] = \mathbb E[\sigma_t^2] = 1$. The endogenous variable is $W_t= Z_t$. The instruments are $X_t= (Z_{t-1}, Y_{t-1},...,Y_{t-L})$. We chose to generate $n = 5000$ samples (some burning period has been thrown away to make sure data are stationary). Note that the model can be used for both NPIV and NPQIV with $\varpi = 0.5$.

We applied a fully-connected $J$-layer ReLU-activated NN with hidden layer width of $K$. The optimization of the unconstrained NPIV or NPQIV objective used vanilla gradient descent. We did not apply mini-batch in gradient descent training as using mini-batches may hurt performance due to insufficient smoothing. The training epoch was as large as $10000$ with learning rate $0.01$ for NPIV and $0.1$ for NPQIV. Furthermore we did not apply any penalty term for this example since the problem is relatively easy and the NN under consideration is of a small scale. The linear sieve bases $(\Psi_1,...,\Psi_{k_n})$ for the instrumental variable space were $\tilde k_n$ cubic B-splines for $X$ and each of the three $Y$ lags concatenated together. For simplicity, no interaction terms between X and Y lags were included. Thus in total, we have $k_n = 4 \tilde k_n - 3$ bases (since all B-spline bases sum up to 1, we remove the last basis for each dimension and finally add the intercept term as another basis). In our simulations, we find that NPQIV requires more number of sieve basis $k_n$ for estimating the instrumental space.

For the NPIV problem, we first optimize the equal weighted quadratic loss to obtain $\widehat h$, which is used to estimate $\Sigma(X_t)$ and $ \Gamma(X_t) $ consistently. In the second step, we optimize the optimally weighted quadratic loss with the weighting matrix $\widehat\Sigma(X_t)^{-1}$ and apply the forward filter to estimate our expectation functional, which in this example is the constant $\vartheta_0=1$. Finally, we carry out the hypothesis testing for $H_0: \phi(h) =\mathbb E[\partial h/\partial Z_t] = \phi_0 = 1$ to check the size of the testing statistic. Specifically, we estimated the forward filtered residuals as $\widehat{\mathcal{W}}_t= \partial\widehat h(W_t)/\partial W_t - \widehat\Gamma_t (Y_t - \widehat h(W_t))$ and estimated $ \Sigma_2 = \operatorname{Var}(\mathcal{W}_t)$ by the Newey-West estimator given $\widehat{ \mathcal{W}}_t$, then solved the constrained optimization of $L_n(h, \phi_0)$ and finally constructed the testing statistic. For NPQIV problem, since the optimal weighting is proportional to equal weighting, we do not need the initial step to estimate $\Sigma(X_t)$. So we directly optimized the optimally weighted quadratic loss and estimated $\Gamma(X_t)$ using the results and then used the forward filter to correct the estimation of the average partial derivative. Finally, similar to NPIV, we conduct the hypothesis testing for $H_0: \phi(\alpha) = 1$ under NPQIV.

table[table omitted — 1,252 chars of source]

As for the computational practice, we find that for NPQIV models, it is helpful to apply truncations to the learned gradients in each step of training the network. Specifically, we smooth the loss function of the NPQIV model and truncate the updated gradient: $$ \theta_{k+1}= \theta_k - \text{lr} *\min\{|\nabla L_{n,k}|, 0.001\}*\text{sgn}(\nabla L_{n,k}) $$ where lr is the learning rate, fixed to be 0.1 for NPQIV; $\nabla L_{n,k} $ is the gradient of the NN at the current step; $\theta_{k+1}$ is the updated neural network coefficients at the current step. The truncation prevents the network from having very large gradients during iterations, helping stabilize the training process empirically.

We repeat each setting for $1000$ times. For the efficient estimation, we report the mean and standard deviation of the forward filtered average gradient for the optimal weighting optimizaiton in Table (ref). For hypothesis testing, we also report in Table (ref) the mean, standard deviation and 95% quantile of the empirical testing statistic. In addition, if we use the theoretical critical value corresponding to 5% significance level, which is 3.84 for $\chi^2_1$, the p-value is also reported.

As we can see from Table (ref), for NPIV, optimal weighting estimates $ \phi(h)$ accurately in the sense that the mean insignificantly differs from the true value $\vartheta_0 = 1$. NPQIV is less efficient with a larger standard deviation, and thus requires more samples to be estimated to the same accuracy. Note that the instrumental space with a step function can be harder to approximate with the cubic B-spline linear sieve bases. In terms of the performance of QLR testing statistic, the p-values are all close to the nominal 5% level for the NPIV and NPQIV models. Admittedly through our experiments the results can be sensitive to some tuning parameters, which is typically the case when applying deep learning for statistical inference: at the moment we still heavily rely on ad-hoc tuning in many problems. In comparison, the estimation of $\phi(h)$ is more stable with respect to different $J$ and $K$ values. Here we only mean to present some results without heavily tuning the parameters. Methods using NN for real applications require more extensive tuning in practice and some rough sense on the model complexity would be useful to determine the balance between the dimensions of the NN sieve and the linear IV sieve.

Conclusion

In this paper we establish neural network estimation and inference on functionals of unknown function satisfies a general time series conditional moment restrictions containing endogenous variables. We consider quasi-likelihood ratio (GN-QLR) based inference, where nonparametric functions are learned using multilayer neural networks. While the asymptotic normality of the estimated functionals depends on some unknown Riesz representer of the functional space, we show that the GN-QLR statistic is asymptotically Chi-square distributed, regardless whether the expectation functional is regular (root-$n$ estimable) or not. This holds when the data are weakly dependent and satisfy the beta-mixing condition.

In addition to estimating partial derivatives in nonparametric endogenous problems as examples, our study is well motivated by the setting of reinforcement learning where data are time series in nature. We apply our method to the off-policy evaluation, by formulating the Bellman equation into the conditional moment restriction framework, so that we can make inference about the state-specific value functional using the proposed GN-QLR method with time series data.