Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
49,356 characters · 10 sections · 1 citation commands
Alternative Asymptotics and the Partially Linear Model with Many Regressors
JEL classification: C13, C31.
Keywords: non-standard asymptotics, partially linear model, many terms, adjusted variance.
\thispagestyle{empty} \setcounter{page}{0}
Many instrument asymptotics, where the number of instruments grows as fast as the sample size, has proven useful for instrumental variables (IV)\ estimators. \citet*{Kunitomo_1980_JASA} and \citet*{Morimune_1983_ECMA} derived asymptotic variances that are larger than the usual formulae when the number of instruments and sample size grow at the same rate, and \citet*{Bekker_1994_ECMA} and others provided consistent estimators of these larger variances. \citet*{Hansen-Hausman-Newey_2008_JBES} showed that using many instrument standard errors provides a theoretical improvement for a range of number of instruments and a practical improvement for estimating the returns to schooling. Thus, many instrument asymptotics and the associated standard errors have been demonstrated to be a useful alternative to the usual asymptotics for instrumental variables.
Instrumental variable estimators implicitly depend on a nonparametric series estimator. Many instrument asymptotics has the number of series terms growing so fast that the series estimator is not consistent. Analogous asymptotics for kernel-based density-weighted average derivative estimators has been considered by Cattaneo-Crump-Jansson_2010_JASA,Cattaneo-Crump-Jansson_2014a_ET . They show that when the bandwidth shrinks faster than needed for consistency of the kernel estimator, the variance of the estimator is larger than the usual formula. They also find that correcting the variance provides an improvement over standard asymptotics for a range of bandwidths.
The purpose of this paper is to show that these results share a common structure, and to illustrate how this structure can be used to derive new results. The common structure is that the object determining the limiting distribution is a V-statistic, which can be decomposed into a bias term, a sample average, and a \textquotedblleft remainder\textquotedblright\ that is an asymptotically normal degenerate U-statistic. Asymptotic normality of the remainder distinguishes this setting from other ones involving V-statistics. Here the asymptotically normal remainder comes from the number of series terms going to infinity or bandwidth shrinking to zero, while the behavior of a degenerate U-statistic tends to be more complicated in other settings. When the number of terms grows as fast as the sample size, or the bandwidth shrinks to zero at an appropriate rate, the remainder has the same magnitude as the leading term, resulting in an asymptotic variance larger than just the variance of the leading term. The many instrument and small bandwidth results share this structure. In keeping with this common structure, we will henceforth refer to such results under the general heading of \textquotedblleft alternative asymptotics\textquotedblright.
The alternative asymptotics that we discuss in this paper applies to statistics that take a specific V-statistic representation, or may be approximated by it sufficiently accurately, and therefore it does not apply broadly to all possible semiparametric settings. Nonetheless, as we illustrate below, this structure arises naturally in several interesting problems in Economics and Statistics. In particular, we show formally that applying this common structure to a series estimator of the partially linear model leads to new results. These results allow the number of terms in the series approximation to grow as fast as the sample size. The asymptotic distribution of the estimator is derived and it is shown to have a larger asymptotic variance than the usual formula, which is in fact a natural and generic consequence of the specific structure that we highlight in this paper. We also find that under homoskedasticity, the classical degrees of freedom adjusted homoskedastic standard error estimator from linear models is consistent even when the number of terms is \textquotedblleft large\textquotedblright \ relative to the sample size. This result offers a large sample, distribution free justification for the degrees of freedom correction when many series terms are employed. Constructing automatic consistent standard error estimator under (conditional) heteroskedasticity of unknown form in this setting turns out to be quite challenging. In \citet*{Cattaneo-Jansson-Newey_2015_HCSEManyCov} , we present a detailed discussion of heteroskedasticity-robust standard errors for general linear models with increasing dimension, which covers the partially linear model with many terms studied herein as a special case.
The rest of the paper is organized as follows. Section (ref) describes the common structure of many instrument and small bandwidth asymptotics, and also shows how the structure leads to new results for the partially linear model. Section (ref) formalizes the new distributional approximation for the partially linear model. Section (ref) reports results from a small simulation study aimed to illustrate our results in small samples. Section (ref) concludes. The appendix collects the proofs of our results.
To describe the common structure of many instrument and small bandwidth asymptotics, let $W_{1},\ldots,W_{n}$ denote independent random vectors. We consider an estimator $\hat{\beta}$ of a generic parameter of interest $\beta_{0}\in\mathbb{R}^{d}$ satisfying
where $u_{ij}^{n}(\cdot)$ is a function that can depend on $i$, $j$, and $n$. We allow $u$ to depend on $n$ to account for number of terms or bandwidths that change with the sample size. Also, we allow $u$ to vary with $i$ and $j$ to account for dependence on variables that are being conditioned on in the asymptotics, and so treated as nonrandom.
We assume throughout this section that there exists a sequence of non-random matrices $\Gamma_{n}$ satisfying $\Gamma_{n}^{-1}\hat{\Gamma}_{n} \rightarrow_{p}I_{d}$ for $I_{d}$ the $d\times d$ identity matrix, and hence we focus on the V-statistic $S_{n}$. (All limits are taken as $n\rightarrow \infty$ unless explicitly stated otherwise.)\ This V-statistic has a well known (Hoeffding-type) decomposition that we describe here because it is an essential feature of the common structure. For notational implicitly we will drop the $W_{i}$ and $W_{j}$ arguments and set $u_{ij}^{n}=u_{ij}^{n} (W_{i},W_{j})$ and $\tilde{u}_{ij}^{n}=u_{ij}^{n}+u_{ji}^{n}-\mathbb{E} [u_{ij}^{n}+u_{ji}^{n}]$.
Letting $\Vert\cdot\Vert$ denote the Euclidean norm, and if $\mathbb{E}[\Vert u_{ij}^{n}\Vert]<\infty$\ for all $i,j,n$, then
where \[ B_{n}=\mathbb{E}[S_{n}],\text{\qquad}\Psi_{n}=\sum_{1\leq i\leq n}\psi_{i} ^{n}(W_{i}),\text{\qquad}U_{n}=\sum_{2\leq i\leq n}D_{i}^{n}(W_{i},...,W_{1}), \] \[ \psi_{i}^{n}(W_{i})=u_{ii}^{n}-\mathbb{E}[u_{ii}^{n}]+\sum_{1\leq j\leq n,j\neq i}\mathbb{E}[\tilde{u}_{ij}^{n}|W_{i}], \] \[ D_{i}^{n}(W_{i},...,W_{1})=\sum_{1\leq j\leq n,j<i}(\tilde{u}_{ij} ^{n}-\mathbb{E}[\tilde{u}_{ij}^{n}|W_{i}]-\mathbb{E}[\tilde{u}_{ij}^{n} |W_{j}]), \] so that $\mathbb{E}[\psi_{i}^{n}(W_{i})]=0$, $\mathbb{E}[D_{i}^{n} (W_{i},...,W_{1})|W_{i-1},...,W_{1}]=0$, and $\mathbb{E}[\Psi_{n}U_{n}]=0$. This decomposition of a V-statistic is well known (e.g., \citet*[Chapter 11]{vanderVaart_1998_Book}), and shows that $S_{n}$ can be decomposed into a sum $\Psi_{n}$ of independent terms, a U-statistic remainder $U_{n}$ that is a martingale difference sum and uncorrelated with $\Psi_{n}$, and a pure bias term $B_{n}$.\footnote{In time series contexts, the exact decomposition is less useful, but approximations thereof with properties similar to those we discuss herein can be developed. For an example and related references see \citet*{Atchade-Cattaneo_2014_SPA}.} The decomposition is important in many of the proofs of asymptotic normality of semiparametric estimators, including \citet*{Powell-Stock-Stoker_1989_ECMA}, with the limiting distribution being determined by $\Psi_{n}$, and $U_{n}$ being treated as a \textquotedblleft remainder\textquotedblright\ that is of smaller order under a particular restriction on the tuning parameter sequence (e.g., when the bandwidth shrinks slowly enough).
An interesting feature of the decomposition ((ref)) in semiparametric settings is that $U_{n}$ is asymptotically normal at some rate when the number of series terms grow or the bandwidth shrinks to zero. To be specific, under regularity conditions and appropriate tuning parameter sequences that we make precise below, it turns out that \[ \left[
\right] \rightarrow_{d}\mathcal{N}(0,I_{2d}). \] In other settings, where the underlying kernel of the U-statistic does not vary with the sample size, the asymptotic behavior of $U_{n}$ is usually more complicated: because it is a degenerate U-statistic, it would converge to a weighted sum of independent chi-square random variables (e.g., \citet*[Chapter 12]{vanderVaart_1998_Book}). However, in semiparametric-type settings as those considered here, the kernel of the underlying U-statistic forming $U_{n}$ changes with the sample size and hence, under particular tuning parameter configurations, the individual contributions $D_{i}^{n}(W_{i},...,W_{1})$ to $U_{n}$ can be made small enough to satisfy a Lindeberg-Feller condition and thus obtain a Gaussian limiting distribution (usually employing the martingale property of $U_{n}$). For an interesting discussion of this phenomenon, see \citet*{deJong_1987_PTRF}. The asymptotic normality property of $U_{n}$ has been shown for certain classes of both series and kernel based estimators, as further explained below.
Alternative asymptotics occurs when the number of series terms grows or the bandwidth shrinks fast enough so that $\mathbb{V}[\Psi_{n}]$ and $\mathbb{V}[U_{n}]$ have the same magnitude in the limit. Because of uncorrelatedness of $\Psi_{n}$ and $U_{n}$, the asymptotic variance will be larger than the usual formula which is $\lim_{n\rightarrow\infty} \mathbb{V}[\Psi_{n}]$ (assuming the limit exists). As a consequence, consistent variance estimation under alternative asymptotics requires accounting for the contribution of $U_{n}$ to the (asymptotic)\ sampling variability of the statistic. Accounting for the presence of $U_{n}$ should also yield improvements when numbers of series terms and bandwidths do not satisfy the knife-edge conditions of alternative asymptotics, since $U_{n}$ is part of the semiparametric statistic. For instance, if the number of series terms grows just slightly slower than the sample size then accounting for the presence of $U_{n}$ should still give a better large sample approximation. \citet*{Hansen-Hausman-Newey_2008_JBES} show such an improvement for many instrument asymptotics. It would be good to consider such improved approximations more generally, though it is beyond the scope of this paper to do so.
Distribution theory under alternative asymptotics may be seen as a generalization of the conventional large sample distributional approximation approach in the sense that under conventional sequences of tuning parameters the asymptotic variances emerging from both approaches coincide. But, the alternative asymptotic approximation also allows for other tuning parameter sequences and, in this case, the limiting asymptotic variance is seen to be larger than usual. Thus, in general, there is no reason to expect that the usual standard error formulas derived under conventional asymptotics will remain valid more generically. From this perspective, alternative asymptotics are useful to provide theoretical justification for new standard error formulas that are consistent under more general sequences of tuning parameters, that is, under both conventional and alternative asymptotics. We refer to the latter standard error formulas as being more robust than the usual standard error formulas available in the literature. For instance, using these ideas, the need for new, more robust standard errors formulas was made before for many instrument asymptotics in IV\ models (\citet*{Hansen-Hausman-Newey_2008_JBES}) and small bandwidth asymptotics in kernel-based semiparametrics (\citet*{Cattaneo-Crump-Jansson_2014a_ET}).
To illustrate these ideas, we show next that both many instrument asymptotics and small bandwidth asymptotics have the structure described above, and we also employ this approach to derive new results in the case of a series estimator of the partially linear model, which we refer to as \textquotedblleft many terms asymptotics\textquotedblright.
The first example is concerned with the case of many instrument asymptotics. For simplicity we focus on the JIVE2 estimator of \citet*{Angrist-Imbens-Krueger_1999_JAE}, but the idea applies to other IV estimators such as the limited information maximum likelihood estimator. See \citet*{Chao-Swanson-Hausman-Newey-Woutersen_2012_ET} for more details, including regularity conditions under which the following discussion can be made rigorous.
Let $(y_{i},x_{i}^{\prime},z_{i}^{\prime})^{\prime}$, $i=1,\ldots,n$, be a random sample generated by the model
where $y_{i}$ is a scalar dependent variable, $x_{i}\in\mathbb{R}^{d}$ is a vector of endogenous variables, $\varepsilon_{i}$ is a disturbance, and $z_{i}\in\mathbb{R}^{K}$ is a vector of instrumental variables.
To describe the JIVE2 estimator of $\beta_{0}$ in ((ref)), let $Q_{ij}$ denote the $(i,j)$-th element of $Q=Z(Z^{\prime}Z)^{-1}Z^{\prime}$, where $Z=[z_{1},\cdots,z_{n}]^{\prime}$. After centering and scaling, the JIVE2 estimator $\hat{\beta}$ satisfies \[ \sqrt{n}(\hat{\beta}-\beta_{0})=(\frac{1}{n}\sum_{1\leq i,j\leq n,j\neq i}Q_{ij}x_{i}x_{j}^{\prime})^{-1}(\frac{1}{\sqrt{n}}\sum_{1\leq i,j\leq n,j\neq i}Q_{ij}x_{i}\varepsilon_{j}). \] Conditional on $Z,$ $\hat{\beta}$ has the structure in ((ref)) with $W_{i}=(x_{i}^{\prime},\varepsilon_{i})^{\prime}$ and \[ \hat{\Gamma}_{n}=\frac{1}{n}\sum_{1\leq i,j\leq n,j\neq i}Q_{ij}x_{i} x_{j}^{\prime},\text{\qquad}u_{ij}^{n}(W_{i},W_{j})= {\rm 1\hspace*{-0.4ex}\rule{0.1ex}{1.52ex}\hspace*{0.2ex}} (i\neq j)Q_{ij}x_{i}\varepsilon_{j}/\sqrt{n}, \] where $ {\rm 1\hspace*{-0.4ex}\rule{0.1ex}{1.52ex}\hspace*{0.2ex}} (\cdot)$ is the indicator function.
For $i\neq j$, $\mathbb{E}[u_{ij}^{n}(W_{i},W_{j})|Z]=0$ and \[ \mathbb{E}[u_{ij}^{n}(W_{i},W_{j})|W_{i},Z]=Q_{ij}x_{i}\mathbb{E} [\varepsilon_{j}|Z]=0,\text{\hspace{0.2in}}\mathbb{E}[u_{ji}^{n}(W_{j} ,W_{i})|W_{i},Z]=Q_{ij}\Upsilon_{j}\varepsilon_{i}/\sqrt{n}, \] where $\Upsilon_{i}=\mathbb{E}[x_{i}|z_{i}]$ can be interpreted as the reduced form for observation $i$. As a consequence, ((ref)) is satisfied with $B_{n}=0$, \[ \psi_{i}^{n}(W_{i})=(\sum_{1\leq j\leq n,j\neq i}Q_{ij}\Upsilon_{j} )\varepsilon_{i}=\Upsilon_{i}(1-Q_{ii})\varepsilon_{i}/\sqrt{n}-(\Upsilon _{i}-\sum_{1\leq j\leq n}Q_{ij}\Upsilon_{j})\varepsilon_{i}/\sqrt{n}, \] \[ D_{i}^{n}(W_{i},...,W_{1})=\sum_{1\leq j\leq n,j<i}Q_{ij}\left( v_{i}\varepsilon_{j}+v_{j}\varepsilon_{i}\right) /\sqrt{n},\text{\hspace {0.2in}}v_{i}=x_{i}-\Upsilon_{i}. \]
Because $\Upsilon_{i}-\sum_{j=1}^{n}Q_{ij}\Upsilon_{j}$ is the $i$-th residual from regressing the reduced form observations on $Z$, by appropriate definition of the reduced form this can generally be assumed to vanish as the sample size grows. In that case, \[ \Psi_{n}=\frac{1}{\sqrt{n}}\sum_{1\leq i\leq n}\Upsilon_{i}(1-Q_{ii} )\varepsilon_{i}+o_{p}(1). \] Furthermore, under standard asymptotics $Q_{ii}$ will go to zero, so the limiting variance of the leading term in $\Psi_{n}$ corresponds to the usual asymptotic variance for IV. The degenerate U-statistic term is \[ U_{n}=\frac{1}{\sqrt{n}}\sum_{1\leq i,j\leq n,j<i}Q_{ij}\left( v_{i} \varepsilon_{j}+v_{j}\varepsilon_{i}\right) . \] \citet*{Chao-Swanson-Hausman-Newey-Woutersen_2012_ET} apply a martingale central limit theorem to show that this $U_{n}$ will be asymptotically normal when $K\rightarrow\infty$ and certain regularity conditions hold. The conditions of the martingale central limit theorem are verified by showing that certain linear combinations with coefficients depending on the elements of $Q$ go to zero as $K\rightarrow\infty$. In the proof, this makes individual terms asymptotically negligible, with a Lindeberg-Feller condition being satisfied. Alternative asymptotics occurs when $K$ grows as fast as $n$, resulting in $\mathbb{V}[\Psi_{n}]$ and $\mathbb{V}[U_{n}]$ having the same magnitude in the limit.
The second example shows that small bandwidth asymptotics for certain kernel-based semiparametric estimators also has the structure outlined above. To keep the exposition simple we focus on an estimator of the integrated squared density, but the structure of this estimator is shared by the density-weighted average derivative estimator of \citet*{Powell-Stock-Stoker_1989_ECMA} treated in \citet*{Cattaneo-Crump-Jansson_2014a_ET} and more generally by estimators of density-weighted averages and ratios thereof (see, e.g., \citet*[Section 2]{Newey-Hsieh-Robins_2004_ECMA} and references therein).
Suppose $x_{i}$, $i=1,\ldots,n$, are i.i.d. continuously distributed $p$-dimensional random vectors with smooth p.d.f. $f_{0}$ and consider estimation of the integrated squared density \[ \beta_{0}=\int_{\mathbb{R}^{p}}f_{0}(x)^{2}\text{d}x=\mathbb{E}[f_{0}(x_{i})]. \] A leave-one-out kernel-based estimator is \[ \hat{\beta}=\sum_{1\leq i,j\leq n,i\neq j}\mathcal{K}_{h}(x_{i}-x_{j})/n(n-1), \] where $\mathcal{K}(u)$ is a symmetric kernel and $\mathcal{K}_{h} (u)=h^{-p}\mathcal{K}(u/h)$. This estimator has the V-statistic form of ((ref)) with $W_{i}=x_{i}$ and \[ \hat{\Gamma}_{n}=1,\text{\hspace{0.2in}}u_{ij}^{n}(W_{i},W_{j})= {\rm 1\hspace*{-0.4ex}\rule{0.1ex}{1.52ex}\hspace*{0.2ex}} (i\neq j)\{\mathcal{K}_{h}(x_{i}-x_{j})-\beta_{0}\}/\sqrt{n}(n-1). \]
Let $f_{h}(x)=\int_{\mathbb{R}^{p}}\mathcal{K}(u)f_{0}(x+hu)\text{d}u$ and $\beta_{h}=\int_{\mathbb{R}^{p}}f_{h}(x)f_{0}(x)$d$x$. By symmetry of $\mathcal{K}(u)$, \[ \mathbb{E}[u_{ij}^{n}(W_{i},W_{j})|W_{i}]=\mathbb{E}[u_{ji}^{n}(W_{j} ,W_{i})|W_{i}]=\{f_{h}(x_{i})-\beta_{0}\}/\sqrt{n}(n-1), \] \[ \mathbb{E}[u_{ij}^{n}(W_{i},W_{j})]=\{\beta_{h}-\beta_{0}\}/\sqrt{n}(n-1), \] so the terms in the decomposition ((ref)) are of the form \[ B_{n}=\sqrt{n}\{\beta_{h}-\beta_{0}\},\text{\qquad}\Psi_{n}=\frac{1}{\sqrt{n} }\sum_{1\leq i\leq n}2\{f_{h}(x_{i})-\beta_{h}\}, \] \[ U_{n}=\frac{2}{\sqrt{n}(n-1)}\sum_{1\leq i,j\leq n,j<i}\{\mathcal{K}_{h} (x_{i}-x_{j})-f_{h}(x_{i})-f_{h}(x_{j})+\beta_{h}\}. \]
Here, $2\{f_{h}(x_{i})-\beta_{h}\}$ is an approximation to the well known influence function $2\{f_{0}(x_{i})-\beta_{0}\}$ for estimators of the integrated squared density. Under regularity conditions, $f_{h}(x_{i})$ converges to $f_{0}(x_{i})$ in mean square as $h\rightarrow0$, so that \[ \Psi_{n}=\frac{1}{\sqrt{n}}\sum_{1\leq i\leq n}2\{f_{0}(x_{i})-\beta _{0}\}+o_{p}(1). \] A martingale central limit theorem can be applied as in \citet*{Cattaneo-Crump-Jansson_2014a_ET} to show that the degenerate U-statistic term $U_{n}$ will be asymptotically normal as $h\rightarrow0$ and $n\rightarrow\infty$, provided that $n^{2}h^{p}\rightarrow\infty$. It is easy to show that $n^{2}h^{p}\mathbb{V}[U_{n}]\rightarrow\Delta=\beta_{0} \int_{\mathbb{R}^{p}}\mathcal{K}(u)^{2}$d$u$ (under regularity conditions). Alternative asymptotics occurs when $h^{p}$ shrinks as fast as $1/n$, resulting in $\mathbb{V}[\Psi_{n}]$ and $\mathbb{V}[U_{n}]$ having the same magnitude in the limit.
The previous two examples show how several estimators share the common structure outlined above. To illustrate how this structure can be applied to derive new results, the third example studies series estimation in the context of the partially linear model. The results will shed light on the asymptotic behavior of this estimator, and the associated inference procedures, when the number of terms are allowed to grow as fast as the sample size.
Let $(y_{i},x_{i}^{\prime},z_{i}^{\prime})^{\prime}$, $i=1,\ldots,n$, be a random sample of generated by the partially linear model
where $y_{i}$ is a scalar dependent variable, $x_{i}\in\mathbb{R}^{d}$ and $z_{i}\in\mathbb{R}^{d_{z}}$ are explanatory variables, $\varepsilon_{i}$ is a disturbance, $g(\cdot)$ is an unknown function, and $\mathbb{E}[\mathbb{V} [x_{i}|z_{i}]]$ is of full rank.
A series estimator of $\beta_{0}$ is obtained by regressing $y_{i}$ on $x_{i}$ and approximating functions of $z_{i}$. To describe the estimator, let $p^{1}(z),$ $p^{2}(z),$ $\ldots$ be approximating functions, such as polynomials or splines, and let $p_{K}(z)=(p^{1}(z),\ldots,p^{K}(z))^{\prime}$ be a $K$-dimensional vector of such functions. Letting $M_{ij}$ denote the $(i,j)$-th element of $M=I_{n}-P_{K}(P_{K}^{\prime}P_{K})^{-1}P_{K}^{\prime},$ where $P_{K}=[p_{K}(z_{1}),\ldots,p_{K}(z_{n})]^{\prime}$, a series estimator of $\beta_{0}$ in ((ref)) is given by \[ \hat{\beta}=(\sum_{1\leq i,j\leq n}M_{ij}x_{i}x_{j}^{\prime})^{-1}(\sum_{1\leq i,j\leq n}M_{ij}x_{i}y_{j}). \] \citet*{Donald-Newey_1994_JMA} gave conditions for asymptotic normality of this estimator using standard asymptotics. See also for example \citet*{Linton_1995_ECMA}, references therein, for related asymptotic results when using kernel estimators.
Conditional on $Z=[z_{1},\ldots,z_{n}]^{\prime}$, $\hat{\beta}$ has the structure outlined earlier:
with \[ \hat{\Gamma}_{n}=\frac{1}{n}\sum_{1\leq i,j\leq n}M_{ij}x_{i}x_{j}^{\prime },\qquad S_{n}=\frac{1}{\sqrt{n}}\sum_{1\leq i,j\leq n}x_{i}M_{ij} (g_{j}+\varepsilon_{j}), \] where $g_{i}=g(z_{i}).$ In other words, $\hat{\beta}$ has the V-statistic form of ((ref)) with $W_{i}=(x_{i}^{\prime},\varepsilon_{i})^{\prime}$ and $u_{ij}^{n}(W_{i},W_{j})=x_{i}M_{ij}(g_{j}+\varepsilon_{j})/\sqrt{n}$.
By $\mathbb{E}[\varepsilon_{i}|x_{i},z_{i}]=0$ we have $\mathbb{E} [x_{i}\varepsilon_{i}|Z]=0$. Therefore, letting $u_{ij}^{n}=u_{ij}^{n} (W_{i},W_{j})$ as we have done previously, we have \[ \mathbb{E}[u_{ij}^{n}|Z]=h_{i}M_{ij}g_{j}/\sqrt{n},\text{\qquad}u_{ij} ^{n}-\mathbb{E}[u_{ij}^{n}|Z]=M_{ij}\left( v_{i}g_{j}+x_{i}\varepsilon _{j}\right) /\sqrt{n}, \] \[ \tilde{u}_{ij}^{n}=M_{ij}\left( v_{j}g_{i}+v_{i}g_{j}+x_{j}\varepsilon _{i}+x_{i}\varepsilon_{j}\right) /\sqrt{n},\qquad\mathbb{E}[\tilde{u} _{ij}^{n}|W_{i},Z]=M_{ij}\left( v_{i}g_{j}+h_{j}\varepsilon_{i}\right) /\sqrt{n}, \] for $i\neq j$, where $h_{i}=h(z_{i})=\mathbb{E}[x_{i}|z_{i}]$ and $v_{i} =x_{i}-h_{i}$. In this case, the bias term in ((ref)) is \[ B_{n}=\frac{1}{\sqrt{n}}\sum_{1\leq i,j\leq n}M_{ij}h_{i}g_{j}, \] which will be negligible under regularity conditions, as shown in the next section. Moreover, \[ \Psi_{n}=\frac{1}{\sqrt{n}}\sum_{1\leq i\leq n}M_{ii}v_{i}\varepsilon _{i}+R_{n},\qquad R_{n}=\frac{1}{\sqrt{n}}\sum_{1\leq i,j\leq n}M_{ij} (v_{i}g_{j}+h_{i}\varepsilon_{j}), \] where $R_{n}$ has mean zero and converges to zero in mean square as $K$ grows, as further discussed below. Under standard asymptotics $M_{ii}$ will go to one and hence the limiting variance of the leading term in $\Psi_{n}$ corresponds to the usual asymptotic variance.
Finally, we find that the degenerate U-statistic term is \[ U_{n}=\frac{1}{\sqrt{n}}\sum_{1\leq i,j\leq n,j<i}M_{ij}\left( v_{i} \varepsilon_{j}+v_{j}\varepsilon_{i}\right) =-\frac{1}{\sqrt{n}}\sum_{1\leq i,j\leq n,j<i}Q_{ij}\left( v_{i}\varepsilon_{j}+v_{j}\varepsilon_{i}\right) . \] Remarkably, this term is essentially the same as the degenerate U-statistic term for JIVE2 that was discussed above. Consequently, the central limit theorem of \citet*{Chao-Swanson-Hausman-Newey-Woutersen_2012_ET} is applicable to this problem. We will employ it to show that $U_{n}$ is asymptotically normal as $K\rightarrow\infty$, even when $K/n$ does not converge to zero.
This example highlights a new approach to studying the asymptotic distribution of semi-linear regression under many terms asymptotics. The alternative asymptotic approximation is useful, for instance, when the number of covariates entering the nonparametric part is large relative to the sample size, as is often the case in empirical applications.
In this section we make precise the discussion given in Example 3, and also discuss consistent standard error estimation under homoskedasticity. The estimator $\hat{\beta}$ described in Example 3 can be interpreted as a two-step semiparametric estimator with tuning parameter $K$, the first step involving series estimation of the the unknown (regression) functions $g(z)$ and $h(z)$. \citet*{Donald-Newey_1994_JMA} gave conditions for asymptotic normality of this estimator when $K/n\rightarrow0$. Here we generalize their findings by obtaining an asymptotic distributional result that is valid even when $K/n$ is bounded away from zero.
The analysis proceeds under the following assumption.
Because $\sum_{i=1}^{n}M_{ii}=n-K$, an implication of part (d) is that $K/n\leq1-C<1$, but crucially Assumption PLM does not imply that $K/n\rightarrow0$. Part (e) is implied by conventional assumptions from approximation theory. For instance, when the support of $z_{i}$ is compact commonly used basis of approximation, such as polynomials or splines, will satisfy this assumption with $\alpha_{g}=s_{g}/d_{z}$ and $\alpha_{h} =s_{h}/d_{z}$, where $s_{g}$ and $s_{h}$ denotes the number of continuous derivatives of $g(z)$ and $h(z)$, respectively. Further discussion and related references for several basis of approximation may be found in \citet*{Newey_1997_JoE}, \citet*{Chen_2007_Handbook} and \citet*{Belloni-Chernozhukov-Chetverikov-Kato_2015_JoE}, among others.
From equation ((ref)), and the discussion in the previous section, we see that the asymptotic distribution of $\hat{\beta}$ will be determined by the behavior of $\hat{\Gamma}_{n}$ and $S_{n}$. The following lemma approximates $\hat{\Gamma}_{n}$ without requiring that $K/n\rightarrow0$.
Because $\sum_{i=1}^{n}M_{ii}=n-K$, it follows from this result that in the homoskedastic $v_{i}$ case (i.e., when $\mathbb{E}[v_{i}v_{i}^{\prime} |z_{i}]=\mathbb{E}[v_{i}v_{i}^{\prime}]$) $\hat{\Gamma}_{n}$ is close to \[ \Gamma_{n}=(1-K/n)\Gamma,\qquad\Gamma=\mathbb{E}[v_{i}v_{i}^{\prime}], \] in probability. More generally, with heteroskedasticity, $\hat{\Gamma}_{n}$ will be close to the weighted average $\Gamma_{n}$. Importantly, this result includes standard asymptotics as a special case when $K/n\rightarrow0 $, where $\sum_{i=1}^{n}(1-M_{ii})/n=K/n$, the law of large numbers and iterated expectations imply
Next, we study \[ S_{n}=\frac{1}{\sqrt{n}}\sum_{1\leq i,j\leq n}M_{ij}v_{i}\varepsilon_{j} +B_{n}+R_{n}. \] The following lemma quantifies the magnitude of the bias term $B_{n}$ as well as the additional variability arising from the (remainder) term $R_{n}$.
Like the previous lemma, this lemma does not require $K/n\rightarrow0$. Interestingly, the bias term $B_{n}$ involves approximation of both unknown functions $g(z)$ and $h(z)$, implying an implicit trade-off between smoothness conditions for $g(z)$ and $h(z)$. The implied bias condition $K^{2(\alpha _{g}+\alpha_{h})}/n\rightarrow\infty$ only requires that $\alpha_{g} +\alpha_{h}$ be large enough, but not necessarily that $\alpha_{g}$ and $\alpha_{h}$ separately be large. It follows that if this bias condition holds, then \[ S_{n}=\frac{1}{\sqrt{n}}\sum_{1\leq i,j\leq n}M_{ij}v_{i}\varepsilon_{j} +o_{p}(1), \] as claimed in Example 3 above.
Having dispensed with asymptotically negligible contributions to $S_{n}$, we turn to its leading term. This term is shown below to be asymptotically Gaussian with asymptotic variance given by \[ \Sigma_{n}=\frac{1}{n}\mathbb{V}[\sum_{1\leq i,j\leq n}M_{ij}v_{i} \varepsilon_{j}|Z]=\frac{1}{n}\sum_{1\leq i\leq n}M_{ii}^{2}\mathbb{E} [v_{i}v_{i}^{\prime}\varepsilon_{i}^{2}|z_{i}]+\frac{1}{n}\sum_{1\leq i,j\leq n,j\neq i}M_{ij}^{2}\mathbb{E}[v_{i}v_{i}^{\prime}\varepsilon_{j}^{2} |z_{i},z_{j}]. \] Here, the first term following the second equality corresponds to the usual asymptotic approximation, while the second term adds an additional term that accounts for large $K$. Once again it is interesting to consider what happens in some special cases. Under homoskedasticity of $\varepsilon_{i}$ (i.e., when $\mathbb{E}[\varepsilon_{i}^{2}|x_{i},z_{i}]=\mathbb{E}[\varepsilon_{i}^{2} ]$), \[ \Sigma_{n}=\frac{\sigma_{\varepsilon}^{2}}{n}\sum_{1\leq i,j\leq n}M_{ij} ^{2}\mathbb{E}[v_{i}v_{i}^{\prime}|z_{i}]=\frac{\sigma_{\varepsilon}^{2}} {n}\sum_{1\leq i\leq n}M_{ii}\mathbb{E}[v_{i}v_{i}^{\prime}|z_{i} ]=\sigma_{\varepsilon}^{2}\Gamma_{n},\qquad\sigma_{\varepsilon}^{2} =\mathbb{E}[\varepsilon_{i}^{2}], \] because $\sum_{j=1}^{n}M_{ij}^{2}=M_{ii}$. If, in addition, $\mathbb{E} [v_{i}v_{i}^{\prime}|z_{i}]=\mathbb{E}[v_{i}v_{i}^{\prime}]$, then $\Sigma _{n}=\sigma_{\varepsilon}^{2}\left( 1-K/n\right) \Gamma$. Also, if $K/n\rightarrow0$, then by $\sum_{1\leq i,j\leq n,i\neq j}M_{ij}^{2}/n\leq K/n$ and the law of large numbers, we have \[ \Sigma_{n}=\frac{1}{n}\sum_{1\leq i\leq n}M_{ii}^{2}\mathbb{E}[v_{i} v_{i}^{\prime}\varepsilon_{i}^{2}|z_{i}]+o_{p}\left( 1\right) =\mathbb{E} [v_{i}v_{i}^{\prime}\varepsilon_{i}^{2}]+o_{p}\left( 1\right) , \] which corresponds to the standard asymptotics limiting variance.
The following theorem combines Lemmas (ref) and (ref) with a central limit theorem for quadratic forms to show asymptotic normality of $\hat{\beta}$.
This theorem shows that $\hat{\beta}$ is asymptotically normal when $K/n$ need not converge to zero. An implication of this result is that inconsistent series-based nonparametric estimators of the unknown functions $g(z)$ and $h(z)$ may be employed when forming $\hat{\beta}$, that is, $K/n\nrightarrow0$ is allowed (increasing the variability of the nonparametric estimators), provided that $K\rightarrow\infty$ (to remove nonparametric smoothing bias). This asymptotic distributional result does not rely on asymptotic linearity, nor on the actual convergence of the matrices $\Gamma_{n}$ and $\Sigma_{n}$, and leads to a new (larger) asymptotic variance that captures terms that are assumed away by the classical result. The asymptotic distribution result of \citet*{Donald-Newey_1994_JMA} is obtained as a special case where $K/n\rightarrow0 $. More generally, when $K/n$ does not converge to zero, the asymptotic variance will be larger than the usual formula because it accounts for the contribution of \textquotedblleft remainder\textquotedblright\ $U_{n}$ in equation ((ref)). For instance, when both $\varepsilon_{i}$ and $v_{i}$ are homoskedastic, the asymptotic variance is \[ \Gamma_{n}^{-1}\Sigma_{n}\Gamma_{n}^{-1}=\sigma_{\varepsilon}^{2}\Gamma _{n}^{-1}=\sigma_{\varepsilon}^{2}\Gamma^{-1}(1-K/n)^{-1}\text{,} \] which is larger than the usual asymptotic variance $\sigma_{\varepsilon} ^{2}\Gamma^{-1}$ by the degrees of freedom correction $(1-K/n)^{-1}.$
Consistent asymptotic variance estimation is useful for large sample inference. If the assumptions of Theorem (ref) are satisfied and if $\hat{\Sigma}_{n}-\Sigma_{n}\rightarrow_{p}0$, then \[ \hat{\Omega}_{n}^{-1/2}\sqrt{n}(\hat{\beta}-\beta_{0})\rightarrow _{d}\mathcal{N}(0,I_{d}),\text{\qquad}\hat{\Omega}_{n}=\hat{\Gamma}_{n} ^{-1}\hat{\Sigma}_{n}\hat{\Gamma}_{n}^{-1}, \] implying that valid large-sample confidence intervals and hypothesis tests for linear and nonlinear transformations of the parameter vector $\beta$ can be based on $\hat{\Omega}_{n}$.\footnote{Another approach to inference would be via the bootstrap. For small bandwidth asymptotics, \citet*{Cattaneo-Crump-Jansson_2014b_ET} showed that the standard nonparametric bootstrap does not provide a valid distributional approximation in general. We conjecture that the standard nonparametric bootstrap will also fail to provide valid inference for other alternative asymptotics frameworks.} Under (conditional)\ heteroskedasticity of unknown form, constructing a consistent estimator $\hat{\Sigma}_{n}$ turns out to be very challenging if $K/n\nrightarrow0$. Intuitively, the problem arises because the estimated residuals entering the construction of $\hat{\Sigma}_{n}$ are not consistent unless $K/n\rightarrow0$, implying that $\hat{\Sigma}_{n}-\Sigma _{n}\nrightarrow_{p}0$ in general. Solving this problem is beyond the scope of this paper. Under homoskedasticity of $\varepsilon_{i}$, however, the asymptotic variance $\Sigma_{n}$ simplifies and admits a correspondingly simple consistent estimator. To describe this result, note that if $\mathbb{E}[\varepsilon_{i}^{2}|x_{i},z_{i}]=\sigma_{\varepsilon}^{2}$ then $\Sigma_{n}=\sigma_{\varepsilon}^{2}\Gamma_{n}$, where $\hat{\Gamma} _{n}-\Gamma_{n}\rightarrow_{p}0$ by Lemma (ref). It therefore suffices to find a consistent estimator of $\sigma_{\varepsilon}^{2}$. Let \[ s^{2}=\frac{1}{n-d-K}\sum_{1\leq i\leq n}\hat{\varepsilon}_{i}^{2},\qquad \hat{\varepsilon}_{i}=\sum_{1\leq j\leq n}M_{ij}(y_{j}-\hat{\beta}^{\prime }x_{j}), \] denote the usual OLS estimator of $\sigma_{\varepsilon}^{2}$ incorporating a degrees of freedom correction. The following theorem shows that $s^{2}$ is a consistent estimator, even when the number of terms is \textquotedblleft large\textquotedblright\ relative to the sample size.
This theorem provides a distribution free, large sample justification for the degrees-of-freedom correction required for exact inference under homoskedastic Gaussian errors. Intuitively, accounting for the correct degrees of freedom is important whenever the number of terms in the semi-linear model is \textquotedblleft large\textquotedblright\ relative to the sample size.
We conducted a Monte Carlo experiment to explore the extent to which the asymptotic theoretical results obtained in the previous section are present in small samples. Using the notation already introduced, we consider the following partially linear model: \[
\] where $d=1$, $\beta=1$, $d_{z}=5$, $z_{i}=(z_{1i},\cdots,z_{d_{z}i})^{\prime}$ with $z_{\ell i}\sim$ i.i.d. $\mathsf{Uniform}(-1,1) $, $\ell=1,\cdots,d_{z}$. The unknown regression functions are set to $g(z_{i})=h(z_{i})=\exp(\Vert z_{i}\Vert^{2})$, which are not additive separable in the covariates $z_{i}$. The simulation study is based on $S=5,000$ replications, each replication taking a random sample of size $n=500 $ with all random variables generated independently. We consider $6$ data generating processes (DGPs) as follows: \[
\] Specifically, Models 1, 3 and 5 correspond to homoskedastic (in $v_{i}$) DGPs, while Models 2, 4 and 5 correspond to heteroskedastic (in $v_{i}$) DGPs. For the latter models, the constant $\varsigma$ was chosen so that $\mathbb{E} [v_{i}^{2}]=1$. The three distributions considered for the unobserved error terms $\varepsilon_{i}$ and $v_{i}$ are: the standard Normal (labelled \textquotedblleft Gaussian\textquotedblright) and two Mixture of Normals inducing either an asymmetric or a bimodal distribution; their Lebesgue densities are depicted in Figure (ref) . We explored other specifications for the regression functions, heteroskedasticity form, and distributional assumptions, but we do not report these additional results because they were qualitative similar to those discussed here.
The estimators considered in the Monte Carlo experiment are constructed using power series approximations. We do not impose additive separability on the basis, though we do restrict the interaction terms to not exceed degree 5. To be specific, we consider the following polynomial basis expansion: \[
\]
Thus, our simulations explore the consequences of introducing many terms in the partially linear model by varying $K$ on the grid above from $K=6$ to $K=277$, which gives a range for $K/n$ of $\{0.012,\cdots,0.554\}$. For each point on the grid of $K/n$, we report average bias, average standard deviation, mean square error and average standarized bias of $\hat{\beta}$ across simulations. We also consider the coverage error rates and interval length for two asymptotic $95\%$ confidence intervals: \[ \text{CI}_{0}=\left[ \hat{\beta}-\Phi_{1-\alpha/2}^{-1}\frac{\hat{\sigma} \hat{\Gamma}_{n}^{-1/2}}{\sqrt{n}}\quad,\quad\hat{\beta}+\Phi_{1-\alpha /2}^{-1}\frac{\hat{\sigma}\hat{\Gamma}_{n}^{-1/2}}{\sqrt{n}}\right] , \] \[ \text{CI}_{1}=\left[ \hat{\beta}-\Phi_{1-\alpha/2}^{-1}\frac{s\hat{\Gamma }_{n}^{-1/2}}{\sqrt{n}}\quad,\quad\hat{\beta}+\Phi_{1-\alpha/2}^{-1} \frac{s\hat{\Gamma}_{n}^{-1/2}}{\sqrt{n}}\right] , \] where $\hat{\sigma}^{2}=(n-d-K)s^{2}/n$, and $\Phi_{u}^{-1}=\Phi^{-1}(u)$ denotes the inverse of the Gaussian distribution function. That is, CI$_{0}$ and CI$_{1}$ are formed employing the t-statistic constructed using the homoskedasticity-consistent variance estimators without and with degrees of freedom correction, respectively.
The main findings from the Monte Carlo experiment are presented in Tables (ref) -- (ref) . All results are consistent with the theoretical conclusions presented in the previous section. First, the results for standard Normal and non-Normal errors are qualitatively similar. This indicates that the Gaussian approximation obtained in Theorem (ref) is a good approximation in finite samples, even when $K$ is a nontrivial fraction of the sample size. Second, as expected, a small choice of $K$ leads to important smoothing biases. This affects the finite sample properties of the point estimators as well as the distributional approximations obtained in this paper. In particular, it affects the empirical size of all the confidence intervals. Third, in all cases the results under homoskedasticity or heteroskedasticity in $v_{i}$ are qualitatively similar, showing that our theoretical results provide a good finite sample approximation in both cases, even when $K$ is a nontrivial fraction of the sample size. Fourth, as suggested by Theorem (ref), confidence intervals without degrees of freedom correction (CI$_{0}$) are under-sized, while the analogue confidence intervals with degrees of freedom correction (CI$_{1}$) have close-to-correct empirical size in all cases. This result shows that the degrees of freedom correction is crucial to achieve close-to-correct empirical size when $K/n$ is non-negligible.
In conclusion, we found in our small-scale simulation study that our theoretical results for the partially linear model with possibly many terms provide good approximation in samples of moderate size. In particular, under homoskedasticity of $\varepsilon_{i}$, we showed that confidence intervals constructed using $s^{2}$ exhibit good empirical coverage even when $K/n$ is \textquotedblleft large\textquotedblright. We also confirmed that the Gaussian distributional approximation given in Theorem (ref) represents well the finite sample distribution of $\hat{\beta}$ even when $K/n$ is \textquotedblleft large\textquotedblright.
This paper showed that the many instrument asymptotics and the small bandwidth asymptotics shared a common structure based on a V-statistic, with a remainder term that is asymptotically normal when the number of term diverges to infinity or the bandwidth shrinks to zero. This feature is particularly useful to obtain new results for other semiparametric estimators. In this paper we employ this common structure to derive a new alternative large-sample distributional approximation for a series estimator of the partially linear model, which implied a new (larger) asymptotic variance formula.
Our results apply to a class of semiparametric estimators $\hat{\beta}$ satisfying \[ \sqrt{n}(\hat{\beta}-\beta_{0})=\hat{\Gamma}_{n}^{-1}S_{n}+o_{p}(1), \] where $\hat{\Gamma}_{n}$ and $S_{n}$ take a particular V-stastistic form, as discussed in Section (ref). This class of semiparametric estimators covers several interesting problems, but it is by no means exhaustive. For example, \citet*{Cattaneo-Jansson_2015_wp-Kernels} show that a large class of (kernel-based) semiparametric estimators admit an expansion of the form \[ \sqrt{n}(\hat{\beta}-\beta_{0})=\hat{\Gamma}_{n}^{-1}S_{n}-\mathcal{B} _{n}+o_{p}(1), \] where the bias term $\mathcal{B}_{n}$ is quantitatively and conceptually distinct from the smoothing bias $B_{n}$ described in Section (ref) and, crucially, dominates the quadratic term $U_{n} $ arising from the V-statistic $S_{n}$; that is, $U_{n}=o_{p} (\mathcal{B}_{n})$ in that setting. Nevertheless, the structure we have considered in this paper is useful, providing new results for the partially linear model and a common structure for disparate literatures on many instruments and small bandwidths.