Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
77,450 characters · 10 sections · 0 citation commands
Inference in Nonparametric Series Estimation with Specification Searches for the Number of Series Terms
We consider the following nonparametric regression model
where $\{ y_i, x_i \}_{ i =1}^n$ is i.i.d., $y_i$ is a scalar response variable, $x_i \in \mathcal{X} \subset \mathbb{R}^{d_x}$ is a vector of covariates, and $g_0(x) = E(y_i | x_i = x)$ is the conditional mean function. The theory of estimation and inference is well developed for nonparametric series (sieve) methods in a large body of econometrics and statistics literature. Series estimators have also received attention in applied economics because they have many appealing features, e.g., they can easily impose shape restrictions such as additive separability and monotonicity. Once the basis function is chosen (e.g., polynomial or regression spline series of fixed order), implementation requires a choice of the number of series terms $K = K_n$ where $K$ denotes the order of the polynomials or the number of knots in the splines. However, this often involves some ad-hoc specification searches over $K \in \mathcal{K}_n$. For example, when $x_i \in \mathbb{R}^{d_x}$ is vector valued, researchers often evaluate the different numbers of terms in each dimension separately and construct a set of bases with different powers and cross-products of covariates. Although specification search seems necessary in some cases, it may lead to misleading inference without considering the first-step specification search or series term selection.\footnote{As a referee noted, the bias and MSE of the series estimator depend on not only $K$ but also the specific bases or sieve spaces, e.g., the order of the splines. In this paper, we fix the basis function, and we do not allow searching over the specific bases or sieve spaces.}
Existing theory for the asymptotic normality of t-statistics and valid inference imposes a so-called undersmoothing (i.e., overfitting) condition that is a faster rate of $K$ than the mean-squared error (MSE) optimal convergence rates, and many papers in the literature typically suggest rule-of-thumb rules that give the desired level of undersmoothing. Among many others, Newey (2013) suggested increasing $K$ until the standard errors are large relative to small changes in objects of interest. Newey, Powell, and Vella (1999) suggested using more terms than that chosen by cross-validation. Horowitz and Lee (2012) suggested increasing $K$ until the integrated variance suddenly increases and then adding additional terms.
In this paper, we formally justify these rule-of-thumb methods or “plug-in” methods with undersmoothed $\widehat{K}$ for valid inference in nonparametric series regression. Specifically, we provide pointwise inference for $g_0(x)$ with possibly data-dependent (undersmoothed) $\widehat{K} \in \mathcal{K}_n$, i.e., constructing $100(1-\alpha) \%$ confidence interval (CI),
with an estimator $\widehat{g}_n(K,x)$, variance $\widehat{V}_n(K, x)$ using $K$ series terms, and critical values $\widehat{c}_{1-\alpha}(x)$ from the supremum of the t-statistics. For this result, we first develop a uniform distributional approximation theory of the absolute value of the supremum of the t-statistics over different series terms to construct asymptotically valid confidence intervals, which are uniform in $K \in \mathcal{K}_n$,
The critical values $\widehat{c}_{1-\alpha} (x)$ can be easily implemented using simple simulation or weighted bootstrap methods.
Furthermore, this paper develops the construction of confidence bands for $g_0(x)$ with asymptotically uniform (in $K \in \mathcal{K}_n$) coverage with critical values $\widehat{c}_{1-\alpha} $ chosen to satisfy
Analogous to the pointwise inference in (ref), we can show the validity of confidence bands with the data-dependent $\widehat{K}$. Even in pointwise inference, deriving a uniform asymptotic distribution theory for all sequences of t-statistics over $K \in \mathcal{K}$ may not be possible unless $p = |\mathcal{K}_n|$ is finite. Allowing $p \rightarrow \infty$ as $n \rightarrow \infty$, results in this paper build on coupling inequalities for the supremum of the empirical process developed by Chernozhukov, Chetverikov, and Kato (2014a, 2016) combined with anti-concentration inequality in Chernozhukov, Chetverikov, and Kato (2014b).
We also provide inference methods in a partially linear model setup focusing on the common parametric part. Unlike the nonparametric object of interest that has a slower convergence rate than $n^{1/2}$ (e.g., regression function or regression derivative), the t-statistics for the parametric object of interest are asymptotically equivalent for all sequences of $K$ under standard rate conditions $K/n \rightarrow 0$ as $n \rightarrow \infty$. To account for the dependency of the t-statistics with the different sequences of $K$ in this setup, we consider a faster rate of $K$ that grows as fast as the sample size $n$, as in Cattaneo, Jansson, and Newey (2018a, 2018b), and develop an asymptotic distribution of the t-statistics over $K \in \mathcal{K}_n$. Then, we discuss methods to construct confidence intervals that are similar to the nonparametric regression setup and provide uniform (in $K \in \mathcal{K}_n$) coverage properties.
We investigate finite sample coverage and length properties of the proposed CIs and uniform confidence bands in various simulation setups. As an illustrative example, we revisit nonparametric estimation of labor supply function using the entire individual piecewise-linear budget set as in Blomquist and Newey (2002). Imposing additive separability, which is derived by economic theory, Blomquist and Newey (2002) estimate the conditional mean of labor supply function using series estimation and report wage elasticity of the expected labor supply as well as other welfare measures with various specifications of the different number of series terms.
Several important papers have investigated the asymptotic properties of series (and sieve) estimators, including papers by Andrews (1991a); Eastwood and Gallant (1991); Newey (1997); Chen and Shen (1998); Huang (2003); Chen (2007); Chen and Liao (2014); Chen, Liao, and Sun (2014); Belloni, Chernozhukov, Chetverikov, and Kato (2015); and Chen and Christensen (2015), among many others. This paper extends inference based on the t-statistic under a single sequence of $K$ to the sequences of $K$ over a set $\mathcal{K}_n$ and focuses both on the pointwise and uniform inferences on $g_0(x)$, which is irregular (i.e., slower than a rate of $n^{1/2}$) and a linear functional, under an i.i.d. setup.
The supremum t-statistics have been used as a correction for multiple-testing problems and to construct simultaneous confidence bands, and the importance of multiple-testing problems (data mining or data snooping) has long been noted in various other contexts (see Leamer (1983), White (2000), Romano and Wolf (2005), Hansen (2005)).
There is also a growing literature on data-dependent series term selection and its impact on estimation and inference in econometrics and statistics. Asymptotic optimality results of cross-validation have been developed, e.g., in papers by Li (1987), Andrews (1991b), and Hansen (2015). Horowitz (2014) develops data-driven methods for choosing the sieve dimension in the nonparametric instrumental variables (NPIV) estimation such that resulting NPIV estimators attain the optimal sup-norm or $L^2$ norm rates adaptive to the unknown smoothness of $g_0(x)$. Although we do not pursue adaptive inference in this paper, there is also a large statistical literature on adaptive inference. For example, Gin\'e and Nickl (2010), Chernozhukov, Chetverikov, and Kato (2014b) construct adaptive confidence bands in the density estimation problem (see Gin\'e and Nickl (2015, Section 8) for comprehensive lists of references). However, once data-driven choice is obtained for adaptive estimation (e.g., Lepski (1990)-type procedures), one still requires an undersmoothing condition for inference to eliminate asymptotic bias terms (see Theorem 1 of Gin\'e and Nickl (2010)), and this may result in similar specification search issues when choosing sufficiently “large" $K$ in practice.
We can, in principle, consider kernel-based estimation where several data-dependent bandwidth selections or explicit bias corrections have been proposed.\footnote{See H\"{a}rdle and Linton (1994), Li and Racine (2007) for references. See also Hall and Horowitz (2013), Calonico, Cattaneo, and Farrell (2018), Schennach (2015) and references therein for various recent works on related bias issues and inference for kernel estimators.} However, there exist many examples estimating $g_0(x)$ using (global) series estimation and imposing shape constraints easily (such as additive separability to reduce dimensionality) that are also interested in both pointwise and uniform inference. Given the issues of specification search, our paper is closely related to a recent paper by Armstrong and Koles\'{a}r (2018) which considers a bandwidth snooping adjustment for kernel-based inference.
Unlike kernel-based methods, little is known about the statistical properties of data-dependent selection rules and explicit bias formulas for general series estimation. Zhou, Shen, and Wolfe (1998) and Huang (2003) are two of the few exceptions. A recent paper, Cattaneo, Farrell, and Feng (2019), develops novel explicit asymptotic bias/integrated mean squared error (IMSE) formulas and asymptotic theory of the bias-correction methods for general partitioning-based series estimators. The results in Cattaneo, Farrel, and Feng (2019) can be used as an alternative to the undersmoothing approach to avoid specification search issues.
The remainder of the paper is organized as follows. Section (ref) introduces the basic nonparametric series regression setup and the candidate set $\mathcal{K}_n$. Section (ref) provides the pointwise inference, and Section (ref) provides uniform inference in $x \in \mathcal{X}$. Section (ref) extends our inference methods to the partially linear model setup. Section (ref) summarizes Monte Carlo experiments in various setups, and Section (ref) illustrates an empirical example as in Blomquist and Newey (2002). Then, Section (ref) concludes the paper. Appendix (ref) includes the main proofs, and Appendix (ref) includes figures and tables. Additional supporting lemmas and simulation results are provided in the Online Supplementary Material available at Cambridge Journals Online (journals.cambridge.org/ect).
$||A ||$ denotes the spectral norm, which equals the largest singular value of a matrix $A$, and $\lambda_{min}(A), \lambda_{max}(A)$ denote the minimum and maximum eigenvalues of a symmetric matrix $A$, respectively. $o_p(\cdot)$ and $O_p(\cdot)$ denote the usual stochastic order symbols, $\overset{d}{\longrightarrow}$ denotes convergence in distribution, and $\Rightarrow$ denotes weak convergence. Let $a \wedge b = \min \{ a,b \}, a \vee b = \max \{ a, b \}$ and denote $\lfloor a \rfloor$ as the largest integer less than the real number $a$. For two sequences of positive real numbers $a_n$ and $b_n$, $a_n \lesssim b_n$ denotes $a_n \leq c b_n$ for all $n$ sufficiently large with some constant $c>0$ that is independent of $n$. $ a_n \asymp b_n$ denotes $a_n \lesssim b_n$ and $b_n \lesssim a_n$. Furthermore, $a_n \lesssim_{P} b_n$ denotes $a_n = O_p(b_n)$. For a given random variable $ \{ X_i \} $ and $1 \leq p <\infty$, $L^p(X)$ is the space of all $L^p$-norm bounded functions with $|| f ||_{L^p} = [E || f(X_i)||^p]^{1/p}$, $\ell^{\infty}(X)$ denotes the space of all bounded functions under the sup-norm, and $|| f ||_{\infty} = \sup_{x \in \mathcal{X}} |f(x)|$ for the bounded real-valued functions $f$ on the support $\mathcal{X}$.
We introduce the nonparametric series regression setup in the model (ref). Given a random sample $\{ y_i, x_i \}_{ i =1}^n$, we are interested in inference on the conditional mean $g_0(x) = E(y_i | x_i = x)$ at a particular point $x \in \mathcal{X} \subset \mathbb{R}^{d_x}$ or uniform in $x \in \mathcal{X}$.
Let $\widehat{g}_{n} (K, x) $ be an estimator of $g_0(x)$ using $K = K_n \geq 1$ series terms $P(K, x) = (p_{1}(x), \cdots, p_{K}(x) )'$, which is a vector of basis functions that can change with $n$. Standard examples for the basis functions are power series, Fourier series, orthogonal polynomials, splines and wavelets. The series estimator is then obtained by the least square (LS) estimation of $y_i$ on regressors $P (K, x_i)$
where $P^K = [P_{K1}, \cdots , P_{Kn}]', P_{Ki} \equiv P(K, x_i) = (p_{1}(x_i), p_{2}(x_i), \cdots, p_{K}(x_i) )', Y = (y_1, \cdots y_n)' $. Define the least square residuals as $\widehat{\varepsilon}_{K i} = y_i - P_{K i}'\widehat{\beta}_K$,
and consider the t-statistic
Under standard regularity conditions (discussed in the next section), the t-statistic can be decomposed as follows:
where $Q_K = E(P_{Ki} P_{Ki}')$, $r_n(K, x)= g_0(x) - P(K, x)' \beta_K$, and $\beta_K \equiv E[P_{Ki}P_{Ki}'])^{-1} E[P_{Ki}y_i]$ is the best linear $L^2$ projection coefficient. The first term in the decomposition (ref) converges to a standard normal distribution for the deterministic sequence $K \rightarrow \infty$ as $n\rightarrow \infty$, and the second term does not necessarily converge to 0 due to approximation errors $ r_n(K, x)$. The second term can be ignored with an undersmoothing assumption, and the asymptotic distribution of the t-statistic, $\widehat{T}_n(K, x) \overset{d}{\longrightarrow} N(0,1)$, is well known in the literature (see, for examples, Andrews (1991a), Newey (1997), Belloni et al. (2015), and Chen and Christensen (2015), among many others). Then, the $100(1-\alpha) \%$ confidence interval for $g_0(x)$ can be easily constructed using the normal critical value $z_{1-\alpha/2}$
However, it is not clear whether the conventional CI using normal critical values (ref) has a correct coverage probability with a possibly data-dependent $\widehat{K}$ such as cross-validation or IMSE-optimal selection. First, $T_n({\widehat{K}},\theta_0) \overset{d}{\rightarrow} N(0,1) $ may not hold with a random sequence of $\widehat{K}$, even if we assume the asymptotic bias is negligible. Second, it is well known that some data-dependent rules $\widehat{K}$ do not satisfy the undersmoothing rate conditions, which can lead to a large asymptotic bias and coverage distortion of the standard CI. For example, suppose that the researcher uses $\widehat{K} = \widehat{K}_{\texttt{cv}}$ selected by cross-validation; then, $\widehat{K}_{\texttt{cv}}$ is typically too “small" and violates the undersmoothing assumption needed to ensure the asymptotic normality without bias terms and the valid inference.
As discussed in the introduction, the undersmoothing assumption involves possibly ad-hoc methods to choose series terms $K$ over a candidate set $\mathcal{K}_n$ for a valid inference, and cross-validation methods naturally involve specification search over a set of the different number of series terms.
The following set assumption on $\mathcal{K}_n$ is constructed to allow a broad range of $K$ such that $\mathcal{K}_n$ can allow (unknown) an optimal MSE rate of $K$ as well as an undersmoothing rate that increases faster than the optimal MSE rate.
Here, we consider a possibly growing set of the number of series terms, and a similar assumption is used in the literature, for example, in Newey (1994a, 1994b). Suppose $g_0(x)$ belongs to the H\"older space of smoothness $s >0$, $\Sigma (s, \mathcal{X})$; then, we obtain optimal $L^2$ convergence rates $O_p(n^{-s/(2s + d_x)})$ with $K \asymp n^{d_x/(d_x + 2s)} $. Assumption (ref) allows having optimal $L^2$ rates of $K$ in a large set of classes of functions. By setting $\mathcal{K}_n = [\underline{K}, \overline{K}] \cap \mathbb{N}$, $\overline{K} \asymp n^{ \overline{\phi}}$ and $\underline{K} \asymp n^{ \underline{\phi}}$ with $\overline{\phi} = d_x/(d_x + 2\underline{s}), \underline{\phi} = d_x/(d_x + 2\overline{s})$, Assumption (ref) contains the number of series terms that obtain an optimal $L^2$ rate of convergence for $g_0(x) \in \bigcup_{s \in S} \Sigma (s, \mathcal{X}) $, $S = [\underline{s}, \overline{s}]$. A similar assumption is used in the literature on adaptive inference, although we do not pursue this direction in the current paper.
Assumption (ref) gives flexible choices of $K$, as we only assume the rates of $K$, for example, $\overline{K} = C n^{\overline{\phi}}, \underline{K} = cn^{\underline{\phi}}$, where $c$ and $C$ can be set arbitrarily small or large. We only require rate restrictions uniformly over $K \in \mathcal{K}$ to guarantee the linearization of the t-statistic in (ref) and the rates of the cardinality $p = |\mathcal{K}_{n}|$. Since $K \in \mathcal{K}_n$ is a positive integer and $p \leq \overline{K}$, $p$ is growing at a rate much slower than $n$ under the rate restrictions in Section (ref).
In this section, we focus on pointwise inference for $g_0(x)$. The goal of this section is to provide a uniform distributional approximation theory of $\widehat{T}_n(K, x)$ over a set $\mathcal{K}_n$ and provide uniform (in $K \in \mathcal{K}_{n}$) coverage properties of confidence intervals for $g_0(x)$ in (ref), (ref) with the construction of critical values.
From the decomposition of the t-statistic in (ref), we first consider the (infeasible) test statistic
where $t_n(K, x) =n^{-1/2} \sum_{i=1}^{n} P (K, x)' Q_K^{-1} P_{ K i } \varepsilon_i/ V_{n}(K, x)^{1/2}$ with the series variance $V_{n}(K, x) = P (K, x) ' Q_K^{-1} \Omega_K Q_K^{-1} P(K, x)$, $\Omega_K = E(P_{Ki} P_{Ki}' \varepsilon_{i}^2)$. In general, $t_n(K, x), K\in \mathcal{K}_{n}$ does not have a limiting distribution because it is not asymptotically tight under Assumption (ref) unless $|\mathcal{K}_n|$ is finite or under the restrictive assumption on $\mathcal{K}_n$.\footnote{In an earlier version of the paper, we provide the weak convergence of a series process under the same rates of $K \in \mathcal{K}_n$ and high-level assumptions. This can be viewed as an analogous result in the kernel estimation literature (see Section 2 of Armstrong and Koles\'{a}r (2018) and other references therein).} However, we show below that there exists a sequence of random variables $\max_{ 1 \leq j \leq p} \sum_{i=1}^{n} |Z_{ij}|$ such that $\big|\max_{K \in \mathcal{K}_{n}} |t_n(K, x)| - \max_{ 1 \leq j \leq p} \sum_{i=1}^{n} |Z_{ij}|\big| = O_p(a_n)$ for a sequence of constants $a_n \rightarrow 0$, where $Z_{i} = (Z_{i1}, ..., Z_{ip})^{\prime}$ is a Gaussian random vector in $\mathbb{R}^{p}$ such that $Z_{i} \sim N(0, \frac{1}{n}\Sigma_{n})$ with $(j, l)$ elements of the variance-covariance matrix
$\Omega_{K_j, K_l} = E ( P_{ K_j i } P_{K_l i}' \varepsilon_{i}^2 )$.
By replacing unknown $\Sigma_n, V_n(K,x)$ with consistent estimators $\widehat{\Sigma}_{n}, \widehat{V}_{n}(K,x)$, we show below that we can approximate $\max_{K \in \mathcal{K}_{n}} |\widehat{T}_n(K, x) |$ by $\max_{ 1 \leq j \leq p} \sum_{i=1}^{n} |Z_{ij}|$ and then obtain critical values by using a simulation-based method to provide valid coverage properties in (ref) and (ref). We define $\widehat{c}_{1-\alpha} (x)$ as follows:
where $\widehat{\Sigma}_n$ is a consistent estimator of the variance-covariance matrix $\Sigma_n$ defined in (ref), $\widehat{V}_n(K,x)$ is the simple plug-in estimator for $V_n(K,x)$ as in (ref), and $\widehat{\varepsilon}_{K i} = y_i - P_{K i}'\widehat{\beta}_K, \forall K \in \mathcal{K}_n$. One can compute $\widehat{c}_{1-\alpha} (x)$ by simulating $B$ (typically $B=1000$ or $5000$) i.i.d. random vectors $\widehat{Z}_{i}^b \sim N(0,\frac{1}{n} \widehat{\Sigma}_n)$ and by taking a $(1-\alpha)$ sample quantile of $\{ \max \limits_{1 \leq j \leq p} \sum_{i=1}^{n} |\widehat{Z}_{ij}^{b}|: b=1, \cdots, B \}$. Alternatively, we can use weighted bootstrap methods. See Section (ref) for the implementation and the validity of bootstrap procedures in the construction of confidence bands.
To establish our main results, we impose mild regularity conditions uniform in $K \in \mathcal{K}_n$. For each $K \in \mathcal{K}_n$, define $\zeta_{K} \equiv \sup_{x \in \mathcal{X}} || P(K, x)||$ as the largest normalized length of the regressor vector and $\lambda_{K} \equiv (\lambda_{min} (Q_{K}))^{-1/2}$ for $K \times K$ design matrix $Q_K = E(P_{K i} P_{K i}')$.
Assumptions (ref)(\romannumeral 2) and (ref)(\romannumeral 1) are similar to those imposed in Belloni et al. (2015) and Chen and Christensen (2015), and all the discussions made there also apply here except that we impose rate conditions of $K$ uniformly over $\mathcal{K}_n$. The rate conditions can be replaced by the specific bounds of $\zeta_K, c_K, \ell_K$ with various sieve bases. For example, when $\mathcal{X} = [0,1]^{d_x}$, the probability density of $x_i$ is uniformly bounded above and bounded away from zero, and $g_0(x) \in \Sigma (s, \mathcal{X})$, i.e., the H\"older space of smoothness $s>0$, then $\lambda_{K} \lesssim 1$, $\zeta_{K} \lesssim \sqrt{K}$, $\ell_K c_K \lesssim K^{-(s \wedge s_0)/d_x}$ for regression spline series of order $s_0$, and Assumption (ref)(\romannumeral 1) is satisfied when $ \sqrt{\overline{K} (\log^{3} \overline{K}) /n}(1+ \overline{K}^{1/2} \underline{K}^{-(s \wedge s_0)/d_x}) + \underline{K}^{-(s \wedge s_0)/d_x} \log \overline{K} \rightarrow 0$. Other standard regularity conditions in the literature (e.g., Newey (1997) and Chen (2007)) can also be used here, and the rate condition can be improved with different pointwise linearization and approximation bounds in Huang (2003) for splines and Cattaneo et al. (2019) for partitioning-based estimators.
Assumption (ref)(\romannumeral 2) imposes either the bounded polynomial moment conditions or sub-exponential moments of the regression errors. Assumption (ref)(\romannumeral 3) imposes the consistency of variance estimator $\widehat{V}_n(K,x)$ uniformly in $K \in \mathcal{K}_n$, and this holds under mild regularity conditions (see Lemma 5.1 of Belloni et al. (2015) and Lemma 3.1-3.2 of Chen and Christensen (2015)).
Theorem (ref) provides a uniform coverage property of the confidence interval over $K \in \mathcal{K}_n$ for the regression function $g_0(x)$. Equation (ref) guarantees the asymptotic coverage of CI for data-dependent $\widehat{K} \in \mathcal{K}_n$ with undersmoothing. Note that standard inference methods in the nonparametric regression setup typically consider a singleton set $\mathcal{K}_n = \{ K \} $ with $K \rightarrow \infty$ as $n \rightarrow \infty$. The rate restriction is mild because it only requires $\overline{K}/n^{1-2/q} \rightarrow 0$, up to $\log n$ terms, in case $(a)$ and $\overline{K}/n \rightarrow 0$, up to $\log n$ terms, in case $(b)$ when $\zeta_{K} \lesssim \sqrt{K}$ for splines and wavelet series. Theorem (ref) builds upon a coupling inequality for maxima of sums of random vectors in Chernozhukov, Chetverikov, and Kato (2014a) combined with the anti-concentration inequality in Chernozukhov, Chetverikov, and Kato (2014b).
This section provides construction of uniform confidence bands for $g_0(x)$ (uniform in $K \in \mathcal{K}_{n}$) given in (ref). We define the following empirical process
over $\mathcal{K}_n \times \mathcal{X}$, and we show below that the supremum of the empirical process $\sup_{(K, x) \in \mathcal{K}_{n} \times \mathcal{X}} |\widehat{T}_n(K, x)|$ can be approximated by a sequence of random variables $ \sup_{(K, x) \in \mathcal{K}_{n} \times \mathcal{X}} |Z_n (K,x)|$, where $Z_n (K,x)$ is a tight Gaussian random process in $\ell^{\infty}( \mathcal{K}_{n} \times \mathcal{X})$ with zero mean and covariance function
Although the Gaussian approximation is an important first step, the covariance function (ref) is generally difficult to construct for the purpose of uniform inference. Thus, we employ weighted bootstrap methods similar to Belloni et al. (2015) and show the validity of the bootstrap procedure for uniform confidence bands.
Let $e_1, ..., e_n$ be a sequence of i.i.d. standard exponential random variables that are independent of $X^{n} = \{ x_1, ..., x_n \}$. For $(K,x) \in \mathcal{K}_n \times \mathcal{X}$, we define a (centered) weighted bootstrap process
where $\widehat{g}_{n}^{e}(K, x) = P(K, x)' \widehat{\beta}_{K}^{e}$, and $\widehat{\beta}_{K}^{e}$ is obtained by the following weighted least squares regression
Define the critical value
and we consider confidence bands of the form
To provide the validity of the bootstrap critical values and confidence bands in (ref), we show below that the conditional distribution of $\sup_{(K, x) \in \mathcal{K}_{n} \times \mathcal{X}} |\widehat{T}_n^{e} (K,x)|$ is “close" to the distribution of $ \sup_{(K, x) \in \mathcal{K}_{n} \times \mathcal{X}} |Z_n (K,x)|$ and that of $ \sup_{(K, x) \in \mathcal{K}_{n} \times \mathcal{X}} |\widehat{T}_n (K,x)|$ using coupling inequalities for the supremum of the empirical process and the bootstrap process as in Chernozhukov et al. (2016). Then, similar to Theorem (ref), this gives bounds on the Kolmogorov distance for the distribution functions of $P(\sup_{K \in \mathcal{K}_n, x \in \mathcal{X}} |\widehat{T}_n (K,x) | \leq u)$ and $P(\sup_{K \in \mathcal{K}_n, x \in \mathcal{X}} |\widehat{T}_n^{e} (K,x) | \leq u | X^{n})$.
The following assumptions are used to establish the coverage probability of confidence bands uniformly over $K \in \mathcal{K}_{n}$. Define $\alpha(K,x) \equiv Q_K^{-1/2} P(K,x) /V_n(K,x)^{1/2}$, and \[ \zeta^{L_1} = \max_{K \in \mathcal{K}_n} \sup_{x, x^{\prime}\in \mathcal{X}, x\neq x^{\prime}} \frac{||\alpha(K,x) - \alpha(K,x^{\prime})||}{||x- x^{\prime} ||}, \ \zeta^{L_2} = \sup_{x\in \mathcal{X}} \max_{K, K^{\prime} \in \mathcal{K}_n: K \neq K^{\prime}}\frac{||\alpha(K,x) - \alpha(K^{\prime},x)||}{|K- K^{\prime} |}. \]
For uniform inference, we require similar but slightly stronger conditions compared to Assumption (ref). We also impose mild rate restrictions on $\zeta^{L_1}, \zeta^{L_2}$ and $\max_{K\in\mathcal{K}_n} \zeta_K$ similar to Chernozhukov et al. (2014a) and Belloni et al. (2015).
Theorem (ref) shows the uniform asymptotic coverage property of the confidence bands defined in (ref) uniformly over $K \in \mathcal{K}_n$. Furthermore, it shows a confidence band with possibly data-dependent $\widehat{K} \in \mathcal{K}_{n}$ having an asymptotic coverage of at least $1-\alpha$. The confidence band constructed in (ref) requires a substantially weaker assumption on the undersmoothing similar to Theorem (ref).
In this section we provide inference methods for the partially linear model (PLM) setup. For notational simplicity, we use similar notation as defined in the nonparametric regression setup. Suppose we observe random samples $\{ y_i, w_i , x_i \}_{ i =1}^n$, where $y_i$ is the scalar response variable, $w_i \in \mathcal{W} \subset \mathbb{R}$ is the treatment/policy variable of interest, and $x_i \in \mathcal{X} \subset \mathbb{R}^{d_x}$ is a set of explanatory variables. For simplicity, we shall assume that $w_i$ is a scalar. We consider the model
We are interested in inference on $\theta_0$ after approximating an unknown function $g_0(x)$ by series terms/regressors $p(x_i)$ among a set of potential control variables. Specification searches can be performed for the number of different approximating terms or for the number of covariates in estimating the nonparametric part.
The series estimator $\widehat{\theta}_{n}(K) $ for $\theta_0$ using the first $K = K_n$ terms is obtained by standard LS estimation of $y_i$ on $w_i$ and $P_{Ki} = P(K, x_i)$ and has the usual “partialling out” formula
where $W = (w_1, \cdots, w_n)', M_K = I_K - P^K(P^{K'}P^K)^{-1}P^{K'}, P^K = [P_{K1}, \cdots, P_{Kn}]', Y = (y_1, \cdots, y_n)'$. The asymptotic normality and valid inference for $\widehat{\theta}_{n}(K)$ have been developed in the literature.\footnote{See also Robinson (1988), Linton (1995) and references therein for the results of the kernel estimators.} Donald and Newey (1994) derived the asymptotic normality of $\widehat{\theta}_{n}(K)$ under standard rate conditions $K/n \rightarrow 0$. Belloni, Chernozukhov, and Hansen (2014) analyzed asymptotic normality and uniformly valid inference for the post-double-selection estimator even when $K$ is much larger than $n $ (see also Kozbur (2018)). Recent papers by Cattaneo, Jansson, and Newey (2018a, 2018b) provided a valid approximation theory for $\widehat{\theta}_{n}(K)$ when $K$ grows at the same rate of $n$.
A different approximation theory using a faster rate of $K$ ($K/n \rightarrow c, 0<c <1 $) than the standard rate conditions ($K/n \rightarrow 0$) is particularly useful for our purpose to establish the asymptotic distribution of t-statistics over $K \in \mathcal{K}_n$. From the results in Cattaneo, Jansson, and Newey (2018a), we have the following decomposition:
where $v_i \equiv w_i - g_{w0}(x_i) $, $g_{w0}(x_i ) \equiv E [w_i | x_i]$ and $ \widehat{\Gamma}_n(K) = W'M_K W/n$. For any deterministic sequence $K \rightarrow \infty$ satisfying standard rate conditions $K/n \rightarrow 0$, $\sqrt{n} (\widehat{\theta}_n(K) - \theta_0) $ is asymptotically normal with variance $V = \Gamma^{-1} \Omega \Gamma^{-1}, \Gamma = E[v_i v_i^{\prime}], \Omega = E[v_i v_i^{\prime} \varepsilon_{i}^{2}]$. Unlike the nonparametric object of interest in the fully nonparametric model, where the variance term increases with $K$, $\widehat{\theta}_{n}(K)$ has a parametric ($n^{1/2}$) convergence rate, and $\widehat{\theta}_{n}(K)$ with all different sequences of $K$ are asymptotically equivalent under $K/n \rightarrow 0$.\footnote{This is also related to the well-known results of the two-step semiparametric estimation; the asymptotic variance of two-step semiparametric estimators does not depend on the type of the first-step estimator or smoothing parameter sequences under certain conditions (see Newey (1994b)).} However, under faster rate conditions, $K/n \rightarrow c$ for $0< c <1 $, the second term in (ref) is not negligible and converges to bounded random variables. Cattaneo, Jansson, and Newey (2018a) apply the central limit theorem of degenerate U-statistics for the second term, similar to the many instrument asymptotics analyzed in Chao, Swanson, Hausman, Newey, and Woutersen (2012). Then, the limiting normal distribution has a larger variance than the standard first-order asymptotic variance, and the adjusted variances generally depend on the number of terms $K$ such that we can provide an asymptotic distribution of the t-statistics with the different sequence of $K$ over $\mathcal{K}_n$.
The following assumption on $\mathcal{K}_n$ is considered, and we impose the regularity conditions that are used in Cattaneo, Jansson, and Newey (2018a, Assumption PLM) uniformly over $K \in \mathcal{K}_n$.
Assumption (ref) does not require $K/n \rightarrow 0$ which is required to obtain asymptotic normality in the literature (e.g., Donald and Newey (1994)). Similar to Assumption (ref)(\romannumeral 3) in the nonparametric setup, Assumption (ref)(\romannumeral 4) holds for the polynomials and spline basis. For example, (ref)(\romannumeral 4) holds with $\gamma_g =s_g/d_x, \gamma_{g_w} = s_w/d_x $ when $\mathcal{X}$ is compact and when the unknown functions $g_0(x)$ and $g_{w0}(x)$ have $s_g$ and $s_w$ continuous derivates, respectively.
Under Assumptions (ref), (ref) and undersmoothing condition ($n\overline{K}^{-2(\gamma_g+ \gamma_{g_w})} \rightarrow 0$), we have a joint asymptotic distribution of the t-statistics $T_n (K, \theta) = \sqrt{n} V_{n}(K)^{-1/2}(\widehat{\theta}_n(K) - \theta_0)$ over $K \in \mathcal{K}_n$:
where
and the variance-covariance matrix $\Sigma$ with $(l, l^{\prime})$ element
for $l, l^{\prime} = 1, ..., p.$ Then, we can similarly define critical values as in (ref) to construct confidence intervals for $\theta_0$ uniform in $K \in \mathcal{K}_{n}$ analogous to the nonparametric setup. Let
where $\widehat{\Sigma}_{n}$ is a consistent estimator for unknown $\Sigma$ defined in (ref).
Theorem (ref) is the main result for the partially linear model setup and provides the asymptotic coverage results of the CIs uniform in $K \in \mathcal{K}_n$ analogous to the nonparametric setup in Section (ref).
This section investigates the small sample performance of the proposed inference methods. We report the empirical coverage and the average length of the confidence intervals/confidence bands considered in Sections (ref) and (ref) with various simulation setups.
We consider the following data generating process:
where $\Phi(\cdot)$ is the standard normal cumulative distribution function needed to ensure compact support, and $\sigma^2(x_{i}^{*})= ((1+ 2x_{i}^{*})/2)^{2}$ (heteroskedastic). We investigate the following three functions for $g(x)$: $g_1(x) = \ln (|6x-3| + 1)sgn(x-1/2)$, $g_2(x) = \frac{\sin (7\pi x/2)}{1+ 2x^2(sgn(x)+1)}$, and $g_3(x) = x-1/2 + 5 \phi(10(x-1/2))$, where $\phi(\cdot)$ is the standard normal probability density function, and $sgn(\cdot)$ is the sign function. $g_1(x)$ is used in Newey and Powell (2003), as well as Chen and Christensen (2018). $g_2(x)$ and $g_3(x)$ are rescaled versions used in Hall and Horowitz (2013). See Figure (ref) for the shapes of all three functions on $[0,1]$. For all simulation results below, we generate 2000 simulation replications for each design with a sample size $n = 200$.
Results for quadratic splines with evenly placed knots are reported where the number of knots $K$ are selected among $\mathcal{K}_n = \{ 6, 7, ..., 12\}$ by setting $\underline{K} = 2n^{1/5}$ and $\overline{K} = 2n^{1/3}$ rounded up to the nearest integer. Then, we calculate a pointwise coverage rate (COV) and the average length (AL) of various 95% nominal CIs, as well as analogous uniform CBs for the grid points of $x$ on the support $\mathcal{X} = [0.05,0.95]$. To calculate critical values, 1000 additional Monte Carlo or bootstrap replications are performed on each simulation iteration. In addition, we investigate results for homoskedastic errors ($\sigma^2(x_{i}^{*})= 1$), different sample sizes $n = \{ 100, 500 \}$, polynomial regressions, and different specifications as in Cattaneo and Farrell (2013) with multivariate and non-normal regressors; however, the results show qualitatively similar patterns and hence are not reported here for brevity. Additional simulation results are reported in the Online Supplementary Material.
Table (ref) reports the nominal 95% coverage of the following pointwise CIs at $x=0.2, 0.5, 0.8, 0.9$: (1) the standard CI in (ref) with $\widehat{K}_{\texttt{cv}} \in \mathcal{K}_n$ selected to minimize the leave-one-out cross-validation; (2) robust CI in (ref) with $\widehat{K}_{\texttt{cv}}$ using the critical value $\widehat{c}_{1-\alpha} (x)$; (3) robust CI using $\widehat{K}_{\texttt{cv+}} = \widehat{K}_{\texttt{cv}} + 2$. Analogous uniform inference results for CBs are also reported. The critical values, $\widehat{c}_{1-\alpha} (x)$ and $\widehat{c}_{1-\alpha}$ are constructed using the Monte Carlo methods and weighted bootstrap method, respectively.
Overall, we find that the coverage of the standard CI with $\widehat{K}_{\texttt{cv}}$ is far less than 95% over the support although it has the shortest length. However, the coverage of robust CIs based on $\widehat{K}_{\texttt{cv}}$ or $\widehat{K}_{\texttt{cv+}}$ with $\widehat{c}_{1-\alpha} (x)$ is close to or above 95% and performs well across the different simulation designs, and this is consistent with theoretical results in Theorem (ref). Using the undersmoothed $\widehat{K}_{\texttt{cv+}}$ (using more terms than the cross-validation) seems to work quite well at most points and for highly nonlinear designs where there exists relatively large bias, e.g., Model 3 $(g_{3}(x))$ at $x=0.5$.\footnote{The possibly poor coverage property of the standard kernel-based CIs for $g_3(x)$ at the single peak ($x= 0.5$) was also described in Hall and Horowitz (2013, Figure 3).} Uniform coverage rates of confidence bands with selected $K$ seem conservative, and this is due to the large critical values based on weighted bootstrap methods to be uniform in both $K \in \mathcal{K}_n$ and $x \in \mathcal{X}$, including boundary points.
In this section, we illustrate inference procedures by revisiting Blomquist and Newey (2002). Understanding how tax policy affects individual labor supply has been a central issue in labor economics (see Hausman (1985) and Blundell and MaCurdy (1999), among many others). Blomquist and Newey (2002) estimate the conditional mean of hours of work given the individual nonlinear budget sets using nonparametric series estimation. They also estimate the wage elasticity of the expected labor supply and find evidence of possible misspecification of the usual parametric model such as maximum likelihood estimation (MLE).
Specifically, Blomquist and Newey (2002) consider the following model by exploiting an additive structure from the utility maximization with piecewise linear budget sets:
where $h_i$ is the hours worked of the $i$th individual and $x_i = (y_1, \cdots, y_J, w_1, \cdots, w_J, \ell_1, \cdots, \ell_J, J)$ is the budget set, which can be represented by the intercept $y_j$ (non-labor income), slope $w_j$ (marginal wage rates) and the end point $\ell_j$ of the $j$th segment in a piecewise linear budget with $J$ segments. Equation (ref) for the conditional mean function follows from Theorem 2.1 of Blomquist and Newey (2002), and this additive structure substantially reduces the dimensionality issues. To approximate $g(x)$, they consider the power series, $p_k(x) = (y_J^{p_1(k)} w_J^{q_1(k)}, \sum_{j=1}^{J-1} \ell_j^{m(k)} (y_j^{p_2(k)} w_j^{q_2(k)} - y_{j+1}^{p_2(k)} w_{j+1}^{q_2(k)})),$ $ p_2(k) + q_2(k) \geq 1$.
From the Swedish “Level of Living" survey in 1973, 1980 and 1990, they pool the data from three waves and use the data for married or cohabiting men of ages 20-60. Changes in the tax system over three different time periods give a large variation in the budget sets. The sample size is $n = 2321$. See Section 5 of Blomquist and Newey (2002) for more detailed descriptions. They estimate the wage elasticity of the expected labor supply
which is the regression derivative of $g(x)$ evaluated at the mean of the net wage rates $\bar{w}$, virtual income $\bar{y}$ and level of hours $\bar{h}$.
Table (ref) is the same table as in Blomquist and Newey (2002, Table 1). They report estimates $\widehat{E}_w$ and standard errors $SE_{\widehat{E}_w}$ with a different number of series terms by adding additional series terms. For example, the estimates in the second row use the term in the first row $(1, y_J, w_J)$ with the additional terms $(\Delta y, \Delta w)$. Here, $\ell^m \Delta y^p w^q$ denotes approximating the term $\sum_j \ell_j^{m} (y_j^{p} w_j^{q} - y_{j+1}^{p} w_{j+1}^{q})$. Blomquist and Newey (2002) also report cross-validation criteria, $CV$, for each specification. In their formula, series terms are chosen to maximize $CV$, which minimizes the asymptotic MSE. In addition to their original table, we add the standard 95% CI for each specification, i.e., $CI (K) = \widehat{E}_w (K) \pm 1.96 SE_{\widehat{E}_w} (K)$. In Table (ref), it is ambiguous as to which large model ($K$) can be used for the inference, and we do not have compelling data-dependent methods for selecting one of the large $K$ for the confidence interval to be reported. Here we want to construct CIs that are robust to specification searches.
Figure (ref) displays pointwise 95% uniform CIs for $K_m \in \{ K_1, K_2, \cdots, K_{11}\}$, where $K_m$ corresponds to each specification in Table (ref) with increasing order of series terms, along with the point estimates and standard 95% confidence interval.\footnote{It is straightforward to construct $\widehat{c}_{1-\alpha}(x)$ using the covariance structure under the homoskedastic error and it only requires estimated variances for different $K\in \mathcal{K}_n$ that are already reported in the table of Blomquist and Newey (2002). Based on 100,000 simulation repetitions, we have $\widehat{c}_{1-\alpha}(x) = 2.503$.} From Figure (ref), we reject a zero wage elasticity of the labor supply for almost all models except $\overline{K}$. Table (ref) also reports robust confidence intervals $CI_{\widehat{E}_w}^{\texttt{sup}} (K) = \widehat{E}_w (K) \pm \widehat{c}_{1-\alpha}(x) SE_{\widehat{E}_w} (K)$ with possibly data-dependent $\widehat{K}$ justified by Theorem (ref) (eq (ref)). Note that cross-validation chooses $ \widehat{K}_{\texttt{cv}} = K_5$, and the standard CI with $\widehat{K}_{\texttt{cv}}$ is $[0.0247, 0.0839]$ and the robust CI is $[0.0165, 0.0921]$. Using $\widehat{K}_{\texttt{cv+}} = K_6$ or $\widehat{K}_{\texttt{cv++}} = K_7$ widens the standard CI, and the robust CIs are $CI_{\widehat{E}_w}^{\texttt{sup}} (\widehat{K}_{\texttt{cv+}})= [0.0166, 0.1152], CI_{\widehat{E}_w}^{\texttt{sup}} (\widehat{K}_{\texttt{cv++}}) = [0.0070, 0.1186]$.
This paper considers nonparametric inference methods given specification searches over different numbers of series terms in the nonparametric series regression model. We provide methods of constructing uniform CIs and confidence bands by adjusting the conventional normal critical value to the critical value based on the supremum of the t-statistics. The critical values can be constructed using simple Monte Carlo simulation or weighted bootstrap methods. Then, we provide an extension of the proposed CIs in the partially linear model setup. Finally, we investigate the finite sample properties of the proposed methods and illustrate uniform CIs in an empirical example of Blomquist and Newey (2002).
While beyond the scope of this paper, there are some potential directions to extend the results established here. First, investigating the coverage property of CIs with data-dependent $\widehat{K}$ using bias-corrected methods is of interest. In particular, it would be of interest to analyze the bias-corrected CI and confidence bands using cross-validation methods combined with the recent results established in Cattaneo, Farrell, and Feng (2019). Second, an extension of the current theory for quantile regression (e.g., Belloni, Chernozhukov, Chetverikov, and Fern\'{a}ndez-Val (2019)) or the nonparametric IV setup would be desirable. In the NPIV setup, for example, one can consider pointwise CIs (or uniform confidence bands) that are uniform in pairs of $ (K_n, J_n) \in \mathcal{K}_n \times \mathcal{J}_n$ with an additional dimension of the instrument sieve and the number of instruments $J =J_n$. This is a difficult problem, and it would require a distinct theory to address the ill-posed inverse problem as well as two-dimensional choices. We leave these topics for future research.