Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
108,892 characters · 21 sections · 81 citation commands
Adaptive Estimation and Uniform Confidence Bands for Nonparametric Structural Functions and Elasticities
\defaultbibliography{refecon} \defaultbibliographystyle{chicago}
With easier access to large data sets, there is increasing interest in estimating flexible, nonparametric structural functions and their derivatives, such as elasticities or other marginal effects. In many applications, the structural function $h_0$ is identified by a conditional moment restriction
where $Y$ (a scalar) and/or some elements of $X$ (a vector) are endogenous, $W$ is a vector of instrumental variables, and the conditional distribution of $(X,Y)$ given $W$ is otherwise unspecified. Examples include consumer demand blundell2007semi,BHP, demand for differentiated products BerryHaile2014,Compiani, and international trade ACD,AAG.\footnote{Other applications include causal inference WangZhiTT2018 and reinforcement learning chen2022well,grettonRL2021. Model ((ref)) also nests nonparametric regression when $W=X$, in which case $h_0$ is the conditional mean of $Y$ given $X$.} Uniform confidence bands (UCBs) are very helpful for inferring the true shape, slope, or curvature of $h_0$, as they graphically convey sampling uncertainty about the estimated structural function and its derivatives.
In applications involving policy counterfactuals, researchers care about estimating and constructing UCBs for $h_0$ or its derivatives. For instance, AAG (AAG, AAG hereafter) derive ((ref)) via a semiparametric gravity equation for the intensive margin of firm exports in a monopolistic competition model based on Melitz2003. In that context, the derivative of $h_0$ is the elasticity of the intensive margin of firm-level exports to changes in bilateral trade costs. Moreover, Compiani performs policy experiments using nonparametric estimates of price elasticities in differentiated product demand models.
As is the case for almost all nonparametric and machine learning (ML) methods, researchers must choose tuning parameters---such as bandwidths, sieve dimensions, or penalty parameters---when estimating or performing inference on $h_0$ and its derivatives. Poor choice of tuning parameters can lead to estimators that converge unnecessarily slowly and confidence bands with poor coverage. But “good” choices of tuning parameters typically require knowledge of key model regularities, such as the smoothness of $h_0$ and the strength of the instruments, which are unknown ex ante. It is therefore important to have data-driven methods that adapt to unknown model regularities and yield estimators and confidence bands with desirable properties. Data-driven methods for choosing tuning parameters also help to improve the transparency of nonparametric and ML methods, removing a degree of freedom with which the researcher can manipulate results. Unfortunately, popular methods for choosing tuning parameters for nonparametric regression, such as standard cross validation, may not be valid in models with endogeneity---see Section (ref).
In this paper, we propose simple, data-driven procedures for choosing tuning parameters for estimating and constructing UCBs for $h_0$ and its derivatives. Our methods are developed for the popular class of sieve nonparametric IV estimators.\footnote{See ai2003efficient, newey2003instrumental, blundell2007semi, and horowitz2011.} That is, $h_0$ is approximated by a linear combination of several basis functions (e.g., B-splines), with the coefficients estimated by Two Stage Least Squares (TSLS) regression of $Y$ on the basis functions of $X$, using functions of $W$ as instruments (see Section (ref) for a detailed description). The key tuning parameter to be chosen by a researcher is the number of basis functions, say $J$, used to approximate $h_0$. If $J$ is too small, then estimators may be badly biased and UCBs may under-cover. But if $J$ is too large, estimators may be very noisy and UCBs may be uninformatively wide. Before precisely stating our theoretical results in Section (ref), we describe our methods and their practical importance.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Our Methods and the Practical Implications.}
Our first contribution is a data-driven choice of sieve dimension, which we denote by $\tilde J$. This choice is simple to compute. Under suitable regularity conditions, we show that sieve estimators implemented with $\tilde J$, which we denote $\hat h_{\tilde J}$, converge at the fastest possible (i.e., minimax) rate in sup-norm.\footnote{We focus on the sup-norm rather than $L^2$ norm (i.e., mean-square error) primarily because our objective is to construct UCBs for $h_0$ and its derivatives. The sup-norm is essential for this purpose, as we require the entire function (or its derivatives) to lie inside the bands with desired coverage probability. The sup-norm also provides a stronger, more informative sense in which the estimator is converging as it measures the maximal, rather than average, error over the support of $X$.} That is, the maximum error over the support of $X$, namely \[ \sup_x |\hat h_{\tilde J}(x) - h_0(x)|, \] vanishes as fast as possible---among all estimators of $h_0$---as the sample size increases, uniformly over a class of data-generating processes (DGPs), for both nonparametric IV and nonparametric regression models. Formally, we refer to $\tilde J$ as sup-norm rate-adaptive: it adapts to features of the DGP that are unknown ex ante, such as the smoothness of $h_0$ and strength of the instruments, so that the resulting estimator $\hat h_{\tilde J}$ converges as fast as possible in sup-norm. We further show that the same data-driven choice $\tilde J$ is sup-norm rate-adaptive for estimating derivatives of $h_0$ as well.\footnote{This is in contrast to kernel estimation, in which different bandwidths must be used for rate-adaptive estimation of a function and its derivatives.} Hence, $\tilde J$ should be very useful for researchers interested in estimating elasticities or other marginal effects. We illustrate this usefulness in our empirical application revisiting AAG, where we use $\tilde J$ to estimate the elasticity of the intensive margin of firm-level exports from aggregate bilateral trade data. We also demonstrate the good performance of $\tilde J$ across a variety of simulation designs for both nonparametric IV estimation and nonparametric regression.
Our second main contribution is a data-driven approach to constructing UCBs for $h_0$ and its derivatives. The term “uniform” indicates that the entire function lies within the bands with desired asymptotic coverage probability. The UCBs for $h_0$ and its derivatives are also simple to compute and have strong theoretical justification. They are honest in the sense that they guarantee coverage for $h_0$ and its derivatives uniformly over a generic class of DGPs, and adaptive in the sense that they contract at, or within a logarithmic factor of, the minimax rate. As such, they provide efficiency improvements relative to UCBs based on the usual approach of undersmoothing, in which a sub-optimally large $J$ is chosen in the hope that bias is negligible relative to sampling variation. Of course, in empirical work, a researcher does not know the true function, and therefore doesn't know which $J$ is truly large enough that sampling uncertainty dominates bias.
Our UCBs for $h_0$ and its derivatives are useful for inferring the true shape of the structural function and its derivatives. They complement existing approaches for testing shape restrictions, as they allow the researcher to read off the shape of the function without imposing a specific null (e.g. monotone increasing) a priori. In our empirical application to AAG we construct UCBs for the elasticity of the intensive margin of firm exports. As emphasized by AAG, this is an important, policy-relevant function yet its shape is not restricted by theory in a nonparametric setting. Our UCBs exclude constant functions and downwards-sloping functions. Hence, they provide evidence against the Pareto specification for unobserved firm productivity used by Chaney, which leads to a constant elasticity, as well as other parameterizations used, e.g., by EKK, HMT, and MelitzRedding, for which the elasticity is downwards-sloping. Empirically-calibrated simulation studies based on the models of Chaney and HMT demonstrate valid coverage of our UCBs for $h_0$ and its derivatives and efficiency improvements relative to undersmoothing.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Related Literature and our Theoretical Contributions.}
Early work on nonparametric IV estimation includes newey2003instrumental, hall2005nonparametric, blundell2007semi, darolles2011nonparametric, horowitz2011 and others.
We complement prior work by horowitz2014adaptive for near-adaptive estimation of $h_0$ in $L^2$ norm, breunig2016adaptive for near-adaptive estimation of linear functionals of $h_0$, and breunigchen2021 for adaptive estimation of quadratic functionals of $h_0$. Our procedure builds on the bootstrap-based implementation of Lepski's method of chernozhukov2014anti for kernel density estimation and spokoiny2019 for linear regression with Gaussian errors. But our procedure does not follow easily from theirs due to several challenges present in the conditional moment restriction ((ref)), in which $h_0$ is identified by $\mathbb{E}[Y|W] = \mathbb{E}[h_0(X)|W]$ (a.s.). The degree of difficulty of inverting $\mathbb{E}[h_0(X) |W]$ to recover $h_0$ is a nonparametric notion of instrument strength and plays an important role in determining minimax rates for estimators of $h_0$ and its derivatives.\footnote{See hall2005nonparametric, chen2011rate, and chen2018optimal for minimax rates for nonparametric IV estimation. When the conditional density of $X$ given $W$ is continuous, these rates are slower than the corresponding rates for nonparametric regression.} While adaptive procedures for nonparametric density estimation or regression deal only with unknown smoothness of the estimand, our procedures must also deal with the unknown degree of difficulty of the inversion problem. The literature has typically classified the difficulty of the inversion problem into “mild” and “severe” regimes. Minimax rates in the mild regime are achieved by a choice of sieve dimension that balances bias and sampling uncertainty, much like standard nonparametric problems. But minimax rates in the severe regime are obtained by a bias-dominating choice of sieve dimension. Our procedure for data-driven choice of sieve dimension delivers the minimax sup-norm rate for $h_0$ and its derivatives across the whole spectrum of models, from nonparametric regression to nonparametric IV models in the severe regime.
Our procedure improves significantly on and supersedes a modified Lepski procedure from Section 3 of chen2015arxiv on sup-norm rate-adaptive estimation of ((ref)). Ours uses a multiplier bootstrap to avoid selection of several constants and performs much better in practice. Moreover, our rate-adaptivity guarantees encompass nonparametric regression and nonparametric IV in both mild and severe regimes.
Recent work on (non data-driven) UCBs for $h_0$ and functionals thereof via undersmoothing includes horowitz2012uniform, chen2018optimal and babii2020. Our UCBs build on prior work on honest, adaptive UCBs for nonparametric density estimation gine2010confidence,chernozhukov2014anti and Gaussian white noise models bull2012honest,gine2016mathematical. But none of these works allows for nonparametric models with endogeneity, and our procedures do not follow easily from these existing methods due to the above-mentioned challenges present in model ((ref)). Our UCBs for $h_0$ and its derivatives apply to nonparametric regression with non-Gaussian, heteroskedastic errors as a special case, which appears to be a new contribution.
Finally, our work also compliments several recent papers on (non data-driven) estimation and inference for nonparametric IV models with shape constraints; see for example BHP, chetverikov2017, FreybergerReeves and CNS. These works all assume a deterministic sequence of tuning parameters satisfying regularity conditions that depend on unknown model features such as the smoothness of $h_0$ and instrument strength. An exception is breunig2020adaptive who study $L^2$ rate-adaptive testing of a specific null hypothesis (e.g., monotone increasing, or a parametric functional form). Our approach is conceptually different from theirs: our UCBs graphically convey sampling uncertainty about an estimate of $h_0$ and its derivatives. Hence, our UCBs are very useful for inferring the true shape of $h_0$ in situations---such as our trade application---where there are no specific prior shape restrictions suggested by economic theory.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Outline.} Section (ref) introduces our methods. Section (ref) presents the application to international trade. Section (ref) contains the main theoretical results. Section (ref) provides additional simulation results for difficult designs. Section (ref) presents extensions to additive and partially linear models, and Section (ref) concludes. Appendix (ref) presents a simplified version of our procedures for nonparametric regression. Appendix (ref) provides additional details for the trade application and simulations. In the online supplement, Appendix (ref) presents additional simulations to an empirically calibrated Engel curve design, Appendix (ref) gives details on basis functions and nonparametric function classes, and Appendix (ref) contains technical results and proofs.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Notation.} Let $\mc X$ be the support of $X$, $d$ the dimension of $X$, and $L^2_X$ and $L^2_W$ the space of functions of $X$ and $W$ with finite second moments. Let $\|h\|_{\infty} :=\sup_{x\in \mc X} |h(x)|$ be the sup-norm of $h : \mc X \to \mb R$. Let $\mb N$ be the set of integers and $\mb N_0 := \mb N \cup \{0\}$ the non-negative integers. Let $\lceil a \rceil = \min\{n \in \mathbb N : n \geq a\}$ and $\lfloor a \rfloor = \max\{n \in \mathbb N_0 : n < a\}$. For a multi-index $a=(a_1,...,a_d) \in (\mb N_0)^d$ with order $|a| = \sum_{i=1}^d a_i$, the $a$-derivative of $h$ is defined as \[ \partial^a h(x) = \frac{\partial^{|a|} h(x)}{\partial^{a_1} x_1 \ldots \partial^{a_d} x_d} \,. \] Let $A^-$ denote the generalized (or Moore--Penrose) inverse of a matrix $A$ and $A^{-1/2}$ the inverse of the positive-definite square root of $A$.
We begin in Section (ref) by briefly reviewing sieve nonparametric IV estimation and UCBs with a deterministic sieve dimension. Section (ref) explains why standard cross validation for regression fails in models with endogeneity. Section (ref) presents our data-driven choice of sieve dimension and Section (ref) presents our data-driven UCBs. These methods extend naturally to partially linear and partially additive models (see Section (ref)). Both procedures apply to nonparametric regression as well (see Appendix (ref)).
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Estimators.} Consider approximating $h_0$ by a linear combination of $J$ basis functions:
where $\psi^J(x) = (\psi_{J1}(x),\ldots,\psi_{JJ}(x))'$ is a vector of basis functions and $c_J = (c_{J1},\ldots,c_{JJ})'$ is a vector of coefficients. Combining ((ref)) and ((ref)), we obtain \[ Y = (\psi^J(X))'c_J + \mathrm{bias}_J + u\,,~~~\mathbb{E}[u|W] = 0\,, \] where $u = Y - h_0(X)$ and $\mathrm{bias}_J = h_0(X) - (\psi^J(X))'c_J$. Provided the bias term is “small” relative to $u$ in an appropriate sense, we have an approximate linear IV model where $\psi^J(X)$ is a $J\times 1$ vector of “endogenous variables” and $c_J$ is a vector of unknown “parameters”. One can then estimate $c_J$ using TSLS or GMM using a $K\times 1$ vector of basis functions $b^K(W)= (b_{K1}(W),\ldots,b_{KK}(W))'$ of $W$ as instruments. Evidently, $K\geq J$ is necessary to estimate $c_J$.
Given data $(X_i,Y_i,W_i)_{i=1}^n$, the TSLS estimator of $c_J$ is simply \[ \hat c_J = \left(\mf \Psi_J' \mf P_K^{\phantom \prime} \mf \Psi_J^{\phantom \prime} \right)^{-} \mf \Psi_J' \mf P_K^{\phantom \prime} \mf Y \,, \] where $\mf \Psi_J = (\psi^J({X_1}),\ldots,\psi^J({X_n}))'$ and $\mf B_K = (b^K({W_1}),\ldots,b^K({W_n}))'$ are $n \times J$ and $n \times K$ matrices, $\mf P_K = \mf B_K^{\phantom \prime} (\mf B_K' \mf B_K^{\phantom \prime})^{-} \mf B_K^{\prime}$ is the projection matrix onto the instrument space, and $\mf Y = (Y_1,\ldots,Y_n)'$ is a $n \times 1$ vector. Estimators of $h_0$ and its derivative $\partial^a h_0$ are given by \[ \hat h_J(x) = (\psi^J(x))' \hat c_J \,,~~~~\mbox{and}~~~~\partial^a \hat h_J(x) =(\partial^a \psi^J(x))' \hat c_J \,, \] where $\partial^a \psi^J(x) = (\partial^a \psi_{J1}(x),\ldots,\partial^a \psi_{JJ}(x))'$.
Sieve Bases. Many linear sieves, such as polynomial splines, B-splines, wavelets, Fourier series, and various polynomials, can be used as the instrument basis $\{b_{Kk}\}_{k=1}^K$. However, only B-splines and Cohen--Daubechies--Vial (CDV) wavelet bases for $\{\psi_{Jj}\}_{j=1}^J$ have been shown to achieve the optimal minimax sup-norm rates under a suitable choice of $J$ chen2018optimal.\footnote{Bases for $h_0$ must have bounded Lebesgue constant to attain the minimax sup-norm rate for nonparametric regression (see, e.g., belloni2015some and chen2015optimal). B-splines and CDV wavelets have this property. Bases without this property, such as polynomials and Fourier series, cannot attain the minimax sup-norm rate and hence cannot lead to sup-norm rate-adaptive estimators or UCBs.} As our objective is to have estimators that converge as fast as possible in sup-norm---which is essential for constructing UCBs that are as narrow and informative as possible---we restrict attention to B-splines and CDV wavelets for $\{\psi_{Jj}\}_{j=1}^J$ in our theory that follows. Moreover, since B-splines are easy to compute, much less collinear than polynomials and polynomial splines, and available in standard software packages, we confine our presentation to B-spline bases for both $\{\psi_{Jj}\}_{j=1}^J$ and $\{b_{Kk}\}_{k=1}^K$ in the main text.
Key tuning parameter $J$. Based on simulations and theoretical studies in blundell2007semi, chen2018optimal and others, the performance of the sieve TSLS estimator for $h_0$ is sensitive to the choice of $J$ and not sensitive to $K$ as long as $K\geq J$. We introduce a data-driven method for choosing $J$ in Section (ref). The choice of $K$ is pinned down by $J$ in our procedure, so we write $K(J)\geq J$, $b^{K(J)}(W)$, $\mf B_{K(J)}$ and $\mf P_{K(J)}$ in what follows. Let $\mf M_J = (\mf \Psi_J' \mf P_{K(J)}^{\phantom \prime} \mf \Psi_J^{\phantom \prime} )^{-} \mf \Psi_J' \mf P_{K(J)}^{\phantom \prime}$ be a $J \times n$ matrix. We can equivalently write
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{“Undersmoothed” UCBs.} We now review the usual approach of constructing “undersmoothed” UCBs for $h_0$ and its derivatives based on a deterministic $J$. Let $\hat{\mf u}_J = (\hat u_{1,J},\ldots,\hat u_{n,J})'$ denote the $n \times 1$ vector of residuals whose $i$th element is $\hat u_{i,J} = Y_i - \hat h_J(X_i)$. Then $\hat h_J (x) - h_0 (x)$ and $\partial^a \hat h_J (x) - \partial^a h_0 (x)$ can be estimated by
and their variances can be estimated by
where $\widehat{\mf U}_{J,J}$ is a $n\times n$ diagonal matrix whose $i$th diagonal entry is $\hat u_{i,J} \hat u_{i,J}$.
Let $\hat{\mf u}_J^* = (\hat u_{1,J}\varpi_1,\ldots,\hat u_{n,J}\varpi_n)'$ denote a multiplier bootstrap version of $\hat{\mf u}_J$, where $(\varpi_i)_{i=1}^n$ are IID $N(0,1)$ draws independent of the data. Then
are bootstrap versions of $D_J (x)$ and $D_J^{a}(x)$. For each independent draw of $(\varpi_i)_{i=1}^n$, compute the sup $t$-statistics:
Let $z_{1-\alpha,J}^*$ and $z_{1-\alpha,J}^{a*}$ denote the $(1-\alpha )$ quantile of these sup statistics across a large number (say $1000$) independent draws of $(\varpi_i)_{i=1}^n$. chen2018optimal construct 100$(1-\alpha)$% UCBs for $h_0$ and $\partial^a h_0$ as follows: \[ C_{n,J}(x) = \bigg[ \hat{h}_{J}(x) - z_{1-\alpha,J}^* \hat \sigma_{J}(x) , ~ \hat{h}_{J}(x) + z_{1-\alpha,J}^* \hat \sigma_{J}(x) \bigg] \,, \] \[ C_{n,J}^a(x) = \bigg[ \partial^a \hat{h}_{J}(x) - z_{1-\alpha,J}^{a*} \hat \sigma_{J}^a(x) ,~ \partial^a \hat{h}_{J}(x) + z_{1-\alpha,J}^{a*} \hat \sigma_{J}^a(x) \bigg] \,. \] The above UCBs are theoretically justified provided $J$ increases faster than the oracle $J_0$ (the optimal sieve dimension for estimating $h_0$ or its derivatives in sup-norm), so that the bias is of smaller order than sampling uncertainty. Unfortunately, $J_0$ is unknown in practice since it depends on the unknown smoothness of $h_0$ and other unknown model regularities of ((ref)). This motivates us to propose the new data-driven UCBs in Section (ref).
We briefly explain why the usual approach of cross validation (CV) for regression is not a valid method for choosing $J$ in models with endogeneity. Consider the standard CV criterion
where $n$ is the sample size and $\hat h_{-i,J}$ denotes version of $\hat h_J$ computed from a sub-sample that excludes the $i$th observation. Let $u_i = Y_i - h_0(X_i)$. We may then expand ((ref)) as \[ \mathrm{CV}(J) = \frac{1}{n} \sum_{i=1}^n (h_0(X_i) - \hat h_{-i, J}(X_i))^2 + \frac{1}{n} \sum_{i=1}^n u_i^2 + \frac{2}{n} \sum_{i=1}^n u_i (h_0(X_i) - \hat h_{-i, J}(X_i)). \] The first term in the expansion is an estimate of the MSE $\mathbb{E}[ (h_0(X) - \hat h_J(X))^2]$ of $\hat h_J$ and the second term is independent of $J$. The third term is an estimate of $\mathbb{E}[u(h_0(X) - \hat h_J(X))]$. This term is asymptotically negligible without endogeneity (i.e., when $\mathbb{E}[u|X] = 0$) as is the case for nonparametric regression, making $\mathrm{CV}(J)$ a suitable sample analogue of the mean-square error of $\hat h_J$ in that case (see, e.g., Li1987). But in models with endogeneity (i.e., when $\mathbb{E}[u|X] \neq 0$), there is no guarantee that $\mathbb{E}[u(h_0(X) - \hat h_J(X))] = 0$ and so this third term---which depends on $J$---may be non-negligible even asymptotically. If so, cross validation gives a biased estimate of the MSE of $\hat h_J$ and is therefore not a meaningful criterion by which to choose $J$ in models with endogeneity. Indeed, a cross-validated choice of $J$ may not even lead to a consistent estimator of $h_0$ in model ((ref)).
In addition, even for nonparametric regression, the $J$ chosen by CV balances bias and sampling uncertainty in $L^2$ norm. Such as choice is not optimal for estimation of $h_0$ and its derivatives in sup-norm, nor is it sutiable for adaptive UCBs for $h_0$ and its derivatives.
We now present our data-driven choice $\tilde J$ of sieve dimension using B-spline bases. B-splines are characterized by their order $r$. In the simulations and empirical application, we use a cubic B-spline ($r = 4$) for $\{\psi_{Jj}\}_{j=1}^J$ and a quartic B-spline ($r = 5$) for $\{b_{Kk}\}_{k=1}^K$.\footnote{In the first submitted version we also used a quadratic B-spline ($r=3$) for $\{\psi_{Jj}\}_{j=1}^J$. In additional simulations we obtained very similar results with a Fourier basis for $\{b_{Kk}\}_{k=1}^K$.}
Let $\mc T = \{J=(2^l + r-1)^d : l \in \mb N_0 \}$ denote a dyadic grid of candidate values of $J$, where the integer $r$ is the order of the B-spline basis for $\{\psi_{Jj}\}_{j=1}^J$ (i.e., each $\psi_{Jj}$ is a piecewise polynomial of degree $r -1$). For example, $\mc T = \{J=2^l + 3: l \in \mb N_0 \} \equiv \{4,5,7,11,19,35,\ldots\}$ for a scalar $X$ ($d=1$) and cubic B-splines $(r=4)$.\footnote{ Letting $J$ vary over $\mc T$ ensures there is enough separation that we can accurately compare the bias and variance of estimators with different $J \in \mc T$. This helps improve the numerical stability of the method, coherent with implementations of Lepski's method in other nonparametric contexts.} The index $l$ is the resolution level. We construct $\{b_{Kk}\}_{k=1}^K$ similarly, using B-splines of order $(r + 1)$ because the reduced form is smoother than $h_0$. Given the resolution level $l$ for the basis for $X$, the resolution level for the basis for $W$ is $l_w = \lceil (l + q) d/d_w \rceil$ for some $q \in \mb N_0$ where $d_w$ is the dimension of $W$. Linking $l_w$ to $l$ in this manner defines a mapping $K(J)$ that satisfies $\lim_{J \to \infty} K(J)/J = c \in [1,\infty)$. We recommend taking $q$ as the second- or third-smallest value for which $K(J) \geq J$ holds for all $J$ (i.e., $q = 1$ or $q = 2$ if both $X$ and $W$ are of the same dimension). We advise against choosing $q$ any larger, as the number of basis functions increases exponentially in the resolution level. Let $J^+ = \min\{j \in \mathcal{T} : j > J\}$ be the smallest sieve dimension in $\mc T$ exceeding $J$.
For $J,J_2 \in \mc T$ with $J_2 > J$, the contrast $D_{J}(x)-D_{J_2}(x)$ is an estimate of $\hat h_J(x) - \hat h_{J_2}(x)$, whose variance can be estimated by
where $\hat \sigma_{J}^2(x)$ is defined in ((ref)) and $\widehat{\mf U}_{J,J_2}$ is a $n\times n$ diagonal matrix whose $i$th diagonal entry is $\hat u_{i,J} \hat u_{i,J_2}$. Moreover, the multiplier bootstrap version of $D_{J}(x)-D_{J_2}(x)$ is \[ D_{J}^*(x)-D_{J_2}^*(x) = (\psi^J(x))'\mf M_J^{\phantom \prime} \hat{\mf u}_J^* - (\psi^{J_2}(x))' \mf M_{J_2}^{\phantom \prime} \hat{\mf u}_{J_2}^*. \] Finally let $\hat s_J$ be the smallest singular value of $(\mf B_{K(J)}'\mf B_{K(J)}^{\phantom \prime})^{-1/2} (\mf B_{K(J)}' \mf \Psi_J^{\phantom \prime}) (\mf \Psi_J'\mf \Psi_J^{\phantom \prime})^{-1/2}$.
\centerline{\bf Procedure 1: Data-driven Choice of Sieve Dimension}
We present the theoretical results on the adaptivity of $\tilde J$ in Section (ref).
Let $\ul p >d/2$ denote the minimal degree of smoothness assumed for $h_0$. For instance, if $X$ is scalar and $h_0$ is Lipschitz, then one could take $\ul p = 1$ even through the true smoothness of $h_0$ is unknown. Let $\hat A = \log \log \tilde J$ and \[ \hat{\mc J}_{-} =
\]
\centerline{\bf Procedure 2: Data-driven UCBs for $h_0$}
\centerline{\bf Procedure 2$^\prime$: Data-driven UCBs for $\partial^a h_0$ ($0<|a|<\ul p$)}
Theoretical properties of these UCBs are presented in Sections (ref) and (ref). We show that the Procedures 2 and 2$^\prime$ UCBs are honest and adaptive for models in the mild regime (including nonparametric regression as a special case). For models in the severe regime, we show that the Procedures 2 and 2$^\prime$ UCBs with critical values corresponding to $\tilde J =\hat{J}_n$ have valid (actually conservative) coverage. Nevertheless, the Engel curve simulation in Appendix (ref) shows that the Procedure 2 UCBs still have valid (actually conservative) coverage for a severe regime design.
\citeauthor*{AAG} (AAG, hereafter AAG) derive semiparametric gravity equations for the extensive and intensive margins of firm exports in a monopolistic competition model of international trade. Importantly, and in sharp contrast with the existing literature Melitz2003,Chaney,EKK,HMT,MelitzRedding, AAG do not impose any parametric assumptions on the distribution of unobserved firm heterogeneity. The gravity equations identify functions which characterize the elasticities of the extensive and intensive margins of firm-level exports to changes in bilateral trade costs. AAG emphasize the importance of these elasticities for counterfactuals.
In this section, we apply our procedures to estimate and construct UCBs for the intensive margin and its elasticity using AAG's baseline model and data. We also present simulation studies based on empirical calibrations of two workhorse trade models to illustrate the sound performance of our procedures.
We begin by briefly summarizing the empirical framework of AAG. They use a monopolistic competition model of international trade---see MelitzRedding2014 for a review. There are a continuum of firms in each country. Firm $\omega$ in country $i$ is characterized by an entry potential $e_{ij}(\omega)$ and a revenue potential $r_{ij}(\omega)$ for selling in country $j$. Firms draw $e_{ij}(\omega)$ from a distribution $H_{ij}^e(e)$ then $r_{ij}(\omega)$ from a (possibly degenerate) distribution $H_{ij}^r(r|e)$. Firm $\omega$ in country $i$ exports to country $j$ if and only if $e_{ij}(\omega)$ exceeds a threshold. The proportion of firms in country $i$ that export to country $j$ is denoted $\pi_{ij}$.
The extensive margin is characterized by the inverse distribution of entry potential, i.e., $\epsilon_{ij}(\pi_{ij}) = (H_{ij}^e)^{-1}(1-\pi_{ij})$. Assuming homogeneity (so $H_{ij}^e = H^e$ and $\epsilon_{ij} = \epsilon$), AAG's gravity equation for the extensive margin is \[ \log \epsilon(\pi_{ij}) = \log (\bar f_{ij} \bar \tau_{ij}^{\tilde \sigma}) + \delta_i^\epsilon + \zeta_j^\epsilon, \] where $\bar \tau_{ij}$ and $\bar f_{ij}$ are variable and fixed trade costs from $i$ to $j$ and $ \delta_i^\epsilon$ and $ \zeta_j^\epsilon$ are exporter and importer fixed effects (FEs). Costs depend linearly on a cost shifter $z_{ij}$: \[
\] where the idiosyncratic error terms $\eta_{ij}^\tau$ and $\eta_{ij}^f$ are conditionally mean-zero and independent of $z_{ij}$ and the FEs. This yields the estimating equation
Note that $\pi_{ij}$ depends (possibly nonlinearly) on $z_{ij}$ and the error terms $\eta_{ij}^f$ and $\eta_{ij}^\tau$.
The intensive margin is characterized by the average revenue potential of exporting firms: \[ \rho_{ij}(\pi) = \frac{1}{\pi} \int_0^\pi \mathbb{E}[r|e = \epsilon_{ij}(v)] \, \mathrm d v\,, \] where the expectation is taken under $H_{ij}^r(r|e)$. Assuming homogeneity (so $H^r_{ij} = H^r$ and $\rho_{ij} = \rho$), AAG's gravity equation for the intensive margin is \[ \log \bar x_{ij} - \log \rho (\pi_{ij}) = \log (\bar \tau_{ij}^{\tilde \sigma}) + \delta_i^\rho + \zeta_j^\rho, \] where $\bar x_{ij}$ are average firm exports and $ \delta_i^\rho$ and $ \zeta_j^\rho$ are FEs. With $\bar \tau_{ij}$ as above, AAG obtain
More concisely,
where $y_{ij} := \log \bar x_{ij} + \tilde \sigma \kappa^\tau z_{ij}$ is the dependent variable,\footnote{AAG construct $y_{ij}$ from data on $\bar x_{ij}$ and $z_{ij}$ based on external estimates of $\tilde \sigma$ and $\kappa^\tau$.} $\tilde \pi_{ij} := \log \pi_{ij}$ is the endogenous regressor, $\log \tilde \rho(\tilde \pi) := \log \rho(e^{\tilde \pi})$ is the unknown structural function, $\delta_i := \delta_i^\rho - \tilde \sigma \delta_i^\tau$ and $\zeta_j := \zeta_j^\rho - \tilde \sigma \zeta_j^\tau$ are exporter and importer FEs, and the idiosyncratic error term $u_{ij} := - \tilde \sigma \eta_{ij}^\tau$ is conditionally mean-zero and independent of the instrumental variable $z_{ij}$.
Our goal is to use ((ref)) to estimate $\log \tilde \rho$ and its derivative, as $\frac{\partial \log \tilde \rho(\tilde \pi)}{\partial \tilde \pi}\equiv \frac{\partial \log \rho(\pi)}{\partial \log \pi}$ characterizes the elasticity of the intensive margin of firm-level exports to changes in bilateral trade costs. We use the same data that AAG use for their baseline estimates, which consists of $\bar x_{ij}$, $z_{ij}$, and $\tilde \pi_{ij}$ for a sample of 1522 country pairs for the year 2012. We refer the reader to AAG for a detailed description of the data and its construction.
Model ((ref)) differs from model ((ref)) due to the presence of FEs. AAG estimate $\log \tilde \rho$ and FEs jointly, using both $z_{ij}$ and exporter and importer country dummies as instruments. As such, they estimate a partially linear model with a large number of linear regressors (due to the country dummies) and, similarly, a large number of instrumental variables.\footnote{These comments are based on the November 2020 version of AAG, which is currently under revision. Some of their implementation and findings may differ in future versions.} Our methods and theoretical results are not formally developed for such a setting.\footnote{Our approach extends to partially linear models---see Section (ref). But with bilateral trade data the number of dummy variables representing origin and destination FEs is increasing with the sample size $n$. This “many regressors/many instruments” asymptotic framework falls outside the scope of our analysis.} Therefore, we maintain their assumption that $z_{ij}$ and origin and destination FEs are exogenous, but we further assume that $\mathbb{E}[\log \tilde \rho(\tilde \pi_{ij})|z_{ij}, \delta_i, \zeta_j] = \mathbb{E}[\log \tilde \rho(\tilde \pi_{ij})|z_{ij}]$ (a.s.). That is, the intensive margin is conditional mean independent of exporter- and importer-specific factors given cost shifters. Note, however, that we are not imposing that average firm exports are conditional mean independent of exporter- and importer-specific factors. The reduced form for $y_{ij}$ is
where $g(z_{ij}) = \mathbb{E}[\log \tilde \rho(\tilde \pi_{ij})|z_{ij}]$ and $\mathbb{E}[e_{ij} | z_{ij}, \delta_i, \zeta_j] = 0$. We estimate $\delta_i$ and $\zeta_j$ from ((ref)) by partially linear series regression. That is, we regress $y_{ij}$ on origin and destination dummies and functions $b_{K1},\ldots,b_{KK}$ of $z_{ij}$ at dimension $K(\hat J_{\max})$. We then apply our procedures using $Y_{ij} = y_{ij} - \hat \delta_i - \hat \zeta_j$ as the dependent variable ($Y$), $\tilde \pi_{ij}$ as the endogenous regressor ($X$), and $z_{ij}$ as the instrumental variable ($W$). We present simulations below for models with and without FEs and show that this first-stage estimation of $\delta_i$ and $\zeta_j$ does not affect the performance of our procedures. Appendix (ref) provides further details on implementation.
We implement our procedures using AAG's data. Our data-driven choice of sieve dimension is $\tilde J = 4$ for this sample. Figure (ref) plots our estimate of $\log \rho$ and the elasticity of the intensive margin, together with their 95% UCBs that are constructed as in displays ((ref)) and ((ref)), respectively. We report results over the interval $[0.1\%,50\%]$, as in AAG.
UCBs for $\log \rho$ and the elasticity of $\rho$ are both narrow and informative. Figure (ref) also plots a linear IV estimate of $\log \rho$ and the corresponding (constant) elasticity estimate.\footnote{For the linear IV estimates, we estimate $\log \rho$ jointly with the FEs as in AAG.} These both lie outside the UCBs for much of the support of $\pi_{ij}$. As such, our UCBs for the elasticity provide evidence against the Pareto specification for unobserved firm productivity used, e.g., by Chaney, under which the elasticity of $\rho$ is constant. Whereas Figure 1 of AAG shows that several conventional parameterizations of the distribution of unobserved firm heterogeneity used by EKK, HMT, and MelitzRedding all imply a decreasing elasticity over $[0.1\%,50\%]$. By contrast, decreasing elasticities necessarily fall outside our 95% UCBs over $[0.1\%,50\%]$, as the right-most point of the lower UCB lies above the upper UCB for smaller values of $\tilde \pi_{ij}$.
To show that our results are not sensitive to first-stage elimination of fixed effects, we also estimate $\log \rho$ and the FEs jointly, using our data-driven choice $\tilde J = 4$ and instrumenting with $b_{K(\tilde J)1}(z_{ij}),\ldots,b_{K(\tilde J)K(\tilde J)}(z_{ij})$ and the origin and destination dummies, and using $y_{ij}$ as the dependent variable. Estimates using this approach are also shown in Figure (ref) (labeled Joint NPIV + FEs). There is a vertical shift in the estimate of $\log \rho$ between the two approaches due to the different treatment of FEs, but the estimated elasticity---which is the focus of AAG---lies entirely within our 95% UCB for the elasticity and is very close to our data-driven elasticity estimate over the whole range $[0.1\%,50\%]$.
We now present simulation studies based on empirical calibrations of two workhorse trade models. The first design is based on HMT who assume a log-normal distribution for latent firm productivity. The second design is based on Chaney who assumes a Pareto distribution. In the first design the elasticity of $\rho$ is decreasing whereas in the second design $\log \rho(\pi) = \rho \log \pi$ and hence the elasticity is constant. For brevity we only present results for elasticity estimates in the log-normal design here. Additional results for the Pareto design and estimation of $\log \rho$ are deferred to Appendix (ref).
We generate data by first sampling $z_{ij}$ independently with replacement from its empirical distribution. We then generate data on $\tilde \pi_{ij}$ and $\bar x_{ij}$ by simulating from equations ((ref)) and ((ref)), using the expressions for $\log \epsilon(\pi)$ and $\log \rho(\pi)$ implied by the log-normal assumption---see Appendix (ref). As the empirical application has $n=1522$, we investigate the performance of our procedures across 1000 samples of size $761$, $1522$, $3044$, and $6088$.
Plots for a representative sample of size 1522 are presented in Figure (ref)(a). We generate the results in Table (ref) and Figure (ref) by implementing our procedures as in the empirical application. That is, the dependent variable is $Y_{ij} = y_{ij} - \hat \delta_i - \hat \zeta_j$, where $\hat \delta_i$ and $\hat \zeta_j$ are first-stage estimates of the exporter and importer fixed effects. We construct basis functions as in the application; see Appendix (ref) for details. We also compute estimates and confidence bands over the range 0.1% to 50% for $\pi_{ij}$ as reported in the application.
The first panel in Table (ref) presents the average and median (across simulations) of \[ \sup_{\pi \in [0.001, 0.5]} \left| \frac{d\, \widehat{\log \rho}(\pi)}{d \log \pi} - \frac{d \log \rho(\pi)}{d \log \pi} \right|, \] which is the maximal error of estimates of the elasticity of $\rho$ for $\pi_{ij}$ over $[0.1\%,50\%]$. We compare estimates using $\tilde J$ to estimates that use a deterministic choice of sieve dimension, namely $J = 4$, $5$, $7$, and $11$ (these are the first few values of $J$ over which our procedure searches). In each simulation, the maximal error is generally smallest with $J = 4$ or $J = 5$. The average $\tilde J$ is between 4.1 and 4.2 depending on the sample size. The maximal error of $\tilde J$ is at least half that with $J = 7$, and ten times smaller than with $J = 11$.
Turning to the coverage properties of UCBs for the elasticity, the second panel of Table (ref) shows our data-driven UCBs have correct but somewhat conservative coverage. Some conservativeness is to be expected, as our UCBs have uniform coverage guarantees over a class of DGPs. We also present coverage of UCBs based on the usual approach of “undersmoothing” from Section (ref). These UCBs use a deterministic $J$ and have valid coverage provided $J$ is chosen sufficiently large that bias is negligible relative to sampling uncertainty. Of course, in any empirical application a researcher does not know the true function, and therefore doesn't know which values of $J$ are sufficiently large that sampling uncertainty dominates bias. As can be seen from Table (ref), $J = 4$ or $J = 5$ seems too small, and consequently these bands under-cover. Bands with $J = 7$ have coverage closer to nominal coverage, but these bands are more than 70% wider than the data-driven bands. Comparing the UCBs in Figures (ref)(a) and (ref)(c), we see the efficiency improvement of our bands relative to undersmoothed bands with $J = 7$, for estimating both $\rho$ and its elasticity.
The fact that our UCBs are based on an optimal choice of $J$, and therefore contract faster than bands based on undersmoothing, has important practical consequences. Consider the data-driven UCBs for the elasticity of $\rho$ reported in Figure (ref)(a). These bands do not contain any constant function because the upper limit of the lower band exceeds the lower limit of the upper band. This provides evidence against the Pareto specification of productivity used by Chaney, for which the elasticity of $\rho$ is constant.\footnote{Table (ref) presents the frequency that such a test rejects the constant elasticity specification.} Note this is despite the fact that our bands tend to be a bit conservative. The undersmoothed bands with $J = 7$ have coverage closer to nominal coverage. But for the sample shown in Figure (ref), the undersmoothed bands with $J = 7$ are sufficiently wide that constant functions lie entirely within the bands. Hence, the researcher could not reject a constant elasticity specification on the basis of the undersmoothed bands in this sample. In fact, the undersmoothed bands with $J = 7$ only reject the constant elasticity specification in 15.8% of simulations with $1522$ observations whereas the rejection rate for the data-driven bands is 34.4%. This difference in rejection rates illustrates the general phenomenon that undersmoothed bands sacrifice efficiency for coverage. The undersmoothed bands are also quite wiggly, making it difficult to infer the shape of the true elasticity.
We note in closing that our procedures can equally be applied to other IV-based nonparametric analyses in international trade; see, e.g., ACD.
We first outline the main regularity conditions in Section (ref). Section (ref) shows that $\tilde J$ leads to minimax convergence rates for estimators of both $h_0$ and its derivatives. We then present the main results for UCBs in Sections (ref) and (ref).
We first state and then discuss the assumptions that we impose on the model and sieve space. We require these to hold for some constants $a_f, \ul c, \ol C, C_T, C_Q, \underline{\sigma}, \overline{\sigma} > 0$ and $\gamma \in (0,1)$. Let $T : L^2_X \to L^2_W$ denote the operator $Th(w) = \mathbb{E}[h(X)|W = w]$. For nonparametric regression we have $W \equiv X$ and so $T$ reduces to the identity.
Let $ \Psi_J$ and $ B_K$ be the closed linear subspaces of $L^2_X$ and $L^2_W$ spanned by $\psi_{J1},\ldots,\psi_{JJ}$ and $b_{K1},\ldots,b_{KK}$, respectively. Define \[ \tau_J = \sup_{h \in \Psi_J : \| h \|_{L^2_X} \neq 0} \frac{\|h \|_{L^2_X}}{\| Th \|_{L^2_W}}\,, \] where $\|\cdot \|_{L^2_X}$ and $\|\cdot\|_{L^2_W}$ denote the $L^2_X$ and $L^2_W$ norms. The sieve measure of ill-posedness $\tau_J$ quantifies the degree of difficulty of inverting $Th_0$ to recover $h_0$. As conditional expectations are (weakly) contractive, we have $\tau_J \geq 1$. Large $\tau_J$ indicate a more difficult inversion problem. The model ((ref)) is said to be mildly ill-posed (or in the mild regime) if $\tau_J \asymp J^{\varsigma/d}$ for some $\varsigma \geq 0$ and severely ill-posed (or in the severe regime) if $\tau_J \asymp \exp(C J^{\varsigma/d}) $ for some $C,\varsigma > 0$, where $d = \dim(X)$. For nonparametric regression models we have $\tau_J = 1$ for all $J$. Hence, nonparametric regression is a special case of the mild regime with $\varsigma = 0$.
Let $\Pi_J : L^2_X \to \Psi_J$ and $\Pi_{K(J)} : L^2_W \to B_{K(J)}$ denote LS projections onto $\Psi_J$ and $B_{K(J)}$: \[
\] Also let $Q_J : L^2_X \to \Psi_J$ denote the TSLS projection onto $\Psi_J$:
Denote the “population” sieve variance of $\hat h_J(x)$ as $\|{\sigma}_{x,J}\|^2_{sd} = L_{J,x}^{\phantom \prime} \Omega_J^{\phantom \prime} L_{J,x}'$ where $L_{J,x} = (\psi^J(x))' [S_J' G_{b,J}^{-1} S_J^{\phantom \prime} ]^{-1} S_J' G_{b,J}^{-1}$ and $\Omega_J = \mathbb{E} [ u^2 b^{K(J)}(W) (b^{K(J)}(W))' ]$ with $u = Y - h_0(X)$, $G_{b,J} = \mathbb{E} [ b^{K(J)}(W) (b^{K(J)}(W))' ]$, and $S_J = \mathbb{E} [ b^{K(J)}(W) (\psi^J(X))' ]$. Also let $\|\sigma_{x,J}\|^2 = (\psi^J(x))'[S_J' G_{b,J}^{-1} S_J^{\phantom \prime}]^{-1} (\psi^J(x))$, which satisfies $\|\sigma_{x,J} \| \asymp \|{\sigma}_{x,J}\|_{sd}$ uniformly in $x$ by Assumption (ref).
Assumptions (ref)(i)(ii) and (ref) are standard conditions on the support of $X$ and $W$ and the conditional variance of the errors (see, e.g., chen2018optimal) that can be relaxed. Assumption (ref)(iii) is an identification condition that is generically satisfied under endogeneity (see andrews2017examples) and is trivially satisfied for nonparametric regression because $T$ reduces to the identity in that case. Assumption (ref) is also trivially satisfied for nonparametric regression with $C_T,C_Q = 1$. Assumption (ref)(i) is imposed to ensure that $\hat s_J^{-1}$ is a suitable sample analog of $\tau_J$. Assumption (ref)(ii) is the usual $L^2$ “stability condition” imposed in the NPIV literature to derive $L^2$-norm rates. Assumption (ref)(iii) is a $L^{\infty}$-norm analogue used to control the bias in sup-norm. chen2018optimal provide a thorough discussion of Assumption (ref)(i) and derive primitive sufficient conditions for it in the context of nonparametric demand estimation. Assumption (ref)(ii) says that $\|{\sigma}_{x,J}\|^2_{sd}$ is increasing in $J \in \mathcal{T}$, uniformly in $x$. We view this as mild because $J$ increases exponentially over $\mathcal{T}$. Indeed, by Assumption (ref) and (ref)(i) and the fact that $J \asymp 2^{Ld}$ for some $L \in \mathbb{N}$, for any $J,J_2 \in \mathcal{T}$ with $J_2 > J$ we have \[ \sup_{x \in \mc X} \frac{\|{\sigma}_{x,J}\|_{sd}}{\|{\sigma}_{x,J_2}\|_{sd}} \asymp \frac{ \tau_J \sqrt{J}}{\tau_{J_2} \sqrt{J_2}} \leq \frac{\tau_{2^{Ld}}}{\tau_{2^{(L+1)d}}} 2^{-d/2}\leq 2^{-d/2}<1\,. \]
We now show $\tilde J$ leads to minimax rate-adaptive estimators of both the structural function $h_0$ and its derivatives. Our results encompass nonparametric regression as a special case.
We first define the parameter space for $h_0$. Let $B_{\infty,\infty}^p(M)$ denote the H\"older ball of smoothness $p$ and radius $M$ (see Appendix (ref) for a formal definition). For given constants $C_T,C_Q,M > 0 $ and $\ol p > \ul p > \frac d2$ with $r \geq \lfloor \ol p \rfloor + 1$, let $ \mathcal{H}^p = \mathcal{H}^p(M,C_T,C_Q)$ denote the subset of $B_{\infty,\infty}^p (M)$ that satisfies Assumption (ref)(ii)(iii) for any distribution of $(X,W,u)$ satisfying Assumptions (ref)-(ref), and let $\mc H = \bigcup_{p \in [\ul p,\ol p]} \mc H^p$. For each $h_0 \in \mc H$, we let $\mb P_{h_0}$ denote the distribution of $(X_i,Y_i,W_i)_{i=1}^\infty$ where each observation is generated by an IID draw from a distribution of $(X,W,u)$ satisfying Assumptions (ref)-(ref) with $Y = h_0(X) + u$.
We now show $\tilde J$ also leads to adaptive estimation of derivatives of $h_0$. Intuitively, estimating the derivative of $h_0$ inflates convergence rate of the (squared) bias and variance terms by the same factor (a power of $J$). Therefore, a rate-optimal choice of $J$ for estimating $h_0$ is also rate-optimal for estimating derivatives of $h_0$.
It is known since low1997 that it is impossible to construct confidence bands that are simultaneously honest and adaptive over H\"older classes of different smoothness. As is standard following picard2000, gine2010confidence, bull2012honest, chernozhukov2014anti, and many others, we establish coverage guarantees over a “generic” subclass $\mc G$ of $\mc H$. To describe $\mc G$, first note by the discussion in Appendix (ref) that there exists a constant $\ol B < \infty$ for which $\sup_{h \in \mc H^p} \|h - \Pi_J h\|_\infty \leq \ol B J^{- \frac pd}$ holds for all $J \in \mc T$ and all $p \in [\ul p, \ol p]$. For any small fixed $\ul B \in (0,\ol B)$ and any $\ul J \in \mc T$, we define \[ \mathcal{G}^p = \left\{ h \in \mc H^p : \ul B J^{- \frac pd} \leq \|h - \Pi_J h\|_\infty \mbox{ for all $J \in \mathcal{T}$ with $J \geq \ul J$} \right\},~~~\mc G = \bigcup_{p \in [\ul p, \ol p]} \mc G^p\,. \] The class $\mc G$ is sometimes called a class of “self-similar” functions. gine2010confidence,gine2016mathematical present several results establishing the genericity of $\mc G$ in $\mc H$. Loosely speaking, their results say $\mc H^p \setminus (\cup_{\ul B > 0,\ul J \in \mc T} \mc G^p)$ is nowhere dense in $\mc H^p$ under the norm topology of $\mc H^p$. Thus, the set of functions in $\mc H^p$ but not in $\mc G^p$ for some $\ul B$ and $\ul J$ is topologically meagre.
We say that a UCB $\{C_n(x) : x \in \mc X\}$ is honest over $\mc G$ with level $\alpha$ if
and adaptive if for every $\epsilon > 0$ there exists a constant $D$ for which \[ \liminf_{n \to \infty} \inf_{p \in [\ul p,\ol p]} \inf_{h_0 \in \mc G^p} \mathbb{P}_{h_0} \bigg( \sup_{x \in \mathcal{X}} |C_n(x)| \leq D r_n(p) \bigg) \geq 1 - \epsilon \,, \] where $|\,\cdot\,|$ is Lebesgue measure and $r_n(p)$ is the minimax sup-norm rate of estimation over $\mc H^p$. Let $C_n(x,A)$ denote the UCB from ((ref)) replacing $\hat A$ with a constant $A > 0$. Our first main result is that $C_n(x,A)$ is honest and adaptive in the mildly ill-posed case:
Theorem (ref) establishes that the UCBs for $h_0$ in Procedure 2 is honest and adaptive in the mildly ill-posed case. We have found that the UCBs in Procedure 2 perform well in terms of coverage across many simulation designs including the severely ill-posed design in Appendix (ref). Nevertheless, for the severely ill-posed case, we can only establish valid coverage of the UCBs in Procedure 2 using the critical value $\mbox{cv}^*(x)$ corresponding to $\tilde J = \hat J_n$ case, i.e.,
The term $\tilde {J}^{-\ul p/d}$ bounds the order of the bias term $\|\Pi_{\tilde J} h_0 - h_0 \|_{\infty}$, which accounts for the fact that the optimal choice of $J$ in severely ill-posed models is bias-dominating. This band reduces to the Procedure 2 UCB when $\theta^*_{1-\hat \alpha} \geq \tilde {J}^{-\ul p/d} /\hat \sigma_{\tilde J}(x)$ for all $x$.
Let $C_n(x,A)$ denote the UCB ((ref)) with the critical value ((ref)), except replacing $\hat A$ with a constant $A > 0$.
Our recommended choice $\hat A = \log \log \tilde J$ ensures that the UCBs are asymptotically valid over $\mc G$ for any $\underline B > 0$ and contract within a $\log \log n$ factor of the minimax rate if the true smoothness is $p = \ul p$, and within a $\log n$ factor of the minimax rate otherwise.
We now present an analogous set of results for data-driven UCBs for derivatives of $h_0$. Here we require an additional regularity condition similar to Assumption (ref)(i), which is only needed for the results in this subsection. Let $\|\sigma^a_{x,J}\|^2 = (\partial^a \psi^J(x))'[S_J' G_{b,J}^{-1} S_J^{\phantom \prime}]^{-1} (\partial^a \psi^J(x))$.
\setcounter{assumption}{3}
We first present results for the mildly ill-posed case. Let $C_n^a(x,A)$ denote the UCB $C_n^a(x)$ from ((ref)) when $\hat A$ is replaced by a constant $A > 0$.
As in the previous subsection, for the severely ill-posed case, we can only establish valid coverage of the UCB ((ref)) using the critical value $\mbox{cv}^{a*}(x)$ corresponding to $\tilde J = \hat J_n$, i.e.,
This band reduces to the Procedure 2$^\prime$ UCB when $\theta^*_{1-\hat \alpha} \geq \tilde {J}^{|a|-\ul p/d} /\hat \sigma_{\tilde J}^a(x)$ for all $x$, which is the case in our empirical application.
Let $C_n^a(x,A)$ denote the band ((ref)) with critical value ((ref)) when $\hat A$ is replaced by a constant $A>0$.
In this section we present two additional simulation studies. The first is a nonparametric IV design with a non-monotonic, non-Lipschitz structural function. The second is a very wiggly nonparametric regression design, which shows that $\tilde J$ can choose a relatively high-dimensional model when needed. Finally, Appendix (ref) presents a third set of simulations in an empirically calibrated Engel curve design which is severely ill-posed.
This design features a non-monotonic, non-Lipschitz structural function. We first draw $(U,V)$ from a bivariate normal distribution with mean zero, unit variance, and correlation $0.75$, and draw $Z \sim N(0,1)$ independent of $(U,V)$. We then set $W = \Phi(Z)$ where $\Phi(\cdot)$ denotes the standard normal CDF, $X = \Phi( D(Z + V) + (1-D) V)$ where $D$ is an independent Bernoulli random variable taking the values $0$ and $1$ each with probability $0.5$, and
The structural function $h_0(x) = \sin (4x) \log(x)$ is plotted in Figure (ref). Note that the derivative of $h_0$ diverges to $-\infty$ as $x \downarrow 0$. Therefore, $h_0$ is H\"older continuous with exponent $p$ for any $p < 1$, but not Lipschitz continuous.
For each simulated data set we compute our data-driven estimator $\hat h_{\tilde J}$ and UCBs from ((ref)). We compare these with estimators and UCBs using deterministic choices of sieve dimensions for $J = 4$, $5$, $7$, and $11$ (the first few dimensions over which our procedure searches). We again use a cubic B-spline basis to approximate $h_0$ and a quartic B-spline for the reduced form.
The first panel of Table (ref) presents the average sup-norm loss of $\hat h_{\tilde J}$ across simulations. These are of similar magnitude to the loss for deterministic-$J$ estimates with $J = 4$ and $5$ and are much smaller than the loss with $J = 7$ and $11$. Our data-driven UCBs demonstrate valid but slightly conservative coverage for smaller $n$ and coverage close to nominal coverage for $n = 10000$. Bands with $J = 4$ have poor coverage while bands with $J = 5$ have valid coverage for the smaller sample sizes but under-cover for $n = 10000$. It seems $J = 7$ or $J = 11$ is required to have valid coverage for $n = 10000$ in this design. Note that while our bands are slightly conservative for smaller $J$, they are only about 10% wider than the $J = 5$ bands, and less than half the width of the $J = 7$ bands.
In Figure (ref) we plot data-driven estimates and UCBs for $h_0$ and its derivative over $[0.01, 0.99]$ for a sample of size 2500, alongside deterministic-$J$ estimates and UCBs. In this sample, $\tilde J = 4$ and our data-driven UCBs contain the true structural function. The data-driven bands are narrower and more accurately convey the shape of $h_0$ than the $J = 7$ bands, which are much more wiggly. Our bands are also of a similar width to (but are less wiggly than) the $J = 5$ bands. Panel (d) of Figure (ref) also presents data-driven estimates and UCBs for the conditional mean of $Y$ given $X$. Here the data-driven choice is again $\tilde J =4$. The true structural function falls outside the UCBs for the conditional mean function over almost all of the support of $X$, highlighting the importance of estimating $h_0$ using IV methods in this design.
Finally, in Table (ref) we present the coverage of our data-driven UCBs $C_n(x,A)$ where we replace $\hat A = \log \log \tilde J$ with a deterministic choice $A$ ranging over $[0,1]$. For this design, $A \geq 0.3$ suffices for correct coverage. In particular, $\hat A = \log \log \tilde J$ yields correct coverage.
For this design we simulate $X \sim U[0,1]$ and $U \sim N(0,1)$ independently, then set
Here $h_0(x) = \sin(15 \pi x) \cos (x)$ is very wiggly over $[0,1]$ and requires a high value of $J$ to be selected in order to well approximate $h_0$ (see Figure (ref)). While $h_0$ is infinitely differentiable, its Lipschitz constant is at least $47.1$, the Lipschitz constant of its derivative is at least $2220$, and Lipschitz constants grow rapidly for higher derivatives.
We again compare our data-driven estimator and UCBs using the procedures described in Appendix (ref) with estimators and UCBs that use deterministic choices of $J$ for $J = 11$, $19$, $35$, and $67$ (these are a subset of values over which our procedure searches). We again use cubic B-splines to approximate $h_0$.
It is clear from the simulation results presented in Table (ref) that $J > 19$ is required to well approximate the true $h_0$. The average sup-norm loss of $\hat h_{\tilde J}$ is similar to that of the deterministic-$J$ estimator for $J = 35$, and is smaller than the average loss for all other $J$ presented in the table. Our data-driven UCBs also deliver valid, but conservative, coverage for the true conditional mean function. UCBs based on a deterministic choice of $J$ have zero coverage for $J = 11$ and $J = 19$ as these dimensions are too small to adequately approximate $h_0$, and tend to under-cover for the remaining $J$, except perhaps for $J = 35$ when $n = 10000$.
In this design a much smaller value of $A$ suffices to deliver valid coverage, as seen in Table (ref). The reason is that the set $\hat{\mc J}_-$ is large and $\hat h_J$ varies a lot across different $J$ due to the wiggliness of $h_0$. Therefore $z_{1-\alpha}^*$, which is the quantile of a sup-statistic over $\mc X \times \hat{\mc J}_-$, is relatively more conservative than for the other designs. This extra conservativeness suffices to deliver valid coverage in this design with smaller $A$.
Figure (ref) plots our data-driven estimator $\hat h_{\tilde J}$ and 95% UCBs for the conditional mean function for a sample of size 2500. In this sample, $\tilde J = 35$. The data-driven estimator well approximates the true conditional mean function $h_0$, which lies entirely within the 95% UCBs, and the same is true for estimates and UCBs for the derivative of $h_0$. Deterministic-$J$ bands with $J = 67$ are of a similar width to our data-driven bands for this sample, even though they use a less conservative critical value which only accounts for sampling uncertainty. The estimator is also much wigglier with $J = 67$ than our data-driven estimator and does not approximate $h_0$ as well.
So far we have assumed the structural function $h_0$ is a general $d$-variate function. As with many other nonparametric estimation problems, minimax rates deteriorate as $d$ increases. This so-called curse of dimensionality applies to any estimator of $h_0$. However, it can be circumvented by imposing additional structure on $h_0$ (when appropriate), such as additivity or partial linearity. In this section, we show how our data-driven procedures extend to additive and partially linear models.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Additive Structural Functions.}
Consider first the additive structural function: \[ h_0(x) = c_0 + h_{10}(x_1) + \ldots + h_{d0}(x_d) \] where $x = (x_1,\ldots,x_d)'$. Here $c_0$ is a constant representing an “intercept” term and the $h_{i0}$ are suitably normalized for identifiability. In the context of nonparametric regression, Stone1985 showed that imposing additivity can yield estimators of $h_0$ that achieve the same (optimal) rate for general $d$ as for $d = 1$.
Our methods may be easily adapted to additive models as follows. We assume for sake of exposition that $X$ is bivariate $(d = 2)$. Let $\psi^J(x) = (1,\tilde \psi_1^J(x_1)',\tilde \psi_2^J(x_2)')'$ where for $i = 1,2$ we have $\tilde \psi_i^J(x_i) = (\tilde \psi_{J1}(x_i),\ldots,\tilde \psi_{JJ}(x_i))'$. Here $J$ represents the dimensions of sieves used to approximate both $h_{10}$ and $h_{20}$. The basis functions $\tilde \psi_{J1},\ldots,\tilde \psi_{JJ}$ are formed by setting $\tilde \psi_{Jj}(x_i) = \psi_{Jj}(x_i) - \int_0^1 \psi_{Jj}(v) \mr d v$ with $\psi_{J1}(x_1),\ldots,\psi_{JJ}(x_1)$ a univariate B-spline basis. We estimate $c_0$ and $c_J^i$, $i = 1,2$, by TSLS regression of $Y$ on $\psi^J(X)$ using $b^{K(J)}(W)$ as instruments: \[ \left(
\right) = \left(\mf \Psi_J' \mf P_{K(J)}^{\phantom \prime} \mf \Psi_J^{\phantom \prime} \right)^{-} \mf \Psi_J' \mf P_{K(J)}^{\phantom \prime} \mf Y =\mf M_J \mf Y\,, \] where the notation is as in Section (ref) but with $\psi^J(x) = (1,\tilde \psi_1^J(x_1)',\tilde \psi_2^J(x_2)')'$. The estimator of $h_{i0}$ is $\hat h_{iJ}(x_i) = (\psi^J_i(x_i))' \hat c_J^i$. Derivatives of $h_{i0}$ are estimated by differentiating $\hat h_{iJ}$.
Our data-driven choice of $J$ is implemented exactly as described in Section (ref) with $\psi^J(x) = (1,\tilde \psi_1^J(x_1)',\tilde \psi_2^J(x_2)')'$. Data-driven UCBs for $h_{10}$ are formed analogously to Section (ref) with two small modifications. First, when computing the critical value $z_{1 - \alpha}^*$ in Step 4 of Procedure 2 we now use the sup-statistic \[ \sup_{(x_1,J) \in [0,1] \times \hat{\mc J}_{-}} \left| \frac{D_{1J}^*(x_1)}{\hat \sigma_{1J}(x_1)} \right| \] where $D_{1J}^*(x_1) = (0,\tilde \psi_1^J(x_1)',0_J')' \mf M_J \hat{\mf u}_J^*$ with $0_J$ a $J$-vector of zeros, and \[ \hat \sigma_{1J}^2(x) = (0,\tilde \psi_1^J(x_1)',0_J') \mf M_J^{\phantom \prime} \widehat{\mf U}_{J,J}^{\phantom \prime} \mf M_J' (0,\tilde \psi_1^J(x_1)',0_J')'\,. \] The 100$(1-\alpha)$% UCB for $h_{10}$ is
with \[ cv^*(x_1) =
\] where $\ul p$ is the minimal smoothness assumed for $h_{10}$ and $h_{20}$. UCBs for derivatives of $h_{10}$ are constructed analogously.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Partially Linear Structural Functions.}
An alternative to additivity is the partially linear specification ai2003efficient \[ h_0(x) = h_{10}(x_1) + x_2'\beta_0 \] where $x$ is partitioned as $x = (x_1',x_2')'$ with $x_1$ of dimension $d_1 < d$, $h_{10}$ is an unknown function, and $\beta_0$ is an unknown vector of parameters. When $X$ is exogenous (so $W \equiv X$) this is the important partially linear regression model of Robinson1988.
Our methods may be adapted to estimate and construct UCBs for $h_{10}$ as follows. First, we let $\psi^J(x) = (\psi^J_1(x_1)',x_2')'$ where $\psi^J_1(x_1) = (\psi_{J1}(x_1),\ldots,\psi_{JJ}(x_1))'$.\footnote{We assume without loss of generality that the $X_2$ variables have mean zero, which permits identification of $h_0$ and $\beta$. In practice these variables can be de-meaned.} We estimate $c_J$ and $\beta$ by TSLS regression of $Y$ on $\psi^J(X)$ using $b^{K(J)}(W)$ as instruments: \[ \left(
\right) = \left(\mf \Psi_J' \mf P_{K(J)}^{\phantom \prime} \mf \Psi_J^{\phantom \prime} \right)^{-} \mf \Psi_J' \mf P_{K(J)}^{\phantom \prime} \mf Y =\mf M_J \mf Y\,, \] where the notation is as in Section (ref) but with $\psi^J(x) = (\psi^J_1(x_1)',x_2')'$. The estimator of $h_{10}$ is $\hat h_{1J}(x_1) = (\psi^J_1(x_1))' \hat c_J$. Derivatives of $h_{10}$ are again estimated by differentiating $\hat h_{1J}$. When $X$ is exogenous, we simply take $w = x$ and $b^K(w) = \psi^J(x)$.
Our data-driven choice of $J$ is implemented analogously to Section (ref), except we form the contrasts $D_J$, $D_{J}(x)-D_{J_2}(x)$, and $D_{J}^*(x)-D_{J_2}^*(x)$ and the variance terms $\hat \sigma_J^2(x)$ and $\hat \sigma_{J,J_2}^2(x)$ using $\psi^J_0(x_1) := (\psi^J_1(x_1)',0_{d_2})'$ in place of $\psi^J(x)$. As such, the $t$-statistics are functions of $x_1$ only and the supremums in the sup-statistics in Steps 2 and 3 of Procedure 1 only need to be computed over the support $\mc X_1$ of $x_1$. UCBs for $h_{10}$ are constructed analogously to Section (ref), where the contrast $D_J^*(x)$ and the variance term $\hat \sigma_{J}^2(x)$ are again formed using $\psi^J_0(x_1)$ in place of $\psi^J(x)$. The 100$(1-\alpha)$% UCB for $h_{10}$ is
with \[ cv^*(x_1) =
\] where $\ul p$ is the minimal degree of smoothness assumed for $h_{10}$. UCBs for derivatives of $h_{10}$ are constructed analogously.
We have introduced data-driven procedures for estimation and inference on a nonparametric structural function $h_0$ and its derivatives using instrumental variables. Our data-driven choice of sieve dimension leads to estimators of $h_0$ and its derivatives that converge at the fastest possible (i.e., minimax) rate in sup-norm. Our data-driven uniform confidence bands (UCBs) for $h_0$ and its derivatives are shown to have coverage guarantees and contract at, or within a logarithmic factor of, the minimax rate. Both procedures have good finite sample performance in various simulation designs, including empirically-calibrated trade and Engel curve designs. Our methods are simple to compute, and are applied to estimate and construct UCBs for the elasticity of the intensive margin of firm exports in a monopolistic competition model of international trade.
Aside from the extensions in Section (ref), it would be straightforward to extend our methods to weakly dependent data, which is relevant for dynamic causal inference and reinforcement learning. It would also be interesting to consider sup-norm rate-minimaxity jointly with respect to both $p$ and the degree of ill-posedness.