Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
102,236 characters · 23 sections · 43 citation commands
Averaging estimation for instrumental variables quantile regression
\onehalfspacing
Since KoenkerBassett1978's (KoenkerBassett1978) seminal work, quantile regression (QR) has become a useful tool to capture unobserved heterogeneous effects in policy analysis or program evaluation. However, in practice endogeneity commonly results in inconsistent estimates of the conventional quantile regression. To address this endogeneity issue in quantile regression, ChernozhukovHansen2005 propose an instrumental variables method to identify the structural quantile function or (conditional) quantile treatment effects. Subsequently, ChernozhukovHansen2006 and then many others have provided estimation methods for the instrumental variables quantile regression (IVQR) model. Although there are important distinctions, for now I simply refer to “the” IVQR estimator.
Although the IVQR estimator has desirable large-sample properties like consistency and asymptotic normality, it can have imprecise estimates due to its large finite-sample variance,\footnote{It is possible that the IVQR variance is infinite, similar to IV Kinal1980. More technically, the word “variance” in this paper means a different measure of dispersion/spread that is never infinite, like trimmed/truncated variance or interquartile range.} just as the two-stage least squares (2SLS) estimator may have a substantial finite-sample dispersion.
The goal of this paper is to propose new estimation methods to improve the finite-sample estimation efficiency of IVQR. I propose two such methods. The first applies the averaging generalized method of moments (GMM) framework of \citet*{ChengLiaoShi2019} (hereafter \citetalias{ChengLiaoShi2019}) to IVQR. The second is a bootstrap averaging method, which uses the bootstrap world's optimal averaging weights on IVQR, QR, and 2SLS estimators.
This paper has four main contributions. First, beyond extending the 2SLS/OLS averaging to the analogous IVQR/QR averaging, this paper shows it is also helpful to include 2SLS in averaging to improve on IVQR. Second, the implementation of \citetalias{ChengLiaoShi2019}/GMM averaging is not trivial for IVQR. In particular, it needs two-step GMM estimation and nonparametric Jacobian matrix estimation, which are discussed or extended in this paper (and my code). Third, the bootstrap IVQR/QR/2SLS averaging method is new and outperforms \citetalias{ChengLiaoShi2019}/GMM averaging in simulations, although it lacks theoretical results like in \citetalias{ChengLiaoShi2019}. Fourth, the simulations compare various methods in a wide range of data-generating process (DGP) types (varying endogeneity, heterogeneity, distributional shapes, etc.), although they are still limited.
\paragraph{Averaging estimation in this paper} My first method follows the framework of \citetalias{ChengLiaoShi2019}. \citetalias{ChengLiaoShi2019} first define a benchmark conservative GMM estimator based on a set of valid moments. Then they add an additional set of possibly misspecified moments to the conservative moments to obtain the “aggressive” GMM estimator. On one hand, the aggressive GMM estimator might be more biased than the conservative GMM estimator since additional moments might not be valid. On the other hand, by adding this additional information, the aggressive GMM estimator might significantly reduce variance and overall mean squared error (MSE) compared to the conservative GMM estimator. \citetalias{ChengLiaoShi2019} propose an optimal weight formula to average the conservative and aggressive GMM estimators. They show that under certain conditions the averaging estimator uniformly dominates the conservative GMM estimator in asymptotic risk.
This paper applies \citetalias{ChengLiaoShi2019} to IVQR as follows. The conservative GMM estimator uses only the IVQR moments. Two types of additional moments are proposed: the conventional QR moments and the 2SLS slope moments (excluding the intercept term). The motivation for proposing these two additional moments comes from the bias--variance tradeoff. When there is not much endogeneity in the model, the conventional QR estimator is little biased; meanwhile, the QR estimates usually have smaller variance than IVQR estimates. Therefore, introducing QR moments as the additional moments can reduce variance and maybe achieve an overall reduction in MSE, although it might increase the bias. When there is not much heterogeneity across quantiles, 2SLS and IVQR at any quantile will usually have similar slope estimates, thus similar bias; at the same time, 2SLS estimates usually have a smaller variance than IVQR estimates, except in cases like a fat-tailed error term. Therefore, using 2SLS slope moments as the additional moments can also improve efficiency by reducing the overall MSE. From another perspective, 2SLS can be viewed as a limiting case of smoothed IVQR estimation as the smoothing bandwidth goes to infinity KaplanSun2017. These two types of additional moments yield two types of aggressive estimators to average with the IVQR estimator. I apply \citetalias{ChengLiaoShi2019}'s empirical averaging weight formula to obtain the averaging estimator. Simulation results demonstrate the averaging estimator has MSE uniformly below or equal to that of the IVQR estimator at all quantiles under certain uniform dominance conditions.
Besides the \citetalias{ChengLiaoShi2019} GMM averaging method, I propose a new bootstrap averaging method. The bootstrap averaging method averages the IVQR, QR, and 2SLS estimators in the bootstrap world with a grid of fixed weights, and picks the weight that minimizes the robust root mean squared error (robust RMSE, or rRMSE) as the bootstrap optimal weight. This optimal weight is then used to average the IVQR, QR, and 2SLS estimators in the original sample to obtain the bootstrap averaging estimator.
The motivation for the bootstrap averaging method is the same as for the additional QR and 2SLS slope moments in the GMM averaging method. The QR and 2SLS estimators might have smaller MSE than the IVQR estimator in some DGPs with little endogeneity or little heterogeneity, respectively. The hope is that the bootstrap world is similar enough to the real world that the bootstrap method places more weight on the 2SLS and/or QR estimator when they perform better than the IVQR estimator, and put more weight on the IVQR estimator when 2SLS and QR have larger MSE than IVQR.
The bootstrap averaging method, although lacking theoretical results, has potential advantages over the \citetalias{ChengLiaoShi2019} averaging GMM method when applied to IVQR. Bootstrap averaging is easier for computation since it avoids highly over-identified quantile GMM, which has a difficult criterion function to minimize. Moreover, the bootstrap averaging estimator has better performance in simulation results. This is partly because the bootstrap can average among three different estimators, whereas \citetalias{ChengLiaoShi2019} only averages between two estimators. It might also be because often in finite samples QR outperforms the aggressive IVQR-QR GMM estimator and 2SLS outperforms the aggressive IVQR-2SLS GMM estimator in situations when the additional moments provide significant variance reduction.
\paragraph{Literature}
Averaging estimation Averaging estimation originates from Stein-like shrinkage estimation and has recently been reinvestigated and extended by many authors to improve estimation efficiency. \Citet{JamesStein1961} propose an estimator that shrinks the least squares estimator toward zero. The James--Stein estimator can be viewed as in the class of averaging estimation, in that it averages the least squares estimator and zero. Under certain conditions, the James--Stein estimator dominates the least squares estimator in terms of a strict reduction of MSE. A limitation of the James--Stein estimator is it restricts to the normal distribution. \Citet{Maasoumi1978} applies the idea of Stein-like shrinkage estimation to simultaneous equations by averaging the three-stage least squares (3SLS) estimator with the least squares (LS) estimator. He shows while the 3SLS and 2SLS estimators have no finite moments (therefore, unbounded risk) in some cases, the averaging estimator has finite moments (therefore, bounded risk). Recently, Hansen2017 applies averaging estimation to a single equation instrumental variables model. He averages 2SLS and OLS estimators with weight depending on the statistic for testing exogeneity in the model. He shows that the averaging estimator uniformly dominates 2SLS in asymptotic risk when the number of endogenous regressors is greater than two. It extends the James--Stein framework to a general error term distribution but still limits to homoskedasticity. The averaging GMM method of \citetalias{ChengLiaoShi2019} works in a general framework with no normality or homoskedasticity restriction. \citetalias{ChengLiaoShi2019} provide supporting simulations to show the uniform dominance results in their averaging GMM framework can hold with both Gaussian and non-Gaussian errors.
IVQR estimation and computation Since ChernozhukovHansen2005's (ChernozhukovHansen2005) seminal work on IVQR identification, many researches focus on IVQR estimation and computational efficiency. The challenging computation of IVQR estimation comes from the non-differentiable IVQR moments and the non-convex GMM objective function. \Citet{ChernozhukovHansen2006} first propose a two-step inverse quantile regression method with a grid search on the endogenous regressors' coefficients to compute the IVQR estimator. This IVQR estimator is asymptotically equivalent to a GMM estimator. However, its computation time scales poorly with the number of endogenous regressors. \Citet{ChenLee2018} propose an exact GMM estimator using mixed-integer quadratic programming, but it also has long computation time. \Citet{Zhu2019} proposes a $k$-step correction approach using mixed integer linear programming. This estimator is asymptotically equivalent to the GMM estimator and has computational efficiency in models with multiple endogenous regressors. \Citet{KaidoWuthrich2019} decompose IVQR estimation into conventional QR sub-problems. \Citet{KaplanSun2017} propose smoothing the IVQR moments, which helps both estimation and computational efficiency. However, their results only apply to an exactly-identified linear model and iid sampling. \Citet{deCastroGalvaoKaplanLiu2019} propose a GMM estimator using the same smoothed IVQR moments, extending results to over-identified nonlinear models and dependent data. I use the deCastroGalvaoKaplanLiu2019 in this paper because \citetalias{ChengLiaoShi2019} require a two-step GMM estimator.
Other identification methods in quantile regression with endogeneity There are three main approaches to address endogeneity in quantile regression. They are based on different sets of assumptions appropriate for different empirical settings; none is strictly “better” or “worse.” First, as noted, ChernozhukovHansen2005 use an instrumental variables approach to identify the structural quantile function and (conditional) quantile treatment effects. Second, the local quantile treatment effect (LQTE) model AbadieEtAl2002 identifies the (conditional) quantile treatment effect for the sub-population of “compliers” in the binary treatment variable case, parallel to the local average treatment effect (LATE) model. Third, triangular models and control functions have been used by Chesher2003, Lee2007, and others. See ChernozhukovHansenWuthrich2017 and MellyWuthrich2017 for more detailed comparisons with additional references. Different from these listed studies that focus on identification, I focus on improving estimation efficiency within the IVQR framework of ChernozhukovHansen2005.
\paragraph{Outline} The organization of this paper is as follows. (ref) presents the model setup. (ref) presents the GMM averaging estimation method. (ref) presents the bootstrap averaging estimation method. (ref) presents simulation results. (ref) concludes. Proofs and additional computational details are collected in the appendix. Code is provided for all methods and simulations.
\paragraph{Notation} For scalar/vector/matrix variable formatting, $\boldsymbol{\mathbf{X}}$ is a random vector with elements $X_j$, $\boldsymbol{\mathbf{x}}$ is a non-random vector with elements $x_j$, $Y$ and $y$ are random and non-random scalars, respectively, and $\underline{\boldsymbol{\mathbf{M}}}$ and $\underline{\boldsymbol{\mathbf{m}}}$ are random and non-random matrices with row $i$, column $j$ elements $M_{ij}$ and $m_{ij}$. For vector/matrix multiplication, all vectors are treated as column vectors. Also, $\operatorname{\mathds{1}}\mathopen{}\mathclose\bgroup\originalleft\{\cdot\aftergroup\egroup\originalright\}$ is the indicator function, $\operatorname{E}(\cdot)$ expectation, $\operatorname{Q}_\tau(\cdot)$ the $\tau$-quantile, $\operatorname{P}(\cdot)$ probability, and $\mathrm{N}(\mu,\sigma^2)$ the normal distribution. Acronyms used include those for instrumental variables (IV), two-stage least squares (2SLS), generalized method of moment (GMM), [smoothed] instrumental variables quantile regression ([S]IVQR), probability density function (PDF), cumulative distribution function (CDF), quantile regression (QR), conditional quantile function (CQF), data-generating process (DGP), [root] mean squared error ([R]MSE), asymptotic mean squared error (AMSE), and interquantile range (IQR).
We are interested in estimating the parameter $\boldsymbol{\mathbf{\beta}}_{0\tau}\in\mathcal{B}\subseteq{\mathbb R}^{d_\beta}$ in a linear quantile model that uniquely satisfies the conditional probability
where $\tau \in (0,1) $ is a given quantile level; $ Y_{i}$ is the outcome variable; $\boldsymbol{\mathbf{X}}_i= \big( \boldsymbol{\mathbf{X}}_{exog, i}' , \boldsymbol{\mathbf{D}}_{i}' \big)' \in\mathcal{X}\subseteq{\mathbb R}^{d_X}$ is the vector of regressors; $ \boldsymbol{\mathbf{D}}_{i}$ is the vector of potentially endogenous explanatory variables; $\boldsymbol{\mathbf{X}}_{exog, i}$ is the vector of exogenous explanatory variables; and $\boldsymbol{\mathbf{Z}}_i=\big( \boldsymbol{\mathbf{X}}_{exog, i}' , \boldsymbol{\mathbf{Z}}_{excl, i}' \big)' \in\mathcal{Z}\subseteq{\mathbb R}^{d_Z}$ is the full vector of instruments, which contains both the exogenous explanatory variables $\boldsymbol{\mathbf{X}}_{exog, i}$ and excluded instruments $\boldsymbol{\mathbf{Z}}_{excl, i}$. The conditional probability in (ref) comes from ChernozhukovHansen2005's (ChernozhukovHansen2005) identification result in their Theorem 1, which states conditions under which the $\boldsymbol{\mathbf{\beta}}_{0\tau}$ satisfying (ref) is a structural parameter or includes a (conditional) quantile treatment effect parameter.
For intuition about identification, suppose a structural random coefficient model
with unobserved scalar $U \sim \textrm{Unif}(0,1)$, and $\boldsymbol{\mathbf{\beta}}(\cdot)$ is the vector-valued function of $U$ satisfying the monotonicity condition that $\boldsymbol{\mathbf{X}}' \boldsymbol{\mathbf{\beta}}(u)$ is increasing in $u$ for any $\boldsymbol{\mathbf{X}}=\boldsymbol{\mathbf{x}}$ in its support $\mathcal{X}$. Monotonicity implies that $Y\le\boldsymbol{\mathbf{X}}'\boldsymbol{\mathbf{\beta}}(\tau)$ is equivalent to $U\le\tau$. If additionally $\boldsymbol{\mathbf{X}}$ is exogenous with $\boldsymbol{\mathbf{X}} \protect\mathpalette{\protect\independenT}{\perp} U$, then $\operatorname{P}( Y \le \boldsymbol{\mathbf{X}}' \boldsymbol{\mathbf{\beta}}(\tau) \mid \boldsymbol{\mathbf{X}} ) = \operatorname{P}( U \le \tau \mid \boldsymbol{\mathbf{X}})$ by monotonicity, $\operatorname{P}( U \le \tau \mid \boldsymbol{\mathbf{X}})=\operatorname{P}(U\le\tau)$ by exogeneity, and $\operatorname{P}(U\le\tau)=\tau$ by the normalization $U\sim\textrm{Unif}(0,1)$. If $\boldsymbol{\mathbf{X}}$ and $U$ are dependent, then we can instead condition on instrument $\boldsymbol{\mathbf{Z}}$ satisfying $\boldsymbol{\mathbf{Z}}\protect\mathpalette{\protect\independenT}{\perp} U$, yielding (ref) where $\boldsymbol{\mathbf{\beta}}_{0\tau} \equiv \boldsymbol{\mathbf{\beta}}(\tau)$.
The conditional probability in (ref) can be written as a conditional expectation,
By the law of iterated expectations, (ref) implies the unconditional moments
Most IVQR estimators use (ref). In principle, the $\boldsymbol{\mathbf{Z}}_i$ in (ref) could be replaced by the (nonparametrically estimated) optimal instruments based on (ref). However, this is not the focus of this paper. This paper estimates structural parameter $\boldsymbol{\mathbf{\beta}}_{0\tau} $ based on the unconditional moments in (ref).
We introduce a few notations. Define the population map $\boldsymbol{\mathbf{M}}_{1} \colon \mathcal{B} \times \mathcal{T} \mapsto {\mathbb R}^{d_Z}$ as
where the subscript $1$ denotes the original IVQR moments, used as the “conservative” moments in GMM averaging later. Population moments (ref) imply
Denote the sample moments as
This section describes the first averaging estimation method. It follows the averaging GMM framework in \citetalias{ChengLiaoShi2019}. It contributes to provide two types of additional moments and apply the averaging estimation method to the instrument variables quantile regression.
\citetalias{ChengLiaoShi2019} define a “conservative” GMM estimator based on a set of valid moments and an “aggressive” GMM estimator based on the set of moments that combines the valid moments and additional, possibly misspecified moments. \citetalias{ChengLiaoShi2019} average between these two GMM estimators and show that under certain conditions the averaging GMM estimator uniformly dominates the conservative GMM estimator.
To apply the \citetalias{ChengLiaoShi2019} averaging GMM framework to the IVQR model, I use deCastroGalvaoKaplanLiu2019's (deCastroGalvaoKaplanLiu2019) smoothed two-step GMM method for computation to obtain the conservative GMM estimator. It replaces the IVQR moment functions
with smoothed moment functions
where $h_n$ is the sequence of smoothing bandwidth and $\tilde{I}(\cdot)$ is the same smoothed indicator function as defined in deCastroGalvaoKaplanLiu2019. The conservative GMM estimator is thus
where $\check{\underline{\boldsymbol{\mathbf{\Sigma}}}}_{1}$ is some consistent estimator of the (long-run) variance of the IVQR sample moments from the first step. With iid sampling,
where $\check{\boldsymbol{\mathbf{\beta}}}$ is some initial consistent estimator of $\boldsymbol{\mathbf{\beta}}_{0\tau}$. Specifically, my $\check{\boldsymbol{\mathbf{\beta}}}$ is deCastroGalvaoKaplanLiu2019's (deCastroGalvaoKaplanLiu2019) method of moments estimator with instrument vector equal to the linear projection of $\boldsymbol{\mathbf{X}}$ onto $\boldsymbol{\mathbf{Z}}$.
In addition to the conservative moments $\operatorname{E}[\boldsymbol{\mathbf{g}}_{1i}(\boldsymbol{\mathbf{\beta}},\tau)]=\boldsymbol{\mathbf{0}}_{d_Z\times1}$, we have “additional moments” based on $\boldsymbol{\mathbf{g}}^{*}( \boldsymbol{\mathbf{\beta}}, \tau)$ that might or might not be valid. If the additional moments are valid (i.e., $\boldsymbol{\mathbf{0}}_{r*} = \operatorname{E}[ \boldsymbol{\mathbf{g}}^{*}(\boldsymbol{\mathbf{\beta}}_{0\tau}, \tau) ]$), then adding additional valid information to estimation will reduce variance and improve efficiency of the IVQR estimator. If the additional moments are misspecified (i.e., $\boldsymbol{\mathbf{0}}_{r*} \ne \operatorname{E}[ \boldsymbol{\mathbf{g}}^{*}(\boldsymbol{\mathbf{\beta}}_{0\tau}, \tau) ]$), then combining these invalid additional moments with the original valid IVQR moments will result in a biased aggressive GMM estimator. However, the misspecified moments could still be helpful, if as a result the aggressive GMM estimator has a large reduction in variance, and an overall reduction in MSE.
In principle, the averaging estimator can always be at least as good as the conservative IVQR estimator, if it puts zero weight on the aggressive estimator when the additional moments are misspecified severely enough. In practice, this desirable result may not exactly hold due to the estimation error of the empirical weight in finite samples.
I propose two different types of additional moments. The first type is the conventional QR moments. If the structural model is a linear conditional $\tau$-quantile function (CQF), then
This conditional quantile restriction can be rewritten as a conditional expectation,
Using the law of iterated expectations, (ref) implies certain unconditional QR moments that are also the first-order condition of the population minimization problem of the expectation of the check function. Writing these unconditional QR moments in two separate parts,
I use (ref) but not (ref) for the additional moments. The first part (ref) is already contained in the IVQR population moments (ref), therefore the “conservative moments,” as the exogenous regressors $\boldsymbol{\mathbf{X}}_{exog}$ are contained in the full instruments $\boldsymbol{\mathbf{Z}}$. I use the second part (ref), the QR moments with potentially endogenous regressors $\boldsymbol{\mathbf{D}}$, as the “additional moments” to compute the “aggressive moments,” the “aggressive estimator,” and the “averaging estimator.” (For computation, I again use a smoothed version of the moment function, like (ref) but with $\boldsymbol{\mathbf{D}}_i$ replacing $\boldsymbol{\mathbf{Z}}_i$.) I call this the IVQR-QR type of averaging.
The motivation for using the conventional QR moments as the additional moments is that when there is little endogeneity (i.e., the additional QR moments are only slightly misspecified), the QR estimator is only a little biased. Meanwhile, the QR estimator usually has lower variance than the IVQR estimator. Therefore the aggressive estimator has a lower variance than the IVQR estimator. When the DGP has severe endogeneity (i.e., the additional QR moments are severely misspecified), the QR estimator and aggressive estimator have larger bias than the IVQR estimator. Ideally, more weight is put on the aggressive estimator when there is little endogeneity, and more weight is put on the conservative IVQR estimator when there is much endogeneity.
As an alternative to the QR type of additional moments, I propose using the 2SLS slope moment functions
where $\boldsymbol{\mathbf{Z}}_{-1}$ is the instruments without the intercept term, i.e., $\boldsymbol{\mathbf{Z}}=(1,\boldsymbol{\mathbf{Z}}_{-1})$. We use the demeaned instruments $(\boldsymbol{\mathbf{Z}}_{-1,i} - \bar{\boldsymbol{\mathbf{Z}}}_{-1})$ in these additional moments, where $\bar{\boldsymbol{\mathbf{Z}}}_{-1} \equiv 1/n \sum_{i=1}^{n} \boldsymbol{\mathbf{Z}}_{-1,i}$, to represent the 2SLS slope condition $\operatorname{Cov}(\boldsymbol{\mathbf{Z}}_{-1}, Y-\boldsymbol{\mathbf{X}}'\boldsymbol{\mathbf{\beta}}_0)=\boldsymbol{\mathbf{0}}$. With the 2SLS slope moments as the additional moments, we obtain the corresponding aggressive and averaging estimators. I call this the IVQR-2SLS type of averaging.
The motivation for using the 2SLS slope moments as the additional moments is that when there is not much heterogeneity across quantiles, 2SLS and IVQR at any quantile have similar slope (but not intercept) estimates, thus similar bias. Moreover, the smoothed IVQR estimator of KaplanSun2017 has slope estimates approach the 2SLS slope estimates as the smoothing bandwidth goes to infinity (see their \S2.2). Meanwhile, the 2SLS estimator usually has smaller variance than IVQR estimator, especially at quantile levels that are away from the median and closer to the tails. \footnote{One exception to 2SLS having smaller variance is with fat-tailed error terms, which I investigate in the simulations in (ref).} For example, even in an empirically-based simulation with substantial heterogeneity (that causes 2SLS to be biased), Table 3 of KaplanSun2017 shows 2SLS to be more efficient than IVQR at four out of five quantile levels. Incorporating the additional 2SLS slope moments, the aggressive estimator usually has smaller variance than the conservative IVQR estimator. The averaging estimator improves efficiency over the IVQR estimator by putting more weight on the aggressive estimator when it has large enough reduction in variance, and putting more weight on the IVQR estimator when the 2SLS slope moments are severely misspecified and the bias increase overwhelms the variance reduction.
Besides the two types of additional moments proposed in this paper, the same idea suggests using IVQR slope moments with other quantile levels to be the additional moments. The intuition is that when there is no or little heterogeneity, there will not be much difference in estimates across quantiles. This is related to the $L$-estimation method in conventional quantile regression by KoenkerPortnoy1986.
Incorporating the additional moments (either the potentially endogenous QR moments or the 2SLS slope moments) to the IVQR moments, we obtain the aggressive moments. Define the aggressive GMM estimator as
where
The aggressive moments contain both the conservative moments and additional moments. The GMM weighting matrix $\check{\underline{\boldsymbol{\mathbf{\Sigma}}}}_{2}^{-1}$ is constructed in the same way as $\check{\underline{\boldsymbol{\mathbf{\Sigma}}}}_{1}^{-1}$, except with aggressive moments (using $\boldsymbol{\mathbf{g}}_2$) instead of conservative moments (using $\boldsymbol{\mathbf{g}}_1$). The subscript “2” denotes “aggressive” here.
Following \citetalias{ChengLiaoShi2019}, define the averaging GMM estimator as
The empirical averaging weight $\hat{w}$ in \citetalias{ChengLiaoShi2019} is the sample analog of the optimal weight:
where
The middle part $\check{\underline{\boldsymbol{\mathbf{\Sigma}}}}_{k}$ is the estimator of the covariance matrix of the conservative sample moments and aggressive sample moments for $k=1$ and $k=2$, respectively; its inverse $\check{\underline{\boldsymbol{\mathbf{\Sigma}}}}^{-1}_{k}$ is the efficient two-step GMM weighting matrix. Define the estimator of the Jacobian matrix of the conservative moments and aggressive moments ($k=1$ and $k=2$, respectively) as
Both $ \check{\underline{\boldsymbol{\mathbf{\Sigma}}}}_{k}$ and $\hat{\underline{\boldsymbol{\mathbf{G}}}}_{k}$, for $k=1$ and $k=2$, are evaluated at the conservative GMM estimator $\hat{\boldsymbol{\mathbf{\beta}}}_1$. Therefore, they are consistent regardless of misspecification of the additional moments. The diagonal matrix $\underline{\boldsymbol{\mathbf{\Upsilon}}} $ measures how much we weight each element in the parameter vector in the (scaled) loss function $(\hat{\boldsymbol{\mathbf{\beta}}}-\boldsymbol{\mathbf{\beta}}_{0\tau})'\underline{\boldsymbol{\mathbf{\Upsilon}}}(\hat{\boldsymbol{\mathbf{\beta}}}-\boldsymbol{\mathbf{\beta}}_{0\tau})$, as in (3.9) of \citetalias{ChengLiaoShi2019}. When $\underline{\boldsymbol{\mathbf{\Upsilon}}}$ is the identity matrix, the expected loss (i.e., risk) becomes the sum of MSEs of each component of the estimator vector, $\sum_{j=1}^{d_X} \operatorname{E}[(\hat\beta_j-\beta_{0\tau j})^2]$.
Averaging estimation can be considered as a bias--variance tradeoff. The averaging weights are crucial for good performance of the averaging estimator. Too much weight on the aggressive estimator can result in large squared bias and large MSE. Too little weight on the aggressive estimator can result in large variance and again large MSE. The optimal averaging weight ideally balances the variance and the squared bias, minimizing MSE. This is like the bandwidth choice for kernel regression and other nonparametric estimators.
The \citetalias{ChengLiaoShi2019} empirical weight formula in (ref) requires estimation of the conservative and aggressive parameters, covariance matrix, and (population) Jacobian matrix, which involves a conditional density that must be nonparametrically estimated. The performance of the averaging GMM method heavily depends on whether the empirical weight is estimated accurately or not.
To estimate the parameters and the covariance matrix with the smoothed GMM approach of deCastroGalvaoKaplanLiu2019, I use the smallest possible smoothing bandwidth, for two reasons. First, KaplanSun2017 note smoothing can reduce MSE. \footnote{\Citet{KaplanSun2017} derive an MSE-optimal bandwidth for estimating the smoothed estimating equation IVQR estimator. This bandwidth is typically much larger than the smallest possible smoothing that makes computation feasible.} However, the goal of this paper is to demonstrate that it is the averaging method, instead of smoothing, that can improve estimation efficiency. Second, the \citetalias{ChengLiaoShi2019} averaging GMM framework assumes the conservative moments are valid and that the conservative estimator is not biased, but smoothing introduces some bias. Using the smallest possible smoothing bandwidth makes the bias of the conservative IVQR estimator small enough to be negligible.
For IVQR and even QR, the population Jacobian matrix involves a conditional PDF, which is commonly estimated by a nonparametric kernel estimator. The usual kernel estimator is actually the same as the standard sample Jacobian when $\hat{\boldsymbol{\mathbf{\beta}}}$ is based on smoothed moments.
To precisely estimate the Jacobian matrix of the IVQR moments and of the aggressive moments, I modify Kato2012's (Kato2012) bandwidth. He provides the optimal bandwidth for conventional QR based on asymptotic mean squared error (AMSE). He also provides a simplified version assuming independent, standard normal regression errors. I extend his Gaussian plug-in bandwidth to allow for any error variance, and I adapt his formulas for QR to IVQR. With QR, the vector $\boldsymbol{\mathbf{X}}$ acts as both the regressors and the instruments; for IVQR, the parts of the bandwidth formulas where $\boldsymbol{\mathbf{X}}$ acts as instruments are replaced by $\boldsymbol{\mathbf{Z}}$. For the QR Gaussian plug-in, I build on Kato2012's (Kato2012) results to prove in (ref) that in the model $Y=\boldsymbol{\mathbf{X}}' \boldsymbol{\mathbf{\beta}}_0 + U$ with $U \mid \boldsymbol{\mathbf{X}} \sim \mathrm{N}( \mu , \sigma^2 )$ and $\operatorname{Q}_{\tau}(U \mid \boldsymbol{\mathbf{X}})=0$, the sample analog of the AMSE-optimal bandwidth is
where $\Phi(\cdot)$ and $\phi(\cdot)$ are the standard normal CDF and PDF. Details are in (ref).
In addition to the theoretically-based averaging GMM estimation in (ref), I also propose a bootstrap averaging estimator for IVQR. Simulation performance of this bootstrap averaging estimator is in (ref).
The bootstrap averaging estimator comes from the same motivating idea in (ref). That is, compared to the IVQR estimator, the QR or 2SLS estimator might have smaller variance, and overall smaller MSE, though larger bias. This is especially true with only mild endogeneity or heterogeneity.
There are some differences with (ref) that may enable the bootstrap's better performance in simulations. (ref) considers averaging between two estimators (conservative and aggressive), whereas the bootstrap method averages among three estimators: IVQR, 2SLS, and QR. Computationally, bootstrap averaging is simpler and easier, not requiring two-step GMM with a large degree of overidentification.
The bootstrap method algorithm is as follows.
The computation time for the bootstrap averaging estimator is reasonably fast even with a large weight grid size like $\num[round-mode=off,group-digits=integer]{13701}$. The actual IVQR, 2SLS, and QR estimators only need to be computed once per bootstrap draw; the $\num[round-mode=off,group-digits=integer]{13701}$ is just arithmetic. In my implementation, $\hat{\boldsymbol{\mathbf{\beta}}}_{\mathrm{IVQR}}$ and $\hat{\boldsymbol{\mathbf{\beta}}}_{\mathrm{IVQR}}^{(b)}$ are obtained by first projecting the regressors onto the instruments, and then solving smoothed versions of the exactly-identified equations, using standard numerical methods. So the additional computation time of bootstrap averaging over GMM averaging is not too big. \footnote{For example, in simulation model 1, running 400 replications with 250 bootstrap draws per replication to compute both the GMM averaging estimator and the bootstrap averaging estimator takes 2.75 times longer than running 400 replications with only GMM averaging.}
This section reports simulation results showing the finite-sample performance of the two types of GMM averaging estimator and the bootstrap averaging estimator, relative to the IVQR estimator.
I consider three different simulation models, with many DGPs within each model. Simulation model 1 presents a case where the uniform dominance condition does not hold. Simulation models 2 and 3 present cases where all the averaging estimators uniformly dominate the IVQR estimator. Simulation model 2 closely follows the \citetalias{ChengLiaoShi2019} simulation model (S2), but with modification to IVQR. Since one important feature of quantile regression is to capture unobserved heterogeneity, simulation model 3 includes slope heterogeneity.
The performance of the averaging estimators are measured by their robust root mean squared error (robust RMSE) relative to that of the IVQR estimator. The robust RMSE is computed by replacing the bias with median bias and replacing the standard deviation with interquartile range (IQR) divided by 1.349. It equals RMSE for normal distributions but is more robust to outliers. More specifically, robust RMSE (rRMSE) is computed as
where the median bias is the median of estimators among the $M$ replications minus the true parameter value,
the IQR is the difference between the $0.75$-quantile and $0.25$-quantile of the estimators among the $M$ replications,
$M$ is the number of replications in total, and $d_\theta$ is the number of parameters.
I normalize the rRMSE of the IVQR estimator to 1 and use the relative rRMSE to see how other estimators perform relative to the IVQR estimator. That is, I divide all rRMSEs by the IVQR rRMSE to get the relative rRMSE. If the relative rRMSE is above (below) 1, that means the estimator has larger (smaller) rRMSE than IVQR, i.e., it performs worse (better) than IVQR. Ideally, an estimator has relative rRMSE (weakly) below 1 in all different DGPs. In this case, we say the proposed estimator uniformly dominates the IVQR estimator: regardless of the true DGP, it performs as good or better than IVQR. Such an estimator is unambiguously preferred to IVQR. An estimator may still be preferred even without uniform dominance, but it would depend on the user's preferences (e.g., the estimator's Bayes risk may be larger or smaller than IVQR's depending on the user's prior over DGPs).
This Job Training Partnership Act-based simulation DGP generalizes DGP 1 in deCastroGalvaoKaplanLiu2019 to a class of DGPs that allow different combinations of endogeneity, heterogeneity, and fat-tail levels. Consider a structural random coefficient model that describes the impact of a job training program $D_i$ on individual $i$'s earnings $Y_i$,
The unobserved scalar $U_i \sim \textrm{Unif}(0,1)$ incorporates other earnings determinants like ability. The individual-specific intercept and slope depend on $U_i$, through functions $\beta(\cdot)$ and $\gamma(\cdot)$.
The slope function is set as $\gamma(U_i)=100c_2 U_i^4$. The nonnegative constant $c_2$ indicates the degree of treatment effect heterogeneity.
The intercept function is set as $\beta(U_i)=60+Q(U_i)$, with two possible $Q(\cdot)$. First, $Q(U_i)$ follows a $\chi^2_3$ distribution. Second, $Q(U_i)$ follows a $t$-distribution with $c_3$ degrees of freedom. Each $c_3$ value represents a different fat-tail degree of earnings. As $c_3$ increases, the earnings distribution becomes less fat-tailed, approaching a normal distribution.
The job training offer (eligibility) $Z_i$ is completely randomized with $\operatorname{P}(Z_i=1)=\operatorname{P}(Z_i=0)=1/2$. This $Z_i$ is a valid instrument for $D_i$.
The relationship between the randomized offer and the self-selection of participation in the program is described as a conditional probability,
where $D_i=1$ if the individual actually takes the training and $D_i=0$ if not. The constant $c_1\in[0,1]$ indicates the endogeneity level of $D_i$, with larger $c_1$ meaning more endogeneity.
Altogether I ran 242 DGPs (combinations of $c_1$, $c_2$, and $c_3$) in this simulation model 1. Since many results are very similar, I selected 14 DGPs representative of different combinations of endogeneity, heterogeneity, and fat-tail level. More details are in (ref).
(ref) presents the finite-sample rRMSE of the estimators proposed in this paper relative to that of the IVQR estimator at the median, $\tau=0.5$.
(ref) shows no averaging estimator uniformly dominates IVQR in simulation model 1. That is, no column has relative rRMSE below $1$ for each DGP. But, they do not suffer much in the least favorable cases, even when 2SLS and/or QR (or their aggressive GMM counterparts) have relative rRMSE well above 1.
As expected, the IVQR-2SLS averaging estimator performs better than (or almost the same as) the IVQR estimator in the DGPs with no slope heterogeneity. The efficiency gain is not much (up to 5%). In the DGPs with slope heterogeneity, the relative rRMSE of IVQR-2SLS averaging is around $1.06$ to $1.13$, while that of IVQR-2SLS aggressive estimator and 2SLS estimator is around 5 to 7. This indicates that although IVQR-2SLS averaging estimator is worse than IVQR, its rRMSE is much closer to 1 compared with the 2SLS estimator and aggressive estimator, due to putting most of the weight on the conservative estimator in these cases.
The IVQR-QR averaging estimator performs around 11--23% better than IVQR estimator in the DGPs with no endogeneity and no heterogeneity. As the endogeneity level of treatment variable increase, the QR moments become more misspecified, and the performance of the IVQR-QR averaging estimator becomes less favorable. In the DGPs with much endogeneity, the IVQR-QR averaging estimator performs worse than the IVQR estimator by around 3--13%. Compared with the relative rRMSE of the QR estimator and that of the IVQR-QR aggressive estimator, the averaging estimator is much closer to 1 and saved from even worse performance by putting most of the weight on the conservative estimator.
The bootstrap estimator that averages among IVQR, 2SLS, and QR performs better than IVQR in all but one DGP in which either 2SLS or QR (or both) has relative rRMSE below 1. In the DGPs in which both 2SLS and QR are worse than IVQR, the bootstrap averaging estimator performs worse than IVQR, but not by as much. For example, in DGPs 7, 8, 13, and 14, the bootstrap averaging estimator has relative rRMSE from 1.06 to 1.20, compared to 1.27 to 29.68 for 2SLS and 2.56 to 3.53 for QR.
Results at other quantiles are shown in (ref). It includes results at $\tau=0.2$ up to $\tau=0.8$.
For IVQR-2SLS averaging, there is a “magic quantile” at which 2SLS performs well across all DGPs, regardless of heterogeneity. In simulation model 1, the slope is $\gamma(U)=100 c_2 U^4$ with $U \sim \textrm{Unif}(0,1)$. The 2SLS population slope is $\operatorname{E}[100 c_2 U^4]=100c_2\operatorname{E}(U^4)=20c_2$, where $\operatorname{E}(U^4)=0.2$ is the fourth moment of a standard uniform distribution. The $\tau$-IVQR slope is $100c_2\tau^4$; since $0.7^4=0.24$, the slope is very close to the 2SLS slope when $\tau=0.7$ ($24c_2$ vs.\ $20c_2$). At a slightly smaller $\tau$ (not run in simulations), the IVQR and 2SLS population slopes are identical. So among $\tau=0.2,0.3,\ldots,0.8$, the “magic quantile” is $\tau=0.7$ in simulation model 1. This is true regardless of $c_2$; it depends only on how close $\tau^4$ is to $\operatorname{E}(U^4)$. Simulation results confirm that at the $0.7$-quantile, 2SLS has relative rRMSE much below 1 even in the DGPs 5--9 with some or much slope heterogeneity. The IVQR-2SLS averaging estimator also has relative rRMSE below 1 in all the DGPs.
As $\tau$ moves away from the “magic quantile,” the 2SLS estimator begins to show the pattern that its relative rRMSE is below 1 in DGPs with no slope heterogeneity, but above 1 in the DGPs with some or much slope heterogeneity. In DGPs 5--9 with slope heterogeneity, the IVQR-2SLS aggressive estimator's relative rRMSE can be as large as 12, but that of averaging estimator is still close to 1 (bounded by 1.19). The averaging estimator's rRMSE is less than 20% worse than IVQR in the least favorable situation.
For the IVQR-QR averaging estimator, the patterns at other quantiles are almost the same as at the median.
(ref) presents the bootstrap averaging results at different quantiles. Although bootstrap averaging does not uniformly dominate the conservative IVQR estimator, it offers significant improvement in a variety of DGPs with relatively small downside. Of the $98$ relative rRMSEs reported in the BS columns in (ref), only five are $1.15$ or above, with the largest being $1.29$, compared to $30$ relative rRMSEs below $0.85$.
This simulation model is close to simulation DGP 2 in \citetalias{ChengLiaoShi2019}. The \citetalias{ChengLiaoShi2019} model considers endogenous regressors at a fixed endogeneity level with a set of valid instruments, while introducing potentially invalid instruments in the additional moments with varying degrees of endogeneity. Since my paper proposes using the IVQR-2SLS and IVQR-QR types of additional moments, no additional instruments are involved. Instead of varying the endogeneity of additional instruments, I let the regressors' endogeneity level vary across DGPs. Otherwise, the DGPs are the same as in \citetalias{ChengLiaoShi2019}. This simulation model 2 has enough continuous endogenous regressors and instruments that uniform dominance holds.
Consider the linear model
where $\theta_1=\theta_2=\cdots=\theta_6=2.5$ and $\theta_0=1 $; $Y$ is the dependent variable; and there are six endogenous regressors $(X_{1}, \ldots, X_{6})$ and twelve excluded valid instruments $(Z_{1}, \ldots, Z_{12})$. The data $(Y_i, X_{1,i}, \ldots, X_{6,i}, Z_{1,i}, \ldots, Z_{12,i})$ are sampled iid for $i=1,\ldots,n$. The regressors are generated by
The $(Z_{1}, \ldots, Z_{12}, \epsilon_1, \ldots, \epsilon_6, u) $ are generated from a multivariate normal distribution with mean zero and covariance matrix $diag(\underline{\boldsymbol{\mathbf{I}}}_{12\times12}, \underline{\boldsymbol{\mathbf{\Sigma}}}_{7\times7})$, where
The $\underline{\boldsymbol{\mathbf{I}}}_{6\times6}$ and $\underline{\boldsymbol{\mathbf{I}}}_{12\times12}$ are the identity matrix. The $\boldsymbol{\mathbf{1}}_{6\times1}$ and $\boldsymbol{\mathbf{1}}_{1\times6}$ are the column and row vector of ones, respectively. In each DGP, $c_0$ is a fixed constant. It measures how much the regressors and the structural error are correlated. For $\underline{\boldsymbol{\mathbf{\Sigma}}}_{7\times7}$ to be positive definite it requires $c_0\le 0.4$. Thus, let $c_0$ take value in $\{0, 0.05, 0.1, \ldots, 0.4 \} $. The higher the $c_0$ value, the more endogenous the regressors. In the special case $c_0=0$, the regressors are generated independently from the structural error. The simulation results demonstrate that in the case $c_0=0.4$, the regressors are endogenous enough that the QR estimator performs much worse than IVQR estimator.
In this data-generating process, the error term $u$ follows a Gaussian distribution. Then, I make a location shift to make its $\tau$-quantile to be zero. This shift is to make the model to satisfy the conditional IVQR restriction that conditional on instruments, the $\tau$-quantile of error term is zero.
The conservative IVQR moment functions use the twelve instruments,
The additional moment functions for IVQR-QR averaging are
The additional moment functions for IVQR-2SLS averaging are
The IVQR estimator is the GMM estimator from the 12 moments using (ref). The IVQR-QR aggressive estimator is the GMM estimator from the 18 moments using (ref). The IVQR-2SLS aggressive estimator is the GMM estimator from the 24 moments using (ref).
In simulation model 2, all three averaging estimators uniformly dominate the IVQR estimator. (ref) shows these three averaging estimators' relative rRMSE is bounded below 1 in a class of DGPs with different endogeneity level. (ref) reports the rRMSE of various estimators relative to that of the IVQR estimator. The second column reports the value of $c_0$, which indicates how much regressor endogeneity there is in the DGP. Columns 3, 6, and 9 are the three averaging estimators proposed in this paper. Columns 4 and 5 are the IVQR-2SLS aggressive estimator and the 2SLS estimator for reference. Columns 7 and 8 are the IVQR-QR aggressive estimator and the QR estimator for reference. Column 10 is the smoothed estimating equations (SEE) estimator proposed by KaplanSun2017, using their code's plug-in bandwidth.
In simulation model 2, the regressors' coefficients are set to be fixed. There is no slope heterogeneity. The 2SLS estimator performs uniformly better than the IVQR estimator, in the sense that its relative rRMSE is less than 1 across all DGPs. The IVQR-2SLS aggressive estimator and averaging estimator both have relative rRMSE less than 1 and uniformly dominate the IVQR estimator.
The more endogeneity of regressors in a DGP, the more misspecified the QR moments, and the worse the performance of QR. In DGPs 1--3 with the least endogeneity, the QR estimator performs better than the IVQR estimator. In DGPs 4--9 with more endogeneity, QR is increasingly worse than IVQR. The IVQR-QR aggressive estimator follows the same pattern as the QR estimator. However, the IVQR-QR averaging estimator has relative rRMSE less than 1 in all DGPs, even in the DGPs where the QR moments are very misspecified and the QR and IVQR-QR aggressive estimators are much worse than the IVQR estimator. This shows the empirical optimal weight formula works as desired, putting more weight on the aggressive estimator when the additional moments reduce variance more relative to the increased bias, and putting more weight on the conservative estimator when the additional moments are severely misspecified.
The bootstrap averaging estimator performs best of all. It not only uniformly dominates the IVQR estimator, up to 45% efficiency gain, but also is uniformly better than either type of GMM averaging estimator. Moreover, the bootstrap averaging estimator usually performs better than the SEE estimator, although not uniformly. As there is no slope heterogeneity in simulation model 2, the IVQR-2SLS averaging estimator's relative rRMSE is always below 1, mostly around 0.9. The IVQR-QR averaging estimator's relative rRMSE is also below 1 in all DGPs, increasing from 0.63 to 0.97 as $c_0$ increases. The bootstrap averaging estimator has even smaller relative rRMSE, ranging from 0.55 to 0.84 as $c_0$ increases.
The simulation results at other quantiles are reported in (ref). The results share the same pattern as that at the median. All three averaging estimators uniformly dominate the IVQR estimator at quantiles from $\tau=0.2$ to $\tau=0.8$.
The uniform dominance in this fixed-coefficient model with multiple endogenous regressors also holds with a non-Gaussian error term. To illustrate this, I set the error term to follow a chi-square distribution with $4$ degrees of freedom. More specifically, I first generate the error term in the same way as in (ref). Second, I transform the error term to $u^{*}= F_{\chi_4^2 }^{-1}(\Phi(u))$ where $\Phi(\cdot)$ is the CDF of the original Gaussian error term $u$ and $F_{\chi_4^2}^{-1}$ is the inverse CDF of a $\chi^2_4$ distribution. Finally, I shift $u^*$ to have its $\tau$-quantile equal zero.
The simulation results have similar patterns as in the Gaussian error case; see the results tables and figures in (ref). All three averaging estimators have relative rRMSE less than 1 in all the DGPs at all quantiles, expect three cases of IVQR-2SLS averaging at $\tau=0.2$ having relative rRMSE between 1.01 and 1.03. However, these three rare cases seem due to simulation error since their relative rRMSEs are less than or equal to 1 when running more simulation replications.
Seeing slope heterogeneity across quantiles is part of the value of quantile regression. Simulation model 3 extends simulation model 2 to allow for slope heterogeneity.
Consider a linear model
with $\theta_0=1 $ and random coefficients for individual $i$ equal to
where $F(\cdot)$ is the CDF of $u$. The slope term (ref) is set to be a function of the rank of the error term in its distribution. It represents the slope heterogeneity feature. The term “$hetero$” is a fixed constant in a DGP, the same as $c_2$ in simulation model 1. When $hetero=0$, the slope is a constant zero, which means the model has no slope heterogeneity. The larger the “$hetero$” value, the more slope heterogeneity in the DGP. In simulation model 3, $hetero$ takes value in $\{0,0.1,\ldots,1\}$. I consider both Gaussian and non-Gaussian errors. The non-Gaussian error follows the same transformation as in (ref). The error term has the same location shift as in (ref) to make its $\tau$-quantile equal zero.
As in simulation model 2, there are six endogenous regressors $(X_{1}, \ldots, X_{6})$ and twelve valid instruments $(Z_{1}, \ldots, Z_{12})$, and sampling is iid. The regressors are generated by
The monotonicity condition in IVQR requires that the structural quantile function is increasing in $u$ given any $\boldsymbol{\mathbf{X}}=\boldsymbol{\mathbf{x}}\in\mathcal{X}$ ChernozhukovHansen2005. For the monotonicity condition to hold in this simulation model, I restrict the regressors to have non-negative support. More specifically, I shift each regressor to the right by 3.1 times its standard deviation. There is less than 0.001 probability that a regressor remains negative, in which case I set it to zero.
As in simulation model 2, the $(Z_{1}, \ldots, Z_{12}, \epsilon_1, \ldots, \epsilon_6, u) $ are generated from a multivariate normal distribution with mean zero and covariance matrix $diag(\underline{\boldsymbol{\mathbf{I}}}_{12\times12}, \underline{\boldsymbol{\mathbf{\Sigma}}}_{7\times7})$, where
Since $c_0 \in \{0, 0.05, 0.1, \ldots, 0.4 \}$, there are $11\times 9$ DGP combinations of $(hetero, c_0)$.
I simulate the rRMSE as in (ref), where the true population parameters are $\theta_j=hetero \times \tau^4$ for all $j=1,\ldots,6$.
In (ref), I fix the endogeneity level indicator $c_0$ at zero, 0.2, or 0.4 (highest value), and vary the slope heterogeneity across all values $hetero=0,0.1,\ldots,1$. Similarly, I fix $hetero$ at zero, 0.5, or 1 (highest value), and vary the endogeneity level across all values $c_0=0,0.05,\ldots,0.4$. Altogether this includes 60 DGPs, covering the boundaries of the 99 DGPs and indicating the patterns. (ref) reports the upper and lower bounds of the relative rRMSE of the three averaging estimators in these six cases (i.e., 3 fixed endogeneity levels, 3 fixed heterogeneity levels).
(ref) summarizes results at the median in these six cases. More detailed results tables are included in the supplemental appendix. We can see that all three averaging estimators have relative rRMSE lower bounds strictly less than 1. The bootstrap averaging estimator has the smallest lower bound in five cases and is essentially tied for smallest (up to simulation error) in the sixth. With no endogeneity and varying heterogeneity, IVQR-QR and bootstrap averaging both have upper bound strictly below 1, while the IVQR-2SLS averaging estimator has upper bound roughly 1 (up to simulation error). With no heterogeneity and varying endogeneity, IVQR-2SLS and bootstrap averaging both have upper bound strictly below 1, while IVQR-QR averaging has upper bound roughly 1. In the other four cases, the three averaging estimators all have upper bound roughly 1. Although some exceed 1 (largest value 1.024), this is believed to be entirely due to simulation error: all the upper bounds reduce to $1.000\pm0.001$ with a larger number of simulation replications.
(ref) is similar to (ref), but it visualizes results for all 60 DGPs (not just bounds). It shows that in simulation model 3, all three averaging estimators uniformly dominate the IVQR estimator. Using any of these three averaging estimators would provide more precise estimation than the IVQR estimator.
(ref) includes results like (ref) at quantiles ranging from $\tau=0.2$ to $\tau=0.8$. The supplemental appendix provides yet more detailed tables for these cases.
For some quantile levels, for some DGPs, the bootstrap averaging estimator's relative rRMSE is above 1, but 1.075 is the maximum among 420 values (i.e., 60 DGPs with 7 values of $\tau$ each). Again, this is partly due to simulation error. The relative rRMSE exceeds 1.05 only 19 times out of 420. In 96 out of 420 DGPs, the relative rRMSE is between 1 to 1.04. In the other 305 out of 420 cases, the bootstrap averaging estimator has relative rRMSE less than 1. At every quantile, the bootstrap averaging estimator's relative rRMSE can reach as low as 0.79, and at some quantiles even 0.574.
The QR estimator and the IVQR-QR averaging estimator each have similar patterns across quantiles in the 60 DGPs shown in the table in the supplemental appendix. One finding is that in the much-endogeneity case (i.e., $c_0\ge0.35$), the IVQR-QR aggressive estimator has a computation problem and performs very poorly. The IVQR-QR averaging estimator, however, still has relative rRMSE of 1. This indicates the empirical weight puts almost all weight on the conservative IVQR estimator and zero weight on the IVQR-QR aggressive estimator. The averaging method works in a desirable way. The IVQR-QR averaging estimators all show uniform dominance over the IVQR estimator.
The “magic quantile” phenomenon for 2SLS and IVQR-2SLS averaging again applies in simulation model 3, as seen when looking at the results across quantiles. In simulation model 3, the slope is $\theta(u_i)=hetero \times [F(u_i)]^4$ where $u \sim N(0,1)$ and $F(\cdot)$ is its CDF. The 2SLS population slope is $\operatorname{E}[hetero \times [F(u_i)]^4 ]$. Since $F(u_i)\sim\textrm{Unif}(0,1)$, this equals $hetero$ times the fourth moment of a standard uniform distribution, which is 0.2. The IVQR slope is $hetero \times \tau^4$. Since $0.68^4=0.20$, the 2SLS and IVQR slopes are equal when $\tau=0.68$. For $\tau$ near $0.68$, even with much slope heterogeneity across units, the 2SLS and IVQR slopes are similar. As $\tau$ gets farther from $0.68$, the IVQR slope increasingly differs from the 2SLS slope, with the rate of increase depending on $hetero$.
The simulation results show that at $\tau=0.6,0.7,0.8$, the 2SLS estimator has relative rRMSE less than 1 uniformly in all 60 DGPs, even in the DGPs with high $hetero$ values. Correspondingly, the bootstrap averaging estimator is much lower than 1 at these three quantiles. (ref) shows the performance of the three averaging estimators at $\tau=0.7$. The bootstrap averaging estimator's relative rRMSE is significantly below 1, and much below that of the IVQR-2SLS averaging estimator. This is probably because the bootstrap method averages the 2SLS estimator directly, rather than through aggressive GMM that may not weight the 2SLS slope moments as heavily as is optimal.
Compared with the “magic tau” results in simulation model 1, in which the 2SLS and IVQR-2SLS averaging estimators have relative rRMSE much lower than 1 at quantile $\tau=0.7$ only, here in simulation model 3, the good performance of the 2SLS and IVQR-2SLS averaging estimators shows up at quantiles $\tau=0.6$, $\tau=0.7$, and $\tau=0.8$. The “magic tau” is the same (0.68) in these two simulated models. The difference in the performance at $\tau=0.6$ and $\tau=0.8$ is because the “magic tau” is only related with the (median) bias part of rRMSE. That is, at $\tau=0.7$, the 2SLS and IVQR slopes are still very close (almost no bias of 2SLS estimates). At $\tau=0.6$ and $\tau=0.8$, the 2SLS and IVQR slopes are somewhat close (little bias of 2SLS estimates). In one simulated model, the bias at $\tau=0.6$ or $\tau=0.8$ is still small compared to the (large) variance, whereas in another simulated model with smaller variance, the bias may start to dominate variance even at $\tau=0.6$ or $\tau=0.8$.
This paper contributes two averaging estimation methods to improve finite-sample efficiency in IVQR estimation. First, I implement the averaging GMM of ChengLiaoShi2019, proposing two types of additional moments based respectively on conventional QR and 2SLS, and considering other important implementation details such as bandwidths, extending a result of Kato2012 along the way. Second, I propose a new method that uses the bootstrap to estimate optimal weights for averaging IVQR, QR, and 2SLS estimators.
This paper provides simulation evidence that these three averaging estimators outperform the IVQR estimator across all kinds of DGPs in large models with multiple endogenous regressors, as well as across quantile levels from $\tau=0.2$ to $\tau=0.8$. Bootstrap averaging offers especially substantial efficiency gains.
Future work could involve developing theory for the practically successful bootstrap method, or investigating averaging across quantiles, or adding non-trivial smoothing into this averaging framework. Additionally, an IVQR application of ArmstrongKolesar2019 may better use additional-but-possibly-misspecified moments when the object of interest is a scalar instead of the full parameter vector.
\singlespacing