EconBase
← Back to paper

A first-stage representation for instrumental variables quantile regression

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

59,520 characters · 15 sections · 55 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A first-stage representation for instrumental variables quantile regression

abstractThis paper develops a first-stage linear regression representation for the instrumental variables (IV) quantile regression (QR) model. The quantile first-stage is analogous to the least squares case, i.e., a linear projection of the endogenous variables on the instruments and other exogenous covariates, with the difference that the QR case is a weighted projection. The weights are given by the conditional density function of the innovation term in the QR structural model, conditional on the endogeneous and exogenous covariates, and the instruments as well, at a given quantile. We also show that the required Jacobian identification conditions for IVQR models are embedded in the quantile first-stage. We then suggest inference procedures to evaluate the adequacy of instruments by evaluating their statistical significance using the first-stage result. The test is developed in an over-identification context, since consistent estimation of the weights for implementation of the first-stage requires at least one valid instrument to be available. Monte Carlo experiments provide numerical evidence that the proposed tests work as expected in terms of empirical size and power in finite samples. An empirical application illustrates that checking for the statistical significance of the instruments at different quantiles is important. The proposed procedures may be specially useful in QR since the instruments may be relevant at some quantiles but not at others. Keywords: Quantile regression, instrumental variables, first-stage. JEL: C13, C23.

\doublespacing

Introduction

Instrumental variables (IV) methods are one of the main workhorses to estimate causal relationships in empirical analysis. Standard IV regression methods stress that for instruments to be valid they must be exogenous and must be related to the endogenous variables. The latter condition is usually evaluated by using a first-stage auxiliary regression, where a linear model is used to make inference on the degree of association of the IV and the endogenous variables. While this is usually accepted as a valid procedure, its representation is in fact specific to the two-stage least squares (2SLS) model for mean models. This paper derives a first-stage representation for quantile regression (QR) models.

Several IV methods have been proposed in QR to solve endogeneity when the covariates are correlated with the error term in a regression model. CH04, CH05,CH06,CH08 (CH hereafter) develop an instrumental variables quantile regression (IVQR) procedure that has been applied in several contexts. It is one of the most prolific approaches in terms of subsequent work, as it provides a general procedure to use IV for endogeneity of regressors CHJ07, CHJ09, Galvao11, Chetverikov16. We refer to CH20 for an overview of IVQR.\footnote{There is also a more recent literature on GMM QR, see e.g., Firpoetal21 and references therein.}

CH comment that their method is a simple solution to a 2SLS analog.\footnote{This has been formally established in GalvaoMontes15.} However, the first-stage of the IVQR estimator has not been explicitly considered, as it is implemented as an inverse QR estimator. The IVQR estimator contrasts to alternative procedures where the first-stage is implemented. For instance, Amemiya82, Powell83, ChenPortnoy96, and KimMuller04 use an explicit first-stage that fits the endogenous variable(s) as a function of exogenous covariates and IV, and this is then plugged in a second-stage. Lee07 also adopts a two-step control-function approach where in first step consists of estimation of the residuals of the reduced-form equation for the endogenous explanatory variable. MaKoenker06 present an estimator for a recursive structural equation model.

This paper builds on the IVQR estimator and shows that a first-stage regression model can be explicitly recovered from the CH IVQR estimator. The first-stage IVQR (FS-IVQR) is a linear projection of the endogenous variables on the instruments and other exogenous variables, with the difference that the QR case is a weighted regression, that is, it has the representation of a weighted least squares (WLS) regression of the endogenous variable(s) on the IV and the exogenous regressors. The weights are given by the conditional density function of the innovation term in the QR structural model, conditional on the endogeneous and exogenous covariates together with the instruments, at a given quantile. This result provides a clear analogy between the first-stage in 2SLS and IVQR. The derivation of the result is simple. We write the IVQR model as a constrained Lagrangian optimization problem and show that one of the restrictions that must be satisfied is the analogue of the first-stage.

The CH IVQR method requires an identification condition that is based on the full-rank of the Jacobian for the exogeneity of the instruments. The lack of identification when the Jacobian is not full rank implies that estimating the parameters can be extremely difficult and the first-order asymptotics can be a poor guide of the actual sampling distributions Dufour97. In this paper, we show that a necessary condition for the Jacobian identification conditions for IVQR models are embedded in the quantile first-stage representation. Hence, the FS-IVQR representation is directly related to the Jacobian requirement of CH IVQR.

We propose a two-step FS-IVQR estimator. The practical implementation of the estimator is straightforward and as follows. First, from the IVQR one estimates the conditional density function at a selected quantile, which produces an estimate of the weights. The weighting factor is estimated from the IVQR errors using, for instance, sparsity or kernel methods (see, e.g., Koenker05). Second, a standard WLS regression is implemented by regressing the endogenous variable on the instruments and exogenous variables with weights from the first step -- this is parallel to the first-stage model used in 2SLS, but using weights. We derive the limiting distribution of the two-step FS-IVQR estimator and show that, under some standard regularity conditions, it is asymptotically normal.

The first-stage regression for conditional average models has been used as a natural framework to evaluate the validity of instruments since one can test for their statistical significance, that is, how the IV impact the endogenous variable(s). Based on the proposed FS-IVQR model, we suggest an analogous test procedure to assess the validity of the IV for given quantiles. In particular, we test the statistical significance of the FS-IVQR coefficients. This is a test for the Jacobian identification condition with the null hypothesis that the required rank condition is not satisfied -- the coefficients are equal to zero. A simple Wald statistic can be used to test this null, and we show that it has an asymptotic Chi-square distribution. Nevertheless, the implementation of the test requires consistent estimation of the weights, which in turn requires at least one valid available instrument. The requirement of consistent estimation of weights in the first-stage QR exposes a caveat with the IVQR model, and hence, we suggest practical use of the test in an over-identification context, that is, when at least one valid instrument is available. In spite of this issue, testing using the first-stage QR allows for a procedure in empirical work to evaluate the degree of association of the IV to the endogenous variable that is parallel to the standard first-stage in two-stage least squares (FS-2SLS).

The requirement of consistent estimation of weights in the QR first-stage leads to two important conclusions of this paper. First, it is difficult to derive an analogous F-statistic type rule-of-thumb for categorizing weak instruments as in, among others, Staiger97, Sanderson16, and Lee20, for ordinary least-squares (OLS) models (see StockYogo05 for an extensive discussion). A complete parallel testing procedure to evaluate the validity of IV in the QR case only works when one is able to estimate the structural parameters -- and consequently the weights -- consistently, which in turn requires at least one valid instrument. Therefore, dropping the valid instrument requirement is a very interesting line of future research for investigating the weak instruments problem. Second, an important conclusion for practical work is that weak identification robust inference procedures, as in CH08, Jun08, CHJ09, and AndrewsMikusheva16, are a very important avenue for empirical applications using QR instrumental variables models. We strongly suggest the use of the QR first-stage and testing proposed here along with the weak identification robust inference procedures in empirical applications.

One important feature of the procedure developed in this paper is that instruments could be statistically insignificant in FS-2SLS, but they could still be related to the endogenous variable in the IVQR set-up. The reason is that the FS-2SLS test only evaluates a mean effect, but the FS-IVQR, because of its specific weighting procedure, allows for different first-stage coefficients across quantiles. As a result, the IV could be relevant at some quantiles but not for the mean (and vice-versa), an issue that has been discussed in Chesher03 and subsequent literature. The test developed here thus allows inference on the validity of the IV for the exogeneity condition across quantiles, rather than only a mean effect.

We use a Monte Carlo exercise to evaluate the finite sample performance of the proposed tests. The tests have correct size in all cases studied, where the structural parameters can be consistently estimated under the null hypothesis. We consider alternative cases where there is no identification under the null. The tests have excellent power properties. In particular these experiments highlight the case where the FS-2SLS test for the mean-based model suggests the instrument is not valid, but the proposed FS-IVQR procedure finds it is for some quantiles.

As an empirical illustration, we apply the FS-IVQR estimator to the Card95 data on instrumenting education using college proximity. The analysis reveals heterogeneity in the significance of the IV across quantiles. In fact, while the 2SLS analysis shows that one instrument (proximity to 2-year college) is not statistically significant in the first-stage, it is indeed for high quantiles.

The paper is organized as follows. Section (ref) briefly reviews the CH IVQR estimator, rewrites that estimator as a constrained minimization problem and derives the first-stage representation for the IVQR. Then it shows that the FS-IVQR estimator is equivalent to the identification condition in CH. Section (ref) presents the first-stage test for validity of instruments. Section (ref) discusses its empirical implementation and derives the estimators' asymptotic distribution. Section (ref) provides finite sample Monte Carlo evidence. Section (ref) applies the proposed tests to an empirical problem. Finally, Section (ref) concludes.

A first-stage representation for IVQR

The IVQR estimator and its variants

Let $(y,d,x,z)$ be random variables, where $y$ is a scalar outcome of interest, $d$ is a $1\times r$ vector of endogenous variables, $x$ is a $1\times k$ vector of exogenous control variables, and $z$ is a $1\times p$ vector of exogenous instrumental variables, with $p\geq r$. Define $w=(x,z)$ and $s=(d,x,z)$.

CH06 developed estimation and inference for a generalization of the QR model with endogenous regressors. A linear representation of the model takes the following form

equation[equation omitted — 119 chars of source]

where $u_{d}$ is the nonseparable error or rank and the subscript indicates the endogenous covariates of the model. Under some regularity conditions, CH establish the following IV identification function

equation[equation omitted — 114 chars of source]

Although each parameter and estimator is indexed by the quantile $\tau \in (0,1)$, throughout the paper we will suppress the dependence on $\tau$.

The restriction in (ref) can be used to estimate the parameters of interest. For a given quantile $\tau$, the population IVQR estimator for model in (ref), is given by

equation*[equation* omitted — 70 chars of source]

where

equation*[equation* omitted — 151 chars of source]

and $\rho_\tau(u)=u(\tau-\bm 1(u<0))$ is the check function, and $\|\cdot\|_A=\cdot'A\cdot$ is the Euclidean distance for any positively definite matrix $A$ of dimension $p\times p$.

As noted by CH06, the IVQR estimator is asymptotically equivalent to a particular GMM estimator where the QR first order conditions are used as moment conditions. In particular, it would involve a Z-estimator solving

align[align omitted — 209 chars of source]

where $\bm{1}(\cdot)$ is the indicator function. Here $\bm{0}_k$ and $\bm{0}_p$ are null vectors with dimensions $k\times 1$ and $p\times 1$, respectively.

Different estimators have been proposed in the GMM framework based on identifying the structural parameters from equations (ref)--(ref). Kaplan17, ChenLee18 and dCGKL19 provide general estimation procedures based on smoothing techniques of the non-differentiable indicator function. However, the constructed estimator differs from the CH IVQR one. This can be seen in the fact that the term $z\gamma$ is not considered altogether from the regression model. Our procedure follows the CH estimator and their specific notation.

The IVQR estimator as a constrained minimization problem

The IVQR estimator proposed by CH06, for a given quantile $\tau$, can be written as a constrained minimization problem, where the constraints are the moment conditions, that is,

equation[equation omitted — 74 chars of source]

subject to

align[align omitted — 203 chars of source]

Now we write this constrained optimization as a Lagrangian problem\footnote{See Pouliot19 and Kaido21 for recent contributions that tackle the problem of practical implementation of the IVQR methods.} as

align[align omitted — 277 chars of source]

where $\lambda_x$ is a $1\times k$ vector and $\lambda_z$ is a $1\times p$ vector. Therefore, the IVQR estimator is given by the empirical counterpart of

equation*[equation* omitted — 111 chars of source]

where $\theta=(\alpha',\beta',\gamma')'$.

The first derivatives of the Lagrangian in equation (ref) are

align[align omitted — 806 chars of source]

where $f:=f_{u_\tau}(0|d,x,z)$ denotes the density function of $u_\tau:=y-d\alpha_0(\tau)-x\beta_0(\tau)$ conditional on $s=(d,x,z)$, evaluated at the $\tau$-th conditional quantile, which is zero. Note that $f$ is specific for each quantile $\tau$. This density function plays a central role in what follows.

The solution should have all equations above equal to zero when assuming an interior solution as in Assumption (ref) below. Thus, from equation (ref),

equation[equation omitted — 139 chars of source]

Then, replacing (ref) in (ref),

equation*[equation* omitted — 166 chars of source]

such that

equation[equation omitted — 185 chars of source]

Finally, replacing (ref) in (ref),

align*[align* omitted — 437 chars of source]

where $\bm{0}_r$ is a $r\times1$ vector of zeros.

Therefore, we can restate the IVQR problem for $(\alpha',\beta',\gamma')'$ as a system of three equations given by

align[align omitted — 629 chars of source]

First-stage IVQR parameters

Given equations (ref)--(ref) above, we can see that (ref) provides a first-stage representation of the IVQR model. This can be written as

equation[equation omitted — 55 chars of source]

where

align[align omitted — 412 chars of source]

Here $\delta$ is a $p\times r$ matrix. Notice that equation (ref) is a least-squares projection coefficient. In particular, the representation in (ref) is a weighted projection, where the endogenous variable(s), $d$, is(are) regressed on the IV, $z$, and the exogenous variables, $x$. This is the analogue to the first-stage in the 2SLS case, with the difference that the QR case is a weighted regression. The weights are given by the conditional density function of the innovation term in the QR structural model, conditional on the endogeneous and exogenous covariates together with the instruments.

Hence, for each endogeneous variable, say $d_j$ for $j=1,2,...,r$, $\delta_{j}$ in equation (ref) can be recovered as the solution to the following optimization problem

equation[equation omitted — 159 chars of source]

Note that the parameter $\delta$ also depends on $\theta=(\alpha',\beta',\gamma')'$, through the conditional density function $f$ at quantile $\tau$. Thus, this first-stage representation depends on the structural (second-stage) parameters, and as such, it is different from the 2SLS case in mean regression models.

We notice that the first-stage in equation (ref) is different from those in the existing literature using two-stage regressions for conditional quantile models. Amemiya82, Powell83, ChenPortnoy96, and KimMuller04 propose different two step procedures in which the first step fits the endogenous variable(s) as a function of exogenous covariates and IV, and this is then plugged in a second-stage. Nevertheless, these papers use least squares without weighting or standard quantile regression in the first-stage. Our procedure derives the first-stage from the IVQR set-up, thus confirming that a first-stage (albeit different) is part of the model.

Relation to the Jacobian condition

Now we consider the relationship between the first-stage derived in the previous section, in particular equation (ref), and the rank identification conditions for the IVQR estimator of CH. For simplification, we consider a model without additional exogenous covariates $x$.

As discussed in CH, the IVQR optimization problem is asymptotically equivalent to solving the following moment condition

equation*[equation* omitted — 115 chars of source]

To establish the asymptotic properties of the IVQR estimator, it is required that the Jacobian matrices, $\frac{\partial}{\partial\alpha}\Pi((\alpha,\gamma),\tau)$ and $\frac{\partial}{\partial\gamma}\Pi((\alpha,\gamma),\tau)$, are continuous and full column rank (see below the conditions for the derivation of the asymptotic properties of the estimator, in particular, Assumption 1, item R3). We show here that these conditions are embedded in the FS-IVQR representation.

The rank Jacobian conditions are

align*[align* omitted — 322 chars of source]

The first equation implies that for the case of one endogeneous variable, $r=1$, $\textnormal{E}[f\cdot z'd]$ has at least one non-zero column, and the second equation requires $p$ noncollinear valid instruments. Now notice that from the FS-IVQR representation given by equation (ref), in the case without exogenous regressors, simplifies to:

equation*[equation* omitted — 77 chars of source]

Therefore, the matrices involved in the rank conditions directly appear in representation (ref). Also, if the FS-IVQR parameter $\delta=0$, then the rank conditions cannot be satisfied. Note that this is a necessary condition, but not a sufficient one. Furthermore, by checking how close $\delta$ is to zero one is in fact evaluating the strength of the identification condition.

Further intuition on the FS-IVQR

The restriction in equation (ref) provides a natural framework to evaluate the relevance of the instruments in IVQR models.

First, the first-stage regression representation in (ref) is a weighted linear projection, where the weights are the conditional density function of the innovation term in the QR structural model, conditional on the endogeneous and exogenous covariates together with the instruments. This exposes a caveat of the QR IV model. In order to estimate the parameters in (ref) consistently, one needs a consistent estimate of the density $f$, and hence at least one valid instrument must be available to the researcher. This is in contrast with the standard conditional average models where the first-stage is a simple OLS regression without weights. The required weights in the QR case will be further discussed below when we suggest a test for the validity of the IV.

Second, notice that the parameter $\delta$ captures the strength of the instrument in the sense it measures the correlation between the instrument $z$ and the endogenous variable $d$ weighted by the density function $f$. This is the QR counterpart of the first-stage partial correlation of $z$ on the endogenous variables $d$ for the 2SLS. As noted by GalvaoMontes15 the CH set-up is equivalent to the 2SLS in least-squares models. In fact the CH estimator is the QR counterpart of a 2SLS estimator. The expression above also shows that there is an implicit first-stage, similar to that in 2SLS problems. As such, this provides an analytical expression to evaluate the relevance of the IV. When the instrument is valid, $\delta\neq\bm{0}_{p\times r}$.

Third, note that the instrument $z$ does not belong in the structural quantile model (ref), hence $\gamma=\bm{0}_{p\times r}$ can be used for identification, a key feature of the CH IVQR estimator. Equation (ref) also shows that when $\delta=\bm{0}_{p\times r}$, the value of $\gamma$ is irrelevant, and therefore it cannot be used in the IVQR procedure to solve endogeneity. As such, $\delta\neq\bm{0}_{p\times r}$ is a necessary condition for the IV to have a purpose in the CH set-up. Therefore, a test for the validity of the instruments can be based on a test for statistical significance of $\delta$.

Finally, another way of gaining intuition on the test is the following. Assume that $r=1$ (i.e. only one endogenous variable), then (ref) is in fact equal to 0, a scalar. If we further assume that $A=I_p$, then

equation[equation omitted — 49 chars of source]

where $\delta=[\delta_1,\ldots,\delta_p]'$ is the column vector that has the first-stage effect of all IV on $d$. Note again that if $\delta=\bm{0}_{p\times 1}$, then the vector $\gamma$ could have any value and its implied restrictions would be irrelevant.

Formulation of the test for validity of the IV

In this section we suggest tests for the validity of the IV using the first-stage representation. The formulation of the test proposed in this paper is based on the condition given in equation (ref) together with the first-stage IVQR representation in equation (ref). A test for validity of the instruments for $p$ instruments can be based on the null hypothesis

equation[equation omitted — 63 chars of source]

against the alternative

equation[equation omitted — 74 chars of source]

We highlight that, differently from the 2SLS, the first-stage IVQR in (ref) is for a given quantile $\tau$. Thus, for the same variables $d$ and instruments $z$, the strength of the instruments may vary across different quantiles. This variation is captured by the weights $f$.

Note that the procedure works for $r\geq 1$, that is for one or more than one endogenous variables. In the $r>1$ case, separate tests could be applied as in 2SLS analysis where there may be a different first-stage for each endogeneous variable. To simplify the procedures below we assume that $r=1$, that is, there is only one endogenous variable.

The expressions of the null and the alternative hypotheses in (ref) and (ref), respectively, lead to the following testing procedure.

When $H_0$ is true, under suitable regularity conditions, $\hat{\delta}$ converges in probability to $\bm{0}_{p\times r}$ for a given $\tau$. On the other hand, when $H_A$ is true, $\hat{\delta} $ converges in probability to $\delta_0\neq \bm{0}_{p\times r}$. Therefore, it is reasonable to reject $H_0$ if the magnitude of $\hat{\delta} $ is suitably large.

A natural choice to test $H_0$ against $H_1$ for the case of $r=1$ is the Wald statistic as

equation[equation omitted — 84 chars of source]

where $V_{\delta}$ is the asymptotic covariance matrix of $\sqrt{n} \hat{\delta}$ under $H_0$. In practice, $V_{\delta}$ is replaced by a suitable consistent estimate. We will discuss the practical implementation as well the limiting distribution in the next section.

Empirical implementation and asymptotic distribution

In this section we propose a two step estimator for the first-stage instrumental variables quantile regression (FS-IVQR), consider its empirical implementation, and derive the estimators' asymptotic distribution. The two steps estimation procedure consists of estimating the conditional density using the IVQR model in the first step, and in the second step employing a weighted least squares (WLS) regression. For simplicity of exposition, we present the case of $r=1$, i.e. one endogenous variable, but as discussed above the case of $r>1$ can be implemented using separate regressions.

FS-IVQR Estimator

The FS-IVQR estimator requires a consistent estimator of $\mu$ in (ref), which will be based on WLS based on the estimator of $f$, at a given quantile of interest $\tau$. The estimator has two steps as following:

1) In the first step we obtain $\hat{\theta}=(\hat{\alpha}, \hat{\beta}', \hat{\gamma}')'$ from the CH estimator,

equation*[equation* omitted — 90 chars of source]

where

equation*[equation* omitted — 191 chars of source]

Provided that the $\tau$th conditional quantile function of $y|s$ is linear, as in (ref), then for $h_{n}\to 0$ we can consistently estimate the parameters of the $\tau \pm h_{n}$ conditional quantile functions by $\hat{\theta}(\tau \pm h_{n})$. And the density $f_{i}:=f_{u_\tau}(0|d=d_{i},x=x_{i},z=z_{i})$ can thus be estimated by the difference quotient

equation[equation omitted — 132 chars of source]

The estimation in (ref) is a natural extension of sparsity estimation methods, suggested by HendricksKoenker92.\footnote{We note that a kernel estimator for the conditional density, as in Powell91, can be used. The procedure would use the error term $\hat{u}_\tau:=y-d\hat{\alpha}(\tau)-x\hat{\beta}(\tau)$ from the first step CH IVQR estimator. We describe the procedure using the sparsity estimation for simplicity.} The estimator is discussed in further details in ZhouPortnoy96 and Koenker05. We introduce the simplifying notation $\hat{f}_{i}:=\hat{f}_{u_\tau}(0|s=s_i)$.\footnote{We are assuming that there is only one endogenous variable, $r=1$. Otherwise the analysis below should be repeated separately for each endogenous variable as there will be a different first-stage for each one.} The bandwidth for the density estimation can be chosen heuristically as a scaled version of HallSheather88:

equation*[equation* omitted — 178 chars of source]

2) In the second step the parameters of interest $\delta$ can be obtained from a feasible WLS as

equation[equation omitted — 196 chars of source]

Equation (ref) produces $\hat{\delta}$ which is the main object of interest.

Define $Y$, $X$, $D$ and $Z$ as the matrices formed from a random sample of $\{y_i,d_i,x_i,z_i\}_{i=1}^n$. Similarly define $W=[X,Z]$. Define the weighting diagonal matrix

equation*[equation* omitted — 110 chars of source]

Then, the estimator in (ref) above can be written in a simple matrix notation as

equation[equation omitted — 76 chars of source]

Notice that if $f_i$ is a constant for all $i$, then the proposed FS-IVQR method should deliver same estimates as FS-2SLS for the mean. This would happen, for example, in the case of $i.i.d.$ innovations in the second-stage structural model. Thus, there will be differences between the two estimators only when $f_i$ varies across $i$, that is, when the weighting factor is not a constant. Example 1 (location model) Appendix B shows a case where the density function is a constant. A typical example where the weights are not constant across individuals is the location-scale model, see Examples 2 and 3 in Appendix B.

Asymptotic distribution

In this subsection, we derive the asymptotic distribution of the proposed estimator. The asymptotic properties of the IVQR estimator can be found in CH06 and the assumptions therein are those required for inference. We consider Assumption 2 in CH06, that we reproduce here for convenience. It imposes conditions for $\theta_0$ to be identified and estimated.

assumptionR1. Sampling. $\{y_i,x_i,d_i,z_i\}$ are $iid$ defined on a probability space and take values in a compact set.\\ R2. Compactness and convexity. For all $\tau\in(0,1)$, $(\alpha,\beta,\gamma\in {\rm int} (\mathcal{A}\times\mathcal{B}\times\mathcal{G})$ is compact and convex.\\ R3. Full rank and continuity. $y$ has bounded conditional density (conditional on $w$), and for $\theta=(\alpha,\beta,\gamma)$, \[\Pi(\theta,\tau):=\textnormal{E}\left[(\tau-\bm{1}(y<d\alpha+x\beta+z\gamma)\cdot[x,z]\right],\] Jacobian matrices $\frac{\partial}{\partial(\alpha',\beta')}\Pi(\theta,\tau)$ and $\frac{\partial}{\partial(\beta',\gamma')}\Pi(\theta,\tau)$ are continuous and have full rank, uniformly over $\mathcal{A}\times\mathcal{B}\times\mathcal{G}$ and the image of $\mathcal{A}\times\mathcal{B}\times\mathcal{G}$ under the mapping $(\alpha,\beta)\mapsto\Pi(\theta,\tau)$ is simply connected. Assume that $\theta_0=(\alpha_0,\beta_0',\gamma_0')'$ is the unique solution to the CH problem.

We impose additional conditions for deriving the limiting properties of the feasible first-stage estimator in (ref) using the sparsity estimation in (ref).

assumptionLet $\varepsilon_{i}:=d_{i}-x_{i}\psi_0- z_{i}\delta_0$, with $\textnormal{E}[\varepsilon_{i} | w_{i} ] = 0$, and $\textnormal{E}[\varepsilon_{i}^{2}|w_{i}] = \sigma_{i}^2$. Also, let $f_{i}:=f_{\theta_0}(y-s\theta_0|s=s_i)$ and assume that $\textnormal{E}[|f_{i}^{-2}w_{i} \varepsilon_{i}|]<\infty$. Let $\Omega_{f \sigma}:=\textnormal{E}[f_{i}^2\sigma_{i}^{2} w_{i} w_{i}']$ and $\Omega_{f }:=\textnormal{E}[f_{i} w_{i} w_{i}']$. The limits $\lim_{n\to\infty}\frac{1}{n}\sum_{i}^{n} f_{i}^2\sigma_{i}^{2} w_{i} w_{i}' = \Omega_{f \sigma}$ and $\lim_{n\to\infty}\frac{1}{n}\sum_{i}^{n} f_{i} w_{i} w_{i}' = \Omega_{f }$ exist and are nonsingular (and hence finite).

Assumption (ref) contains conditions for establishing consistency and asymptotic normality of the proposed estimator. The next result presents an intermediate result.

lemmaUnder Assumptions (ref)--(ref), as $n\rightarrow\infty$, $h_n\rightarrow0$ and $nh_n^2\rightarrow\infty$, \begin{equation} \sqrt{n}\left( \hat{\mu}- \mu_0 \right) \stackrel{d}{\rightarrow} \mathcal{N}\left(\bm{0}_{k+p}, V(\mu_0) \right), \end{equation} where $\mu_0:=(\psi_0,\delta_0)=\underset{\psi,\delta}{\rm argmin} \;\textnormal{E} \left[f \cdot (d-x\psi- z\delta)^2\right]$ and $V(\mu_0) = \Omega_{f}^{-1} \Omega_{f \sigma} \Omega_{f}^{-1}$ is the asymptotic covariance matrix.
proofIn Appendix A.

Asymptotic distribution of the test statistic

Consider a subset of the instruments, $p_1<p$, and consider a partition of $\delta=[\delta_1',\delta_2']'$ of the corresponding first-stage parameters of interest, with dimensions $p_1$ and $p_2$ (with $p=p_1+p_2$), respectively. Consider a $p_1\times (k+p)$ matrix $R=[\bm{0}_{p_1\times k}, \bm{I}_{p_1},\bm{0}_{p_1\times p_2}]$ where $\bm{I}_{p_1}$ is an identity matrix of dimension $ p_1\times p_1$. Thus, $R\mu=\delta_1$ is the subvector of interest. Let $\hat{V}(\hat{\mu})$ be a consistent estimator of $V(\mu_0)$, which can be obtained from the WLS procedure. The next result derives the limiting distribution of the test statistic in equation (ref).

propositionConsider Assumptions (ref)--(ref), $n\rightarrow\infty$, $h_n\rightarrow0$ and $nh_n^2\rightarrow\infty$. Furthermore, assume that $dim(z)=p>p_1\geq 1$. Then, under $\delta_{2}\neq\bm{0}_{p_2}$ and $H_0:\delta_{1}=\bm{0}_{p_1}$ and local alternatives $H_A:\delta_{1}=\bm{a}_{p_1}/\sqrt{n}$, \begin{equation} T_n=n \left(R \hat{\mu} \right)' \{R\hat{V}(\hat{\mu})R'\}^{-1} \left(R \hat{\mu} \right) \stackrel{d}{\rightarrow} \chi^2_{p_1}(\bm{a}_{p_1}). \end{equation}
proofIn Appendix A.

Computation of the test statistic (ref) requires a non-parametric estimator of $f$, the conditional density of $u_\tau|d,x,z$ evaluated at the specific quantile of interest $\tau$. Given that the weights need to be estimated, the proposed FS-IVQR has specific properties when testing under the null hypothesis of an invalid instrument. The condition on the number of IV being larger than the number of parameters tested in the null hypothesis is required for consistent estimation of $\theta$ under the null, which in turn, is used for the consistent estimation of $f$.

remarkAs noted above, testing for the null (ref) that the instruments are not valid requires a consistent estimate of the density $f$, which in turn requires at least one valid instrument, and hence the restrictions on the availability of at least one valid instrument as well as on the number of IV being larger than the number of parameters tested in the null hypothesis. Hence, due to the required estimation of weights in the first stage, it is difficult to establish a complete analogous F-statistic type rule-of-thumb for categorizing weak instruments as in Staiger97, and subsequent variants as Sanderson16, Lee20 and others for OLS models (see StockYogo05 for an extensive discussion). Nevertheless, dropping the multiple instruments requirement is a very interesting line of future research for investigating the weak instruments problem. Therefore, we suggest to practitioners to use the methods for the QR first-stage together with the robust inference procedures for weak identification in QR models proposed in CH08, Jun08, CHJ09 and AndrewsMikusheva16.

Monte Carlo experiments

We analyze in this section the performance of the proposed test with finite samples through a series of Monte Carlo simulation exercises. The data generating process (DGP) has the following location-scale model:

align[align omitted — 151 chars of source]

where $x_i$, $z_{1i}$ and $z_{2i}$ are three independent variables with distribution $U(0,1)$; $u_i$ and $v_i$ have standard bivariate normal distribution with correlation $0.50$. The constant parameter $c_1=10$ is set to a large value to satisfy the monotonicity assumption. Equations (ref)--(ref) specify a model where there could be pure location or location-scale specifications in the first stage, thus allowing the instruments to have different effects on the endogenous variable. Note that the parameters $a$ and $b$ determine the type of effect that the instrument $z_{1}$ has on the endogenous covariate $d$. For example, if $a\neq0$ and $b = 0$ the instrument $z_{1}$ has a pure location effect on $d$ (pure location shift model), while if $a = 0$ and $b\neq0$ the effect is only on the variance of the endogenous covariate (pure scale shift model).

In all cases we consider tests for $H_0:\delta_1=0$ where this is the first-stage parameter associated with the $z_1$ instrument defined in the previous sections. We consider two different cases to investigate the numerical properties of the tests. In the first case, $\phi=1$, there is a second instrument, $z_{2}$, such that the model correctly identifies the parameters in the structural equation (ref) for all possible values of $a$ and $b$, even under the case that $a=b=0$. In the second case, we set $\phi=0$, and therefore, under the null hypothesis the consistent estimation of the weights $f$ is problematic. Also, in this case, when $a=b=0$, there is no valid available instrument.

We will consider three different test statistics from different estimators. First, for comparison purposes, we present a Wald test for the coefficient in $z_1$ using a simple regression model of $d$ on $(x,z_1,z_2)$ in a standard 2SLS framework, denoted FS-2SLS. Second, we test for $H_0:\delta_1=0$ using the true density function, $f$, as weights, that is, using the true $\theta_0$, denoted FS-IVQR (true density). We note that this is not observed in practice, and we include these results for comparison purposes. Our proposed test studied in the previous section is the third one, denoted FS-IVQR (sparsity), where we use the sparsity function estimation described above. Note that the three tests differ only in the weighting procedure used in the regression of $d$ on $(x,z_1,z_2)$ or $(x,z_1)$.

Table (ref)--(ref) show the empirical size (i.e. $a=b=0$) of the computed test with 2000 simulations for $ n = \{500, 1000 \}$ and for the quantiles $\tau = \{0.25, 0.50, 0.75\}$.

Consider first the case where there is a second instrument, $\phi=1$ in Table (ref). The tests have approximately correct empirical size in all cases. As such they clearly evaluate if the instrument $z_1$ exerts an effect on the endogenous variable $d$. In all cases they have a similar performance to the FS-2SLS case.

Now consider the case where there is no available second instrument, $\phi=0$ in Table (ref). The idea of this experiment is to evaluate the test performance when there is lack of identification under the null. In this case, the weights in the structural model cannot be estimated consistently under the null. Since the proposed test evaluates the relationship between $z_1$ and $d$, the main issue is whether this relationship can be evaluated in other than the OLS model. The simulations show that the size is correct for the sparsity estimator. This result suggests that the test can be used even when the structural parameters cannot be estimated under the null (because $z_1$ does not solve the endogeneity problem).

This is not a general result, however, but it illustrates the role of the density function as a weighting factor. In order to explore this, we consider three examples in Appendix B closely related to the DGP used in the Monte Carlo experiments. When the instrument $z$ is available we should be estimating the correct structural parameters and $f_{u_\tau}(0|d,x,z)$ where $u_\tau=y-Q_\tau(y|d,x,z)$. However, the case where, under the null, $z$ is invalid would be equivalent to the case where there are no instruments available. That is, we would not be able to solve the endogeneity in the second-stage. Note that, for this case, the density function that will be implicitly used is that of $u^*_\tau=y-Q_\tau(y|d,x)$. The examples in Appendix B compare $f_{u^*_\tau}(0|d,x)$ with $f_{u_\tau}(0|d,x,z)$.\footnote{Let $\tilde{\alpha}$ and $\tilde{\beta}$ be the parameters that result from the estimation of the biased structural model without instruments, $Q_\tau(y|d,x)=d\tilde{\alpha}+x\tilde{\beta}$. Note that $u_\tau=y-Q_\tau(y|d,x,z)=y-d\alpha_0-x\beta_0$ can be written as $y-d\tilde\alpha-x\tilde\beta-bias(d,x)$, where $bias(d,x)=d(\alpha_0-\tilde{\alpha})+x(\beta_0-\tilde{\beta})$ such that $u_\tau=u^*_\tau-bias(d,x)$.} The partial results suggest that if $f_{u_\tau}(0|d,x,z)$ and $f_{u^*_\tau}(0|d,x)$ are proportional to each other when they vary with $d$, we could implement the first-stage test under the null of all IV being invalid.

table[table omitted — 1,954 chars of source]
table[table omitted — 1,949 chars of source]

To analyze the empirical power of the tests, we performed 2000 simulations only for the case with $n = 1000$ and we calculated the rejection rates of the proposed procedure for the quantiles $\tau = \{0.25, 0.50, 0.75\}$. As benchmark we also use the test rejection rates obtained in the FS-2SLS method, i.e., the Wald test of an OLS regression of $d$ on $z_1$. The results appear in Figures (ref) and (ref). For each figure we have two blocks, (i) and (ii), where in (i) we evaluate a pure location first-stage model of $z_1$ on $d$ using $a = \{0, 0.10, ..., 0.90,1\}$ and $b=0$, and in (ii) we set $a=0$ and we vary $b=\{0, 0.10, ..., 0.90,1\}$ such that $z_1$ has only a scale effect on $d$.

We first consider the case where there is a second valid instrument $\phi=1$. Figure (ref), block (i) pure location first-stage, shows that the FS-IVQR power computed with true and estimated densities behaves similarly to FS-2SLS. That is, they correctly reject as $a$ increases. The estimated density model has slightly less power than the one with the true density. For block (ii), the results of the FS-IVQR differ when we are in the presence of a pure-scale model for $d|z_1$. Note that in this case there is no relationship between $d$ and $z_1$ at the mean (FS-2SLS), but it does affect the other points of the conditional distribution. Therefore, the first-stage of 2SLS does not find any relationship between the endogenous variable and the instrument while the FS-IVQR estimators (both true and estimated weights) are able to correctly detect it.

Finally, consider the last case when $\phi=0$ in Figure (ref). The FS-IVQR tests also work in this case. In both (i) and (ii) cases, the tests detect an association between the instrument and the endogenous variable. In case (ii) the FS-IVQR rejects as $b$ increases while FS-2SLS does not. As noted in Table (ref) the test works even for the case where $a=b=0$ and the endogeneity problem in the structural estimators cannot be solved.

figure[figure omitted — 160 chars of source]
figure[figure omitted — 165 chars of source]

Empirical application: Card (1995) college proximity as an instrument for education

In this section we show an application of the proposed test to a Mincer equation to estimate returns to schooling. The data used is taken from Card95 and correspond to 3010 individuals of the US National Longitudinal Survey of Young Men.\footnote{Downloaded from \url{http://davidcard.berkeley.edu/data_sets/proximity.zip}} Following the same specification of that paper, the model describes wages as a function of the years of education and other exogenous controls such as work experience, race and a set of geographic and regional variables. A classic problem with this model is that ability is unobservable and therefore its omission induces a potential bias due to endogeneity of the OLS estimator. Specification errors have analogous consequences on QR estimators, as analyzed by Angrist06. Card95 proposes to implement an IV strategy using two measures of proximity to the university as external variables to the wage equation. For this application we set $A$, the weighting matrix in the CH-IVQR estimator, equal to the inverse of the asymptotic covariance matrix of $\hat{\gamma}$, as suggested by CH. The grid used in the minimization problem is $\alpha(\tau) = \{0,0.004,0.008,...,0.992,0.996,1\}$.

table[table omitted — 2,258 chars of source]

Table (ref) shows the results of the first-stage to check if the IV are valid, together with the estimated second-stage results. The first column corresponds to the 2SLS mean model and the next ones are the regressions proposed for IVQR for $\tau\in\{0.25,0.50,0.75\}$. The results shows that the first instrument (lived near 2-year college in 1966) is not relevant for the low quantiles and the mean but it is significant for middle and high quantiles. Also, note that although the second instrument (lived near 4-year college in 1966) rejects the null hypothesis for the mean, this variable has different degree of significance across quantiles. In particular, this is for $\tau=0.75$ where the instrument is relevant only at 10% significance. These results are very important since although the proximity to the university seems to be a valid instrument to identify the causal effect of education on the mean, our test also indicates a certain limitation when the object of study is to evaluate the impact on the lower part of conditional distribution of wages. Therefore, this alerts for the quality of the asymptotic properties of the IVQR estimates in the presence of invalid instruments.

Conclusions

This paper proposes a first-stage model and a testing procedure to evaluate the degree of association between the IV and the endogenous regressor(s) in the IVQR estimator. The procedure developed here allows to evaluate instruments in a similar vein to that in 2SLS models for the conditional average, that is, by looking at the statistical significance of the instruments in the first-stage regression. In turn, this will allow to investigate IV validity for specific quantiles. Nevertheless, due to the requirement of consistent estimation of the weights in the first stage, it is important to notice that the testing requires the availability of at least of instrument. This caveat for QR IV models leads to two conclusions. First, it is difficult to derive a complete analogous F-statistic type rule-of-thumb for categorizing weak instruments. We leave this problem to future research. Second, we strongly suggest the use of weak identification robust inference procedures for QR models for practical work applying QR instrumental variables. Monte Carlo experiments clearly illustrate that one may encounter cases where the IV are not valid for the mean, but are still valid for some quantiles. The same issue appears in the empirical application.

The analysis may be extended in the following two directions. First, this approach can be used to identify quantile-specific treatment effects, where an IV estimate being significant at some quantiles corresponds to a particular effect of a treatment. Second, the procedure outlined here could be combined with the second-stage inference to produce statistics similar to the Staiger97 F-statistics rule-of-thumb. In particular, to study weak instruments issues in QR models.