Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
74,391 characters · 16 sections · 0 citation commands
Quantile Models with Endogeneity
Quantile regression is a tool for estimating conditional quantile models that has been used in many empirical studies and has been studied extensively in theoretical econometrics; see \citeasnoun{koenker:1978} and \citeasnoun{koenker:book}. One of quantile regression's most appealing features is its ability to estimate quantile-specific effects that describe the impact of covariates not only on the center but also on the tails of the conditional outcome distribution. While the central effects, such as the mean effect obtained through conditional mean regression, provide interesting summary statistics of the impact of a covariate, they fail to describe the full distributional impact unless the conditioning variables affect the central and the tail quantiles in the same way. In addition, researchers are interested in the impact of covariates on points other than the center of the conditional distribution in many cases. For example, in a study of the effectiveness of a job training program, the effect of training on the lower tail of the earnings distribution conditional on worker characteristics may be of more interest than the effect of training on the mean of the distribution.
In observational studies, the variables of interest (e.g. education or prices) are often endogenous. Just as with the conventional linear model, endogeneity of covariates renders the conventional quantile regression inconsistent for estimating the causal (structural) effects of covariates on the quantiles of economic outcomes. One approach to addressing this problem is to generalize the instrumental variables framework to allow for estimation of quantile models. In this paper, we review developments in instrumental variables approaches to modeling and estimating quantile treatment (structural) effects (QTE) in the presence of endogeneity.
We focus our review on the modeling framework of \citeasnoun{iqr:ema} which provides conditions for identification of the QTE without functional form assumptions. The principal identifying assumption of the model is the imposition of conditions which restrict how rank variables (structural errors) may vary across treatment states. These conditions allow the use of instrumental variables to overcome the endogeneity problem and recover the true QTE. This framework also ties naturally to simultaneous equations models, corresponding to a structural simultaneous equation model with non-additive errors. Within this framework, estimation and inference procedures for linear quantile models have been developed by \citeasnoun{iqr:joe}, \citeasnoun{iqr:joe2}, \citeasnoun{fsqr}, and \citeasnoun{JunWeakIdIVQR}; nonparametric estimation has been considered by \citeasnoun{CIN:JoE}, \citeasnoun{HorowitzLeeNonparametricIVQR}, and \citeasnoun{GagliardiniScaillet}; and inference with discrete outcomes has been explored by \citeasnoun{ChesherDiscrete}. Moreover, the modeling framework provides a foundation for other estimation methods based on IV median-independence and more general quantile-independence conditions as in \citeasnoun{abadie:1997}, \citeasnoun{mcmc}, \citeasnoun{chen:linton}, \citeasnoun{hong:tamer}, \citeasnoun{honore:hu}, and \citeasnoun{sakata}. It is also important to note that the modeling framework we review can be used to study nonparametric identification of structural economic models in cases where quantile effects are not necessarily the chief objects of interest. \citeasnoun{BerryHaileDiscreteChoiceId} provide an excellent example of this in the context of discrete choice models with endogeneity.
We also briefly review other modeling approaches for quantile effects with endogenous covariates. \citeasnoun{abadie} consider a QTE model for the sub-population of “compliers" which applies to binary endogenous variables with binary instruments. \citeasnoun{imbens:newey}, \citeasnoun{chesher}, \citeasnoun{SLeeTriangularIVQR}, and \citeasnoun{koenker:ma} use models with triangular structures and show how control functions can be constructed and used to estimate structural objects of interest. While these models share some features with the model of \citeasnoun{iqr:ema}, the three approaches are non-nested in general.
Quantile models with endogeneity have been used in many empirical studies in economics. See \citeasnoun{abadie}; \citeasnoun{401k}; \citeasnoun{hausman:sidak}; \citeasnoun{IVQRApp:AirTraffic}; \citeasnoun{IVQRApp:Union}; \citeasnoun{IVQRApp:Agriculture}; \citeasnoun{IVQRApp:Insurance}; \citeasnoun{IVQRApp:BirthWeight}; \citeasnoun{IVQRApp:Columbia}; \citeasnoun{IVQRApp:Autor}; and \citeasnoun{IVQRApp:Somainiy} among others. We do not provide a review of empirical applications but note these papers provide further discussion of how the instrumental variables quantile model relates to their specific framework and illustrate some of the rich effects that one can estimate using quantile methods.
In this section, we present an instrumental variable model for quantile treatment effects (QTE), its main econometric implication, and the principal identification result.
Our model is developed within the conventional potential (latent) outcome framework, e.g. \citeasnoun{heckman:robb}. Potential real-valued outcomes which vary among individuals or observational units are indexed against potential treatment states $d \in \mathcal{D}$ and denoted $Y_d$. The potential outcomes $\{ Y_{d}\}$ are latent because, given the selected treatment $D$, the observed outcome for each individual or observational unit is only one component $$ Y := Y_{D}$$ of the potential outcomes vector $\{Y_d\}$. Throughout the paper, capital letters denote random variables, and lower case letters denote the potential values they may take. We do not explicitly state various technical measurability assumptions as these can be deduced from the context.\footnote{ For simplicity, we could assume that $d$ takes on a countable set of values $\mathcal{D}$ or make separability assumptions which imply that the stochastic process $\{Y_d, d \in \mathcal{D}\}$ is defined from its definition over a countable subset $\mathcal{D}_0 \subset \mathcal{D}$. See \citeasnoun{vdV-W}.}
The objective of causal or structural analysis is to learn about features of the distributions of potential outcomes $Y_d$. Of primary interest to us are the $\tau$-th quantiles of potential outcomes under various treatments $d$, conditional on observed characteristics $X=x$, denoted as $$ q(d, x, \tau).$$ We will refer to the function $q(d,x,\tau)$ as the quantile treatment response (QTR) function. We are also interested in the quantile treatment effects (QTE), defined as $$ q(d_1, x, \tau) - q(d_0, x, \tau),$$ that summarize the differences in the impact of treatments on the quantiles of potential outcomes (\citeasnoun{lehmann:1974}, \citeasnoun{doksum:1974}).
Typically, the realized treatment $D$ is selected in relation to potential outcomes, inducing endogeneity. This endogeneity makes the conventional quantile regression of observed $Y$ on observed $D$, which relies upon the restriction $$ P[ Y \leqslant \theta(D, X, \tau) | X, D] = \tau \ \text { a.s., } $$ inappropriate for measuring $q(d,x,\tau)$ and the QTE. Indeed the function $\theta(d, x, \tau)$ solving these equations will not be equal to $q(d,x,\tau)$ under endogeneity. The model presented next states conditions under which we can identify and estimate the quantiles of latent outcomes through the use of instruments $Z$ that affect $D$ but are independent of potential outcomes and the nonlinear quantile-type conditional moment restrictions $$ P[ Y \leqslant q(D, X, \tau) | X, Z] = \tau \ \text { a.s. } $$
Having conditioned on the observed characteristics $X=x$, each latent outcome $Y_d$ can be related to its quantile function $q(d,x,\tau)$ as\footnote{This follows by Fisher-Skorohod representation of random variables which states that given a collection of variables $\{\zeta_d\}$, each variable $\zeta_d$ can be represented as $ \zeta_d = q(d, U_d)$, for some $U_d\sim U(0,1)$, cf. \citeasnoun{durrett}, where $q(d, \tau)$ denotes the $\tau$-quantile of variable $\zeta_d$.}
is the structural error term. We note that representation ((ref)) is essential to what follows.
The structural error $U_d$ is responsible for heterogeneity of potential outcomes among individuals with the same observed characteristics $x$. This error term determines the relative ranking of observationally equivalent individuals in the distribution of potential outcomes given the individuals' observed characteristics, and thus we refer to $U_d$ as the rank variable. Since $U_d$ drives differences in observationally equivalent individuals, one may think of $U_d$ as representing some unobserved characteristic, e.g. ability or proneness.\footnote{\citeasnoun{doksum:1974} uses the term proneness as in “prone to learn fast" or “prone to grow taller".} This interpretation makes quantile analysis an interesting tool for describing and learning the structure of heterogeneous treatment effects and accounting for unobserved heterogeneity; see \citeasnoun{doksum:1974}, \citeasnoun{heckman:smith}, and \citeasnoun{koenker:book}.
For example, consider a returns-to-training model, where $Y_d$'s are potential earnings under different training levels $d$, and $q(d,x,\tau)$ is the conditional earnings function which describes how an individual having training $d$, characteristics $x$, and the latent “ability" $\tau$ is rewarded by the labor market. The earnings function may be different for different levels of $\tau$, implying heterogeneous effects of training on earnings of people that have different levels of “ability". For example, it may be that the largest returns to training accrue to those in the upper tail of the conditional distribution, that is, to the “high-ability" workers.\footnote{It is important to note that the quantile index, $\tau$, in $q(d,x,\tau)$ refers to the quantile of potential outcome $Y_d$ given that exogenous variables are set at $X = x$ and not to the unconditional quantile of $Y_d$. For example, suppose that one of the control variables in the earnings example is years of schooling. An individual at the 30$^{\textnormal{th}}$ percentile of the distribution of $Y_d$ given say 20 years of schooling is not necessarily low income as even a relatively low earner with that level of education may still earn above the median earnings in the overall population.}
Formally, the IVQT model consists of five conditions (some are representations) that hold jointly.
Main Conditions of the Model: Consider a common probability space $(\Omega, F, P)$ and the set of potential outcome variables $(Y_d, d \in \mathcal{D})$, the covariate variables $X$, and the instrumental variables $Z$. The following conditions hold jointly with probability one:
The following is the main econometric implication of the model.
The first result states that the main consequence of A1-A5 is a simultaneous equation model ((ref)) with non-separable error $U$ that is independent of $Z,X$, and normalized so that $U \sim U(0,1)$. The second result considers econometric implications when $\tau \mapsto q(D,X,\tau)$ is strictly increasing, which requires that $Y$ is non-atomic conditional on $X$ and $Z$. In this case, we obtain the conditional moment restriction ((ref)). This implication follows from the first result and the fact that $$ \{ Y \leqslant q(D,X,\tau) \} \text{ is equivalent to } \{ U \leqslant \tau\}, $$ when $q(D,X,\tau)$ is strictly increasing in $\tau$. The final result deals with the case where $Y$ may have atoms conditional on $X$ and $Z$, e.g. when $Y$ is a count or discrete response variable. The first two results were obtained in \citeasnoun{iqr:ema}, and the third result is in the spirit of results given in \citeasnoun{ChesherRosenSmolinski}; \citeasnoun{ChesherDiscrete}; and \citeasnoun{ChesherSmolinski}. The latter results are related to random set/optimal transport methods for identification analysis; see \citeasnoun{BMM:2011}; \citeasnoun{EGH:2010}; \citeasnoun{GH:2009}; and \citeasnoun{GH:2011}.
The model and the results of Theorem 1 are useful for two reasons. First, Theorem 1 serves as a means of identifying the QTE in a reasonably general heterogeneous effects model. Second, by demonstrating that the IVQT model leads to the conditional moment restrictions ((ref)) and ((ref)), Theorem 1 provides an economic and causal foundation for estimation based on these restrictions.
The conditions presented above yield the following identification region for the structural quantile function $(d,x,\tau) \mapsto q(d,x, \tau)$. The identification region for the case of strictly increasing $\tau \mapsto q(d,x, \tau)$ can be stated as the set $\mathcal{Q}$ of functions $(d,x,\tau) \mapsto m(d,x,\tau)$ that satisfy the following relations, for all $\tau \in (0,1]$
This representation of the identification region $\mathcal{Q}$ is implicit. Nevertheless, statistical inference about $q \in \mathcal{Q}$ can be based on ((ref)) and can be carried out in practice using weak-identification robust inference as described in \citeasnoun{iqr:joe2}, \citeasnoun{SakataMarmer}, \citeasnoun{JunWeakIdIVQR}, \citeasnoun{SantosPartialIDIVQR}, or \citeasnoun{fsqr}. Under conditions that yield point identification, these regions collapse to a singleton, and the aforementioned weak-identification-robust inference procedures retain their validity.
The identification region for the case of weakly increasing $\tau \mapsto q(d,x, \tau)$ can be stated as the set $\mathcal{Q}$ of functions $(d,x,\tau) \mapsto m(d,x,u)$ that satisfy the following relations: For any closed subset $I$ of $(0,1]$,
where $m(D,X,I)$ is the image of $I$ under the mapping $\tau \mapsto m(D,X,\tau)$. The inference problem here falls in the class of conditional moment inequalities and approaches such as those described in \citeasnoun{AndrewsShi} or \citeasnoun{CLRIntersectionBounds}, for example, can be used. The sets $I$ to be checked could be reduced by determining approximate core-determining subsets; see \citeasnoun{ChesherRosenSmolinski}, \citeasnoun{GH:2009}, \citeasnoun{GH:2011} for further discussion.
Condition A1 imposes monotonicity on the structural function of interest which makes its relation to the QTR apparent. Condition A2 states that potential outcomes are independent of $Z$, given $X$, which is a conventional independence restriction. Condition A3 is a convenient representation of a treatment selection mechanism, stated for the purposes of discussion. In A3, the unobserved random vector $V$ is responsible for the difference in treatment choices $D$ across observationally identical individuals. Dependendence between $V$ and $\{U_d\}$ is the source of endogeneity that makes the conventional exogeneity assumption $U \sim U(0,1)|X,D$ break down. This failure leads to inconsistency of exogenous quantile methods for estimating the structural quantile function. Within the model outlined above, this breakdown is resolved through the use of instrumental variables.
The independence imposed in A2 and A3 is weaker than the commonly made assumption that both the disturbances $\{U_d\}$ in the outcome equation and the disturbances $V$ in the selection equation are jointly independent of the instrument $Z$; e.g. \citeasnoun{heckman:robb} and \citeasnoun{late}. The latter assumption may be violated when the instrument is measured with error as discussed in \citeasnoun{hausman:1977} or the instrument is not assigned exogenously relative to the selection equation as in Example 2 in \citeasnoun{late}.
Condition A4 restricts the variation in ranks across potential outcomes and is key for identifying the QTR and associated QTE. Its simplest, though strongest, form is rank invariance, when ranks $U_d$ do not vary with potential treatment states $d$:\footnote{Notice that under rank invariance, condition A3 is a pure representation, not a restriction, since nothing restricts the unobserved information component $V$.}
For example, under rank invariance, people who are strong (highly ranked) earners without a training program ($d=0$) remain strong earners having done the training ($d=1$). Indeed, the earnings of a person with characteristics $x$ and rank $U=\tau$ in the training state “0" is $Y_0 = q(0, x, \tau)$ and in the state “1" is $Y_1 =q(1, x, \tau)$.\footnote{Rank invariance is used in many interesting models without endogeneity. See e.g. \citeasnoun{doksum:1974}, \citeasnoun{heckman:smith}, and \citeasnoun{koenker:geling}.} Thus, rank invariance implies that a common unobserved factor $U$, such as innate ability, determines the ranking of a given person across treatment states.
Rank invariance implies that the potential outcomes $\{ Y_d\}$ are jointly degenerate which may be implausible on logical grounds, as pointed out by \citeasnoun{heckman:smith}. Also, the rank variables $U_d$ may be determined by many unobserved factors. Thus, it is desirable to allow the rank $U_d$ to change across $d$, reflecting some unobserved, asystematic variation. Rank similarity A4 achieves this property while managing to preserve the useful moment restriction ((ref)).
Rank similarity A4 relaxes exact rank invariance by allowing asystematic deviations, “slippages" in the terminology of \citeasnoun{heckman:smith}, in one's rank away from some common level $U$. Conditional on $U$, which may enter disturbance $V$ in the selection equation, we have the following condition on the slippages\footnote{Conditioning is required to be on all components of $V$ in the selection equation A3. }
In this formulation, we implicitly assume that one selects the treatment without knowing the exact potential outcomes; i.e. one may know $U$ and even the distribution of slippages, but does not know the exact slippages $U_d-U$. This assumption is consistent with many empirical situations where the exact latent outcomes are not known before receipt of treatment. We also note that conditioning on appropriate covariates $X$ may be important to achieve rank similarity.
In summary, rank similarity is an important restriction of the IVQT model that allows us to address endogeneity. This restriction is absent in conventional endogenous heterogeneous treatment effect models. However, similarity enables a more general selection mechanism, A3, and weaker independence conditions on instruments than often are assumed in nonseparable IV models. The main force of rank similarity and the other stated assumptions is the implied moment restriction ((ref)) of Theorem 1, which is useful for identification and estimation of the quantile treatment effects.
We present some examples that highlight the nature of the model, its strengths, and its limitations.
The purpose of this section is to examine the identifying power of conditional moment restrictions ((ref)). Specifically, we give various conditions for point identification in this section, summarizing and updating some of the results known in the literature. We remark here that point identification is not required in applications in principle as there exist inference methods that apply without point identification. However, it is useful to know and understand conditions under which moment conditions are informative enough that the identification region shrinks to a single point; in such cases the inference methods will also produce very informative confidence sets. We present point-identifying conditions first for the binary case, $D \in \{0,1\}$ and $Z \in \{0,1\}$, and then consider the case of $D$ taking a finite number of values, and finally consider the continuous case.
Here we consider the cases where $D \in \{0,1\}$ and $Z \in \{0,1\}$. The following analysis is all conditional on $X=x$ and for a given quantile $\tau \in (0,1)$, but we suppress this dependence for ease of notation. Under the conditions of Theorem 1, we know that there is at least one function $q(d):= q(d, x, \tau)$ that solves $ P[Y \leqslant q(D)|Z] = \tau \ \ \text{a.s}.$ The function $q(\cdot)$ can be equivalently represented by a vector of its values $q=(q(0), q(1))'$. Therefore, for vectors of the form $y = (y_0, y_1)'$, we have a vector of moment equations
where $y_D := (1-D)\cdot y_0 + D \cdot y_1$. We say that $q$ is identified in some parameter space, $\mathcal{L}$, if $y=q$ is the only solution to $\Pi(y)=0$ among all $y \in \mathcal{L}$.
We require that the Jacobian $\partial \Pi(y)$ of $\Pi(y)$ with respect to $y=(y_0, y_1)'$ exists and that it takes the form
For local identification, we take $\mathcal{L}$ as an open neighborhood of $q = (q(0), q(1))'$. For global identification, we shall use some definitions from Mas-Collell to define $\mathcal{L}$. In what follows, for every proper (non-null) subspace $L \subset \mathbb{R}^l$, let $\textrm{proj}_L : \mathbb{R}^l \mapsto L$ denote the perpendicular projection map. A convex, compact polytope is a bounded convex set formed by an intersection of a finite number of closed half-spaces. Such a polytope is of full dimension in $\mathbb{R}^l$ if it has a non-empty interior in $\mathbb{R}^l$. A face of a polytope $\mathcal{L}$ is the intersection of any supporting hyperplane of $\mathcal{L}$ with $\mathcal{L}$, so that faces of a polytope necessarily include the polytope itself. For instance, a rectangle in $\mathbb{R}^2$ has one 2-dimensional face given by itself, four 1-dimensional faces given by its edges, and four 0-dimensional faces gives by its vertices. A subspace spanned by a non-empty face of $\mathcal{L}$ is the translation to the origin of the minimal affine space containing that face.
The first result is a simple local identification condition of the type considered in \citeasnoun{rothenberg:1971} which we provide to fix ideas. The second result is a global identification condition which extends the result in \citeasnoun{iqr:ema} by allowing non-rectangular sets $\mathcal{L}$. This result is based on the global univalence theorems of \citeasnoun{mas-colell:convex}. As explained below, the positive determinant condition requires the impact of instrument $Z$ on the joint distribution of $(Y,D)$ to be sufficiently rich. In particular, the instrument $Z$ should not be independent of the endogenous variable $D$. We note that existence of the conditional density $f_{Y}( y |D=d,Z=z)$ is only required for $(d,z)$ in the support of $(D,Z)$. Outside the support we can define the conditional density as 0, so the existence condition is not very restrictive. Moreover, the condition is formulated so that $\mathcal{L}$ can take on relatively rich shapes that can carry useful economic restrictions. For instance, in the training context, a useful restriction on the parameters is that training weakly increases the potential earning quantiles. This restriction can be implemented by taking some natural parameter space and intersecting it with the half-space $H= \{ (y_0, y_1) \in \mathbb{R}^2: y_1 \geqslant y_0\}$. Specifically, a cube $C = \{ y \in \mathbb{R}^l: \| y\|_{\infty} \leqslant K\}$ intersected with the halfspace $H$ is an example of a region $\mathcal{L}$ permitted by the global identification result (ii).
We generalize the result of Theorem 2 to more general discrete treatments with discrete instruments. Consider the case when $D$ has the support $\{1,...,l\}$ and $Z$ has the support $\{1,...,r\}$ ($l \leqslant r < \infty$). Note that function $q(\cdot)$ can be represented by a vector $q=(q(1),...,q(l)) '\in \mathbb{R}^l$. Under the conditions of Theorem 1, there is at least one function $q(d)$ that solves $ P[Y \leqslant q(D)|Z] = \tau \ \ \text{a.s}.$ Therefore, for vectors of the form $y = (y_1,...,y_l)'$ and the vector of moment equations
where $y_{D} := \sum_{d} 1[D=d]\cdot y_d$, the model is identified if $y=q$ uniquely solves $\Pi (y)=0$.
We define matrix $\partial \Pi(y)$ as the $r \times l$ matrix with $(d,z)$ element given by $f_{\scriptscriptstyle Y}(y_d|D=d,Z=z) P[D=d|Z=z]$ where $z=1,...,r$ and $d =1,...,l$. We require this to be the Jacobian matrix of the map $y \mapsto \Pi(y)$ and impose full-rank-type conditions on submatrices of this Jacobian. To this end, let $m$ denote any permutation of $l$ distinct integers from $\{1,...,r\}$, called $l$-permutations, and $\mathcal{M}$ be a collection of all such permutations. Let $\Pi_m : = (\Pi_j)_{j \in m}$, which maps $\mathbb{R}^l$ to $\mathbb{R}^l$, be a subvector of $\Pi$ formed by selecting $j$-th elements of $\Pi$ according to their order in $m$.\footnote{Note that this formulation allows reordering elements of $\Pi$ which may be needed to achieve the required positive determinant condition as discussed in the binary case.} Let $\partial \Pi_m$ denote the corresponding $l \times l$ Jacobian matrix of $\Pi_m$. The following theorem generalizes Theorem 2.
We note that in the theorem existence of the conditional density $f_{Y}( y |D=d,Z=z)$ is only required for $(d,z)$ in the support of $(D,Z)$. This density can be defined to take on an arbitrary value for $(d,z)$ outside the support. The first result is a simple local identification condition provided to fix ideas. The second result is a global identification condition based on Global Univalence Theorem 1 of \citeasnoun{mas-colell:convex}. This result complements a similar result given in \citeasnoun{iqr:ema} based on Global Univalence Theorem 2 of \citeasnoun{mas-colell:convex}. The positive determinant condition requires the impact of instrument $Z$ on the joint distribution of $(Y,D)$ to be sufficiently rich.
Finally we consider conditions for point identification in the case of more general $D$ and $Z$ that may take on a continuum of values. We let $d$ denote elements in the support of $D$ and $z$ denote elements in the support of $Z$. Without loss of much generality, we restrict attention to the case where both $Y$ and $D$ have bounded support. We require the parameter space $\mathcal{L}$ to be a collection of bounded (measurable) functions $m: \mathbb{R}^k \mapsto \mathbb{R}$ containing $q(\cdot)$. We say that $q(\cdot)$ such that $P[Y \leqslant q(D)|Z]= \tau$ a.s. is identified in $\mathcal{L}$ if for any other $m(\cdot) \in \mathcal{L}$ such that $P[Y \leqslant m(D)|Z]=\tau$ a.s., $m(D)=q(D)$ a.s. Below, we use $\|\cdot \|_{p,P}$ to denote the $L^p(P)$ norm.
Condition (i), mentioned in \citeasnoun{iqr:ema}, states a non-linear bounded completeness condition for global identification. The condition ((ref)) required is not primitive, but it highlights a useful link with the linear bounded completeness condition: $E\left[ \Delta(D)|Z\right]=0 \text{ a.s. } \Rightarrow \Delta(D)=0 \text{ a.s.}$ used by \citeasnoun{newey:powell}. The latter condition is needed for identification in the mean IV model $E [Y - q(D)|Z ] = 0$ under the assumption of a bounded structural function $q$. The latter condition is known to be quite weak, as shown in \citeasnoun{Hault}, and there are many primitive sufficient conditions that imply this condition. \citeasnoun{andrews:genericity} shows that linear completeness is generic under some conditions. Although condition ((ref)) is not primitive, it is not vacuous either since the previous theorems provide primitive conditions for its validity. The local identification condition (ii), obtained by \citeasnoun{CCLN}, provides yet another sufficient condition for condition (i). The result (ii) replaces the nonlinear completeness condition ((ref)) by the linear completeness condition ((ref)) which is easier to check. The result (ii) also implicitly requires that the set $\mathcal{L}$ is a sufficiently small neighborhood of $q$ and that functional deviations $m(\cdot)-q(\cdot)$ and the conditional density $f_{\epsilon}(\cdot|D,Z)$ are sufficiently smooth. This is explained in detail in \citeasnoun{CCLN} where further primitive smoothness and completeness conditions are also provided.
There are, of course, other sets of modeling assumptions that one could employ to build a quantile model with endogeneity. In this section, we briefly outline two other approaches that have been taken in the literature. The first, due to \citeasnoun{abadie}, extends the local average treatment effect (LATE) framework of \citeasnoun{late} to quantile treatment effects. The second, considered in \citeasnoun{imbens:newey} and \citeasnoun{SLeeTriangularIVQR}, uses a triangular structure to obtain identification.
In fundamental work, \citeasnoun{abadie} develop an approach to estimating quantile treatment effects within the LATE framework of \citeasnoun{late} in the case where both the instrument and treatment variable are binary. The use of the LATE framework makes this approach appealing as many applied researchers are familiar with LATE and the conditions that allow identification and consistent estimation of this quantity. Importantly, the extension proceeds under exactly the same monotonicity requirement as needed for LATE.
Specifically, \citeasnoun{abadie} show that the QTE for a subpopulation is identified if
The subpopulation for whom the QTE is identified is the set of “compliers,” those individuals with $D_1 > D_0$. In other words, the compliers are the set of individuals whose treatment is altered by switching the instrument from zero to one. Monotonicity is key in this framework. The monotonicity condition rules out “defiers,” individuals who would receive treatment in the absence of the intervention represented by the instrument but would not receive treatment if placed into the treatment group. The effects for individuals who would always receive treatment or never receive treatment regardless of the value of the instrument are unidentified.
Looking at these conditions, we see that the model of \citeasnoun{abadie} replaces the monotonicity assumption (A1), the independence assumption (A2), and the similarity assumption (A4) with a different type of monotonicity and a stronger independence assumption and identifies a different quantity: the QTE for compliers. The LATE-style approach has not yet been extended beyond cases with a binary treatment and a single binary instrument while the instrumental variable quantile model of \citeasnoun{iqr:ema} applies to any endogenous variables and instruments. Note that neither set of conditions nests the other, and neither framework is more general than the other. Thus, the frameworks are best viewed as complements, providing two sets of conditions that can be considered when thinking about a strategy for estimating heterogeneous treatment effects.
Of course, the two sets of conditions may be mutually compatible. One such case is discussed in \citeasnoun{401k}. In this example, the pattern of results obtained from the two estimators is quite similar, and the difference between the estimates appears small relative to sampling variation. Further exploration of these two approaches and their similarities and differences may be interesting to consider.
Another compelling framework is based on assuming a triangular structure as in \citeasnoun{imbens:newey}. See also \citeasnoun{chesher}, \citeasnoun{koenker:ma}, and \citeasnoun{SLeeTriangularIVQR} for related models and results. The triangular model takes the form of a triangular system of equations
where $Y$ is the outcome, $D$ is a continuous scalar endogenous variable, $\epsilon$ is a vector of disturbances, $Z$ is a vector of instruments with a continuous component, $\eta$ is a scalar reduced form error, and we ignore other covariates for simplicity. It is important to note that the triangular system generally rules out simultaneous equations which typically have that the reduced form relating $D$ to $Z$ depends on a vector of disturbances. For example, in a supply and demand system, the reduced form for both price and quantity will generally depend on the unobservables from both the supply equation and the demand equation. Outside of $\eta$ being a scalar, the key conditions that allow identification of quantile effects in the triangular system are
The variable $V$ is thus the “control function” conditional on which changes in $D$ may be taken as causal. \citeasnoun{imbens:newey} use $V = F_{D|Z}(d,z) = F_{\eta}(\eta)$, where $F_{\eta}(\cdot)$ represents the CDF of $\eta$, as the control function and show that this variable satisfies the independence condition under the additional condition that $(\epsilon,\eta)$ is independent of $Z$. They show that one may use $D = h(Z,\eta)$ to identify $V$ under the assumed monotonicity of $h(Z,\eta)$ in $\eta$. Using $V$ obtained in this first step, one may then construct the distribution of $Y|D,V$. Then integrating over the distribution of $V$ and using iterated expectations, one has
It then follows that the $\tau^{th}$ quantile of $Y_d$ is $G^{-1}(\tau,d)$.
As with the framework of \citeasnoun{abadie}, the triangular model under the conditions given above is neither more nor less general than the model of \citeasnoun{iqr:ema}. The key difference between the approaches is that \citeasnoun{iqr:ema} uses an essentially unrestricted reduced form but requires monotonicity and a scalar disturbance in the structural equation. The triangular system on the other hand relies on monotonicity of the reduced form in a scalar disturbance. In addition, the triangular system, as developed in \citeasnoun{imbens:newey}, requires a more stringent independence condition in that the instruments need to be independent of both the structural disturbances and the reduced form disturbance. That the approaches impose structure on different parts of the model makes them complementary with a researcher's choice between the two being dictated by whether it is more natural to impose restrictions on the structural function or the reduced form in a given application.
The triangular model and the model of \citeasnoun{iqr:ema} can be made compatible by imposing the conditions from the triangular model on the reduced form and the conditions from \citeasnoun{iqr:ema} on the structural model. \citeasnoun{torgovitsky:id} considers identification and estimation when both sets of conditions are imposed and shows that the requirements on the instruments may be substantially relaxed relative to \citeasnoun{iqr:ema} or \citeasnoun{imbens:newey} in this case.
In the previous sections, we have outlined results that are useful for identifying quantile treatment effects and structural functions that are monotonic in a scalar unobservable. In the following, we briefly review the literature on estimation and inference. We focus on estimation of the model of \citeasnoun{iqr:ema} presented in Section 2 using the moment conditions derived in Theorem 1. For estimation of the triangular model, see \citeasnoun{imbens:newey} for nonparametric estimation and \citeasnoun{SLeeTriangularIVQR} for a semiparametric approach. \citeasnoun{abadie} provides results for estimating the QTE for compliers within the LATE-style framework. Also, we only review approaches for estimating parametric quantile functions: $q(D,X,\tau) = g(D,X,\tau;\theta)$ for $\theta \in \Theta \subset \mathbb{R}^m$. \citeasnoun{HorowitzLeeNonparametricIVQR} and \citeasnoun{GagliardiniScaillet} present nonparametric estimation and inference results for the IVQT model using condition ((ref)).
There are two practical issues that make estimation and inference based on condition ((ref)) challenging. The first is that the sample analog to condition ((ref)) is non-smooth, and the GMM objective function that would be formed by using ((ref)) as the moment conditions is also generically non-convex, even for linear quantile models. The second problem is that the model may suffer from weak identification as in the standard linear IV model; \citeasnoun{stock:survey} provides a useful introductory survey to weak identification and related inference methods in the linear IV model. In the quantile case, the problem of weak identification is more subtle than in the linear model in that some quantiles may be weakly identified while others may be strongly identified. The relevant object for defining the strength of identification of a given quantile is the covariance between $D$ and $Z$ weighted by the conditional density function of the unobservable at the given quantile. See \citeasnoun{iqr:joe2} for a formal definition of this object and related discussion.
While the non-smoothness and non-convexity of the GMM criterion complicates optimization, it does not render the approach infeasible, especially when the dimension of $D$ and $X$ is not too large. \citeasnoun{abadie:1997} considered this approach for estimating an income model and provides further discussion. One could also estimate the model parameters using the Markov Chain Monte Carlo (MCMC) approach of \citeasnoun{mcmc}. This approach bypasses the need for optimization, instead relying on sampling and averaging to estimate model parameters. Note that this approach is not a cure-all since MCMC requires careful tuning in applications. It is also worth noting that standard samplers may perform poorly in even simple linear instrumental variables models when identification is not strong; see \citeasnoun{vanDijkIVPosteriors}. In an approach related to optimizing the GMM criterion function directly, \citeasnoun{sakata} proposes estimating the parameters of an instrumental variables quantile model by optimizing a different non-smooth, non-convex criterion function.
To partially circumvent the numerical problems in optimizing the full GMM criterion, \citeasnoun{iqr:joe} suggest a different procedure termed the inverse quantile regression for the linear quantile model $q(D,X,\tau) = D'\alpha(\tau) + X'\beta(\tau)$. The basic intuition for the inverse quantile regression comes from the observation that if one knew the true value of the coefficient on $D$, $\alpha(\tau)$, the $\tau^{\textnormal{th}}$ quantile regression of $Y - D'\alpha(\tau)$ onto $X$ and $Z$ would yield zero coefficients on the instruments $Z$. This observation allows one to effectively concentrate $\beta(\tau)$ out of the problem and leaves a non-smooth, non-convex optimization problem over only the parameters $\alpha(\tau)$. Since $D$ is low-dimensional in many applications, one can usually solve this optimization problem using highly robust optimization procedures such as a grid-search.
Algorithmically, the inverse quantile regression estimates for a given probability index $\tau$ of interest can be obtained as follows using a grid search over $\alpha(\tau)$:
1. Define a suitable set of values $\{\alpha_j, j=1,...,J\}$, and estimate the coefficients $\beta(\alpha_j,\tau)$ and $\gamma(\alpha_j,\tau)$ from the model $Y-D'\alpha_j = X'\beta(\alpha_j,\tau) + Z'\gamma(\alpha_j,\tau) + \epsilon$ by running the ordinary $\tau$-quantile regression of $Y-D'\alpha_j$ on $X$ and $Z$. Call the estimated coefficients $\widehat \beta(\alpha_j, \tau)$ and $\widehat \gamma(\alpha_j, \tau)$.
2. Save the inverse of the variance-covariance matrix of $\widehat \gamma(\alpha_j, \tau)$, which is readily available in any common implementation of the ordinary QR. Denote this variance-covariance matrix $\widehat A(\alpha_j,\tau)$. Form $W_n(\alpha_j,\tau) = \widehat\gamma(\alpha_j,\tau)'\widehat A(\alpha_j,\tau)^{-1}\widehat\gamma(\alpha_j,\tau)$. Note $W_n(\alpha_j)$ is the Wald statistic for testing $\gamma(\alpha_j, \tau)=0$.
3. Choose $\widehat \alpha (\tau)$ as a value among $\{\alpha_j, j=1,...,J\}$ that minimizes $W_n(\alpha,\tau)$. The estimate of $\beta(\tau)$ is then given by $\widehat \beta(\widehat \alpha (\tau), \tau)$.
\citeasnoun{iqr:joe} and \citeasnoun{iqr:joe2} provide conditions under which the resulting estimator for $\alpha(\tau)$ and $\beta(\tau)$ is consistent and asymptotically normal and provide a consistent variance estimator. \citeasnoun{SakataMarmer} provide a similar multi-step algorithm that circumvents the same numeric problems using the objective function of \citeasnoun{sakata}.
The good behavior of the asymptotic approximations obtained in \citeasnoun{iqr:joe} and \citeasnoun{iqr:joe2} rely on strong identification of the model parameters just as in the linear IV case. Intuitively, strong identification for a quantile of interest requires that a particular density-weighted covariation matrix between $D$ and $Z$ is not local to being rank deficient and that the impact of $Z$ is rich enough to guarantee that the moment equations have a unique solution. The first condition is analogous to the usual full rank condition in linear IV analysis, and the second condition is required because of the nonlinearity of the problem. Checking these conditions in practice may be difficult, and it is therefore useful to have inference procedures that are robust to violations of these conditions.
Fortunately, there are several inference procedures that remain valid under weak identification. A nice feature of the algorithm defined for estimating $\alpha(\tau)$ above is that it produces a weak-identification-robust inference procedure naturally as a byproduct. \citeasnoun{iqr:joe2} show that the Wald statistic, $W_n(\alpha,\tau)$ converges in distribution to $\chi^2_{\dim(Z)}$ under the null that $\alpha = \alpha_0$ where we let $\alpha_0$ denote the true value of $\alpha(\tau)$ without needing either of the conditions discussed in the preceding paragraph. Thus a valid $(1-p)\%$ confidence region for $\alpha(\tau)$ may be constructed as the set:
where $c_{1-p}$ is such that $\textnormal{Pr}(\chi^2_{\dim(Z)} > c_{1-p}) = p$, and the set is approximated numerically by considering $\alpha$'s in the grid $\{\alpha_j, j =1,...,J\}$. \citeasnoun{iqr:joe2} show that confidence region in equation ((ref)) is valid when the model parameters are strongly identified and remains valid when the model is weakly identified or even unidentified. \citeasnoun{SakataMarmer} provide a similar procedure and result for their procedure as well. \citeasnoun{JunWeakIdIVQR} provides yet a different approach to performing weak-identification-robust inference in models defined by conditions ((ref)). Finally, \citeasnoun{fsqr} show that one can form statistics for inference about the entire parameter vector $\theta$ that are condtionally pivotal in finite-samples for models defined by quantile restrictions such as ((ref)). Since the statistics do not depend on unknown nuisance parameters in finite samples, the exact distributions of these statistics can be calculated and inference can proceed without relying on asymptotic approximations or statements about the strength of identification. The distributions produced in \citeasnoun{fsqr} are not standard and so must be calculated by simulation.
In this paper, we have reviewed approaches for building quantile models in the presence of endogeneity, focusing on conditions that can be used for identification. We have also briefly reviewed some of the practical issues that arise in estimation of instrumental variables quantile models and approaches to dealing with these issues. The models and estimation strategies outlined and cited in this review have already seen use in empirical economics where they have mostly been used for their ability to uncover interesting distributional effects. In this review, we have also noted that the identification strategy employed in this paper can be used to uncover structural objects even if quantile effects are not the chief objects of interest as in \citeasnoun{BerryHaileDiscreteChoiceId}.
While the results reviewed in this paper are useful in a variety of contexts, there remain interesting areas for research in quantile models with endogeneity. In some applications, features of the conditional distribution are not the chief objects of interest and researchers are interested in effects of treatments on unconditional quantiles. Given the set of conditional quantiles, such unconditional effects may be uncovered. In recent work, \citeasnoun{FroelichMellyUnconditionalIVQR} propose a different approach, related to \citeasnoun{abadie}, to estimating structural effects of endogenous variables on unconditional quantiles directly. It would also be interesting to think about quantile-like quantities for multivariate outcomes with endogenous covariates. The results reviewed in this paper offer one possible approach for quantile modeling with endogeneity, but there remain many interesting directions and other approaches to be explored in further research.