Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
129,665 characters · 32 sections · 52 citation commands
Distributional Instruments: Identification and Estimation with Quantile Least Squares
A central promise of public health insurance is insurance, not just subsidies: it should protect households against the risk of large, unpredictable medical expenses. In the United States this promise is especially salient for older adults. Nearly all Americans over age 65 are enrolled in Medicare, yet many still face substantial exposure to out-of-pocket (OOP) medical spending, particularly for prescription drugs.\footnote{See, for example, finkelstein2008did on the initial impact of Medicare on OOP spending and jones2018lifetime on lifetime medical spending of retirees.} The introduction of Medicare Part D in 2006 was explicitly designed to reduce this financial risk by expanding subsidized prescription drug coverage and introducing catastrophic protection. A large empirical literature has since documented how Part D and related coverage changes affect drug utilization, adherence, and OOP spending ketcham2008medicare,engelhardt2011medicare,park2017medicare,carvalho2019impact, and a parallel literature shows that Medicare eligibility and supplemental coverage sharply reduce the probability of catastrophic medical expenditures and financial strain among the elderly barcellos2015effects,scott2021assessing,jones2019predicting.
Much of this work treats financial exposure itself---for example, the level, variance, or tail probability of OOP spending---as an outcome of policy. From a policy design perspective, however, we often want the reverse object: the causal effect of financial risk exposure on behavior and welfare. Does facing a higher risk of catastrophic prescription drug expenses change adherence, portfolio choice, or retirement decisions? Does improved financial risk protection from Part D or supplemental policies translate into better mental health or reduced financial distress?\footnote{See ayyagari2015does and ayyagari2016medicare for related evidence on mental health and portfolio choice.} Answering such questions requires treating a measure of risk exposure---for example, a catastrophic-spending indicator or a predicted OOP risk index---as an endogenous regressor and using policy variation as an instrument.
In practice, this approach runs into a familiar econometric obstacle. The most natural instruments for a risk index are policy indicators (e.g.\ Part D eligibility or plan generosity) and features of the plan menu. These instruments primarily shift the distribution of OOP spending---compressing right tails, changing dispersion, and altering the frequency of catastrophic events---while often having relatively modest effects on the mean of the scalar risk measure itself. Empirically, first-stage regressions of a single risk index on these policy variables can have very low $F$-statistics, even when the underlying policy clearly reshapes the full distribution of spending engelhardt2011medicare,barcellos2015effects,scott2021assessing. \footnote{ Table (ref) summarizes some representative studies showing mean effect of insurance on mean OOP spending. }Standard IV diagnostics based on mean shifts therefore label the design as “weak” and push researchers toward reduced-form analyses, regression discontinuity designs at age 65, or structural models of plan choice abaluck2011choice, rather than causal IV estimates for the effect of financial risk on outcomes.
This paper starts from the observation that, in such settings, the instruments are not weak in any substantive sense: they are strongly relevant for the distribution of the endogenous variable, even if they appear weak for its mean. We formalize this idea by replacing the classical notion of mean relevance---variation in $\mathbb{E}[X\mid Z]$---with a weaker and more general notion of distributional relevance: the conditional distribution $F_{X\mid Z}(\cdot\mid Z)$ of the endogenous variable $X$ shifts with the instrument $Z$ in an $L^2$ sense, even when $\mathbb{E}[X\mid Z]$ is constant. In the Medicare example, $X$ may be a scalar index of financial exposure constructed from the OOP distribution, while $Z$ captures policy-induced variation in coverage generosity. Part D and related reforms satisfy distributional relevance whenever they alter the shape (tails, variance, skewness) of the spending distribution, regardless of whether they strongly move its mean.
Our first contribution is to show that distributional relevance is a meaningful notion of instrument strength, not just a descriptive property. Within a non-separable triangular model, we characterize a class of purely distributional instruments: variables $Z$ for which $\operatorname{Var}(\mathbb{E}[X\mid Z])=0$ but $F_{X\mid Z}(\cdot\mid Z)$ is non-degenerate in $L^2$. Building on the control-function framework of cfnsv20_cfa, we prove that such instruments still identify average structural effects via the control variable $V=F_X(X\mid Z)$, and we relate this notion to existing concepts of instrument strength based on heteroskedasticity lewbel1997constructing. Thus, even when policy variables are “purely distributional” for a risk index, they can support point identification of its causal effect on outcomes.
Our second, and main, contribution is to develop a practical IV estimator tailored to these designs, which we call Quantile Least Squares (Q--LS).\footnote{Our notion of “distributional instruments" is orthogonal to the shift-share (Bartik) literature (see adao2019shift, borusyak2025practical, goldsmith2020bartik for a detailed review). There, the focus is on constructing a composite instrument from many shocks and exposure shares. Here, we take the instrument $Z$ as given and show how to optimally exploit its distributional impact on $X$ via conditional quantiles. In particular, Q–LS does not impose a shift–share structure on $Z$; it can be applied whether $Z$ is a simple policy indicator or a more complex shock.} Instead of summarizing the first stage with a conditional mean, Q--LS builds instruments from the entire conditional quantile process of $X$ given $Z$. Concretely, we: (i) estimate a series of first-stage quantile regressions of $X$ on $Z$; (ii) form a data-driven linear combination of these conditional quantiles that best predicts $X$ in mean square error; and (iii) use this optimal quantile-aggregated prediction as the instrument in an otherwise standard linear IV estimator. Under distributional relevance and a mild non-degeneracy condition, the resulting optimal Q--LS instrument is always relevant in the classical sense, even when the conditional mean of $X$ is flat in $Z$. In a “best-case” world where $X\mid Z$ is a pure location shift, we show that Q--LS collapses to optimal 2SLS and attains the same efficiency bound. In more general cases, Q--LS remains reliable precisely when mean-based instruments become weak.
Methodologically, we characterize the asymptotic properties of Q--LS in a linear triangular model with a single endogenous regressor. We show that (i) under distributional relevance and mild regularity conditions, the Q--LS estimator is consistent and asymptotically normal; (ii) standard heteroskedasticity-robust 2SLS standard errors computed with the generated Q--LS instrument are valid for inference; and (iii) ill-posedness in the optimal-weight problem can be handled with ridge or LASSO regularization across quantiles, yielding stable finite-sample performance. We propose simple diagnostics and a decision chart that place Q--LS alongside conventional weak-IV tools: researchers can test for mean relevance, test for distributional relevance, and then choose between classical 2SLS and Q--LS (and between its regularized variants) based on these diagnostics. In this sense, Q--LS is designed to complement, not replace, the existing 2SLS toolkit.
Our Monte Carlo experiments are calibrated to the Medicare context and compare Q--LS to conventional 2SLS and a control-function estimator in designs where instruments primarily shift the distribution of risk rather than its mean. In designs with strong mean relevance, Q--LS and 2SLS deliver nearly identical point estimates and standard errors, confirming that Q--LS does not sacrifice efficiency when classical IV works well. In designs where traditional first-stage $F$-statistics indicate severe weak-instrument problems, Q--LS produces well-centered estimates and near-correct coverage, while 2SLS exhibits substantial bias and undercoverage. Regularized Q--LS variants further stabilise performance when the quantile dictionary is rich relative to sample size.
Finally, we illustrate the usefulness of distributional instruments in the Medicare Part D setting. Using the Health and Retirement Study (HRS), we construct individual-level measures of OOP medical spending risk in real (2015) dollars and exploit Part D as a source of exogenous variation in risk exposure. We show, first, that standard scalar risk measures, such as mean OOP spending or catastrophic-spending indicators, can render an otherwise informative policy design “weak” by classical first-stage diagnostics, even when the upper tail of the OOP distribution clearly shifts. Second, we demonstrate how Q--LS, built from a finite dictionary of conditional quantiles of OOP spending, recovers the same causal effects as linear 2SLS in specifications where the real OOP first stage is strong, while delivering much tighter confidence intervals and more stable estimates in specifications where mean relevance is weak. Substantively, our estimates indicate that higher OOP risk is associated with worse mental health among older Americans. To our knowledge, this is the first empirical illustration of how distributional instruments can materially improve the precision and credibility of IV estimates in applied work on medical spending risk.
\paragraph{Contributions.} Our findings contribute to two main parts of the instrumental‐variables literature. First, on the identification side, building on the control–function results of cfnsv20_cfa, we highlight that their argument continues to apply even when instruments are purely distributional: even when $\operatorname{Var}(\mathbb{E}[X\mid Z]) = 0$ and the conditional mean of the endogenous variable is completely flat in the instrument, nontrivial variation in the conditional distribution $F_{X\mid Z}(\cdot\mid Z)$ is sufficient to identify average structural effects in a nonseparable triangular model via a control function based on $F_{X\mid Z}(X\mid Z)$. Our contribution is to make this “purely distributional’’ case explicit and to connect it to weak–IV diagnostics in linear models, where instruments may look weak in the mean yet be strong in the distribution. This formalizes a notion of distributional relevance that extends classical mean relevance in a direction that is natural for risk and tail–outcome applications.
Second, on the linear IV side, we develop a simple, implementable estimator, Quantile Least Squares (Q--LS), which constructs an optimal distributional instrument from conditional quantiles and then uses it in a conventional linear IV step. Under a strong-distributional-relevance condition, Q--LS shares the same large-sample behavior as 2SLS when instruments are strong in the mean, but in designs where policy reforms mainly reshape the distribution of risk it can extract a much stronger effective first stage from the same variation, complementing---rather than replacing---existing weak-IV diagnostics and robust inference tools. We characterize its finite–sample behavior through Monte Carlo designs.
Our empirical application illustrates how Q--LS can extract a strong first stage from policy variation that looks weak in classical mean-based diagnostics, while collapsing back to conventional IV when the instrument is strong in the mean.
The remainder of the paper is organized as follows. Section (ref) discusses weak identification, quantile regression, and some existing ways to circumvent weak identification. Section (ref) introduces the linear IV setup and formalizes distributional relevance. Section (ref) defines the Q--LS estimator, presents its asymptotic properties, and compares Q--LS to classical 2SLS and shows that they coincide in a Gaussian/location setting. Section (ref) discusses regularization, diagnostics, and practical implementation. Section (ref) outlines the simulation designs and results. Section (ref) presents the Medicare Part D application. Section (ref) concludes.
The aim of this research is to develop a general framework that circumvents the classical notion of weak identification of average structural functions in an instrumental variables design. In this section, we first describe and characterize weak identification in triangular models. Second, we summarize quantile regression and discuss how it has been used to address endogeneity in structural functions. Finally, we discuss how our proposed method relates to the existing literature that imposes strong functional form restrictions on the underlying model to circumvent the problem of weak identification.
Let $(Y,X,Z)$ be random variables that follow a triangular system \[ X = g(Z,\eta), \qquad Y = f(X,\varepsilon), \qquad (\eta,\varepsilon)\perp Z, \] where $f$ and $g$ are unknown measurable structural functions, and $\varepsilon,\eta$ are unobserved disturbances. The vector $Z$ collects the instruments. The average structural function (ASF) is \[ \mu(x) := \mathbb{E}\big[f(x,\varepsilon)\big], \] so identification of $\mu(\cdot)$ hinges on how variation in $Z$ propagates to variation in $X$ through $g(\cdot)$ and, ultimately, to variation in $Y$.
In nonseparable triangular models, identification of $\mu(\cdot)$ often proceeds via control-function or distribution-regression arguments that construct a control variable $V=V(X,Z)$ (for example $V=F_{X\mid Z}(X\mid Z)$) such that $Y \perp Z \mid (X,V)$ and suitable completeness conditions hold; see, among others, newey2003instrumental,newey2021control, florens2008identification, imbens2009identification, cfnsv20_cfa, and dunker2023nonparametric. In this framework, weak identification arises when the mapping from $Z$ to the distribution of $X$ is too “flat’’ to generate informative variation in the control function. Intuitively, if changes in $Z$ induce only very small or nearly collinear shifts in $F_{X\mid Z}(\cdot\mid z)$, then the associated conditional moment conditions are nearly redundant and the inverse problem that recovers $\mu(\cdot)$ becomes ill–posed or nearly unidentified. This is the nonlinear analogue of weak instruments: the conditional distribution of $X$ given $Z$ is only weakly sensitive to $Z$, so the control function has little effective variation.
Closest in spirit to our work is the recent “Distributional Instrumental Variable (DIV)” method of holovchak2025div. They also exploit distributional shifts in $X\mid Z$ and show how to identify and estimate the entire interventional distribution of the outcome using flexible generative models in a nonlinear IV setting. DIV provides conditions under which the interventional distribution is identified, and illustrates designs where 2SLS fails but distributional information suffices for identification and accurate estimation of mean and quantile treatment effects. Relative to this approach, we focus on average structural effects in a triangular model and on the construction of an optimal scalar instrument for linear IV estimands. Our Q--LS procedure can be viewed as a low-dimensional, analytically tractable counterpart to DIV: it concentrates the distributional information in a finite-dimensional quantile dictionary and delivers an optimal $L^2$-projection that can be used within the classical IV toolkit. Moreover, when the endogenous regressor is (or can be reduced to) a binary treatment, the Q--LS index provides a scalar monotone instrument, so that the resulting 2SLS coefficient admits a standard LATE interpretation as a positive-weighted average of complier effects across index increments.
To connect with the more familiar linear IV setting, consider the linear special case \[ Y = X\beta + W'\gamma + u, \qquad X = Z'\pi + W'\delta + v, \] where $Z$ are excluded instruments, $W$ are included exogenous regressors, and $\mathbb{E}[u\mid Z,W]\neq 0$ while $\mathbb{E}[v\mid Z,W]=0$. Classical weak-instrument problems arise when the first-stage coefficients $\pi$ are small, so that the concentration parameter is close to zero and the instruments explain only a tiny fraction of the variation in $X$; see staiger1997instrumental, stock2005asymptotic, and dufour1997some for canonical analyses. In this case, 2SLS and related estimators exhibit large finite-sample bias toward the OLS estimand, heavy-tailed or multimodal sampling distributions, and nonstandard asymptotic behavior andrews2014weak,andrews2019weak. These issues are amplified in many-instrument settings bekker1994alternative,hansen2008estimation, where the effective first-stage signal per instrument becomes small even if the joint $F$-statistic is sizable.
The weak-identification phenomena in triangular models are conceptually similar but operate at the level of distributions rather than means. Instead of small first-stage coefficients on $Z$ in a linear projection of $X$ on $Z$, the problem is that the instrument-induced shifts in $F_{X\mid Z}(\cdot\mid z)$ are too limited or nearly fail completeness, so the mapping from structural objects to observables is nearly singular. As emphasized by l19_id_zoo, both linear and nonlinear models admit a continuum of identification strengths, ranging from strong point identification with regular asymptotics, through weak or set identification with nonstandard limits, to complete lack of identification.
Our Q--LS estimator can be viewed as a two-step device that first maps the original instrument vector $Z$ into a scalar distributional instrument \[ h(Z) \in \mathcal{H} = \Big\{ \int_0^1 \omega(\tau) Q_{X\mid Z}(\tau\mid Z)\,d\tau : \omega \in L^2(0,1)\Big\}, \] and then applies conventional linear IV using $h(Z)$ in place of $Z$. From this perspective, Assumption (ref) is the analogue of a strong-IV condition on the generated instrument $h(Z)$: it rules out sequences of designs in which the covariance between $X$ and $h(Z)$ drifts to zero at a $\sqrt{n}$–local rate.
We therefore do not claim that Q--LS is generally robust in the weak-IV sense of staiger1997instrumental or that it uniformly dominates Anderson--Rubin or conditional tests andrews2019weak,olea2013robust when instruments are arbitrarily weak. Rather, the contribution of our distributional relevance framework is to highlight a different margin of strength: designs in which $\operatorname{Var}(\mathbb{E}[X\mid Z])$ is small or zero, so that instruments appear weak for the mean, may still be strong for the distribution, with $\mathbb{E}[h_{\mathrm{opt}}(Z)^2]$ bounded away from zero. In such cases, Q--LS leverages distributional variation in $X\mid Z$ to construct an effective instrument $h(Z)$ that satisfies a strong-IV condition even when the original mean-based instrument fails conventional first-stage diagnostics. Our notion of distributional relevance thus specifies a “strong-instrument’’ regime at the distributional level and complements both the generative DIV approach and the classical weak-IV toolkit. This allows us to focus on designs where instruments may be weak for linear first stages but remain informative about the distribution of $X$, and hence can still support identification and efficient estimation once the first stage is recast in distributional terms.
Quantile regression was first formalised as an optimization problem minimizing asymmetrically weighted absolute residuals by kb78_qr, to estimate any chosen quantile $\tau\in(0,1)$ of the distribution of an outcome variable $X$ conditional on a set of exogenous regressors $Z$,
where $F_{X|Z}(x|z)$ is the conditional cumulative distribution function (CDF) of $X$ given $Z=z$.
The linear quantile regression model specifies
where $\pi(\tau)$ is estimated by solving the expected check-loss function
This objective function is convex but non-differentiable, yielding a linear programming problem that can be solved efficiently in moderately large samples.
Identification of quantile regression relies on mild regularity conditions, including correct specification of the conditional quantile function, sufficient variation in the regressors, and a rank condition ensuring uniqueness of the conditional quantile. These conditions do not require parametric assumptions on the error distribution
aai02_qiv were the first to combine quantile methods with instrumental variables to estimate distributional treatment effects—that is, how a binary treatment affects quantiles of the outcome distribution—under endogeneity. ch05_qiv generalized this approach to estimate treatment effects at different points of the outcome distribution (quantile treatment effects) under endogeneity and introduced the notion of the structural quantile function. See w20 for a detailed comparison of ch05_qiv and aai02_qiv. Since these seminal contributions, many variants of quantile instrumental variable methods have been proposed. For example, p20_uncon_qiv develop an instrumental-variable approach to estimate unconditional quantile treatment effects, while clp16_gqiv propose methods for group-level quantile treatment effects.
There is also a growing literature that integrates over a grid of quantiles to estimate endogenous structural functions. cfk15_qiv propose a Censored Quantile Instrumental Variables (CQIV) estimator that uses a control-function approach to estimate structural quantiles of a censored outcome with endogenous regressors. In their framework, the control function is estimated by integrating over a grid of quantiles. cfnsv20_cfa extend the control-function framework of cfk15_qiv and develop a unified approach for estimating average, quantile, and distributional structural functions. The control-function estimator we propose for nonseparable triangular models estimates the control function using the same strategy as cfnsv20_cfa. See Appendix (ref) for further details.
There is a substantial literature that seeks to circumvent the problem of weak identification by imposing specific functional form restrictions. In this section, we discuss several key strands of this literature and how it relates to our proposed notion of distributional relevance.
There is a growing literature that imposes functional form restrictions on higher moments. lewbel1997constructing show that, in linear models with measurement error and endogeneity, instruments can be constructed from second-moment variation: \[ Z_i^{\text{Lev}} = (Z_i - \bar Z)(X_i - \bar X), \] provided that the variation of $X$ conditional on $Z$ changes with $Z$ even when $\mathbb{E}[X\mid Z]$ is constant. This identification strategy exploits variance shifts. Several studies have extended lewbel1997constructing and used second-moment restrictions to achieve identification under weaker functional form restrictions; see, for example, kv10_iv_cfa_hetro,l12_hetro. These restrictions can be viewed as a specific form of distributional relevance (variance shifts) and are therefore encompassed by our framework.
t23_iv_icm also impose a functional form restriction that ensures identification in models of the form $\mathbb{E}[Y - X\beta \mid Z]=0$ without excluded instruments. Specifically, t23_iv_icm assume a nonlinear completeness condition requiring injectivity of the conditional mean operator:
\[ \mathbb{E}[X\mid Z] \tau = 0 \quad\Rightarrow\quad \tau = 0. \]
Completeness relies on sufficiently rich nonlinear mean dependence of $X$ conditional on $Z$, that is, on the mapping $X \mapsto \mathbb{E}[X\mid Z]$.
Our approach is fundamentally different from t23_iv_icm. Identification does not rely on the conditional mean operator or its injectivity. Instead, distributional variation in $X$ suffices to identify the average structural function $\mu(x)$, even when $\mathbb{E}[X\mid Z]$ is constant and completeness fails. Consequently, our results do not require completeness, mean dependence, or any functional-form restrictions of the type imposed by t23_iv_icm.
babii2020completeness study estimation in nonidentified linear inverse problems of the from $K\varphi = r$, where $K$ is a compact operator between Hilbert spaces. In nonparametric IV models, $K$ corresponds to the conditional expectation operator $(K\varphi)(z)=\mathbb{E}[\varphi(X)\mid Z=z]$. When completeness (injectivity of $K$) fails, $\varphi$ is not point-identified, but spectral regularization methods converge to the “best approximation” $\varphi_1$ in the orthogonal complement of the null space of $K$. Their work characterizes risk bounds and asymptotic distributions under varying degrees of identification.
Our framework differs in two central respects. First, identification is not formulated as an inverse problem with operator $K\varphi=r$. Instead, we use a structural triangular model with distributional relevance, which yields full nonparametric identification of the average structural function $\mu(x)$ when the usual nonparametric IV completeness condition fails. Second, whereas babii2020completeness obtain convergence to a best approximation under nonidentification, our results deliver point identification of $\mu(x)$ by exploiting the rank structure of the first stage and the distributional variation of $X$. Consequently, our contribution is complementary: we provide an identification and estimation framework that does not rely on operator injectivity or completeness, and remains applicable with instruments that are mean-irrelevant but distributionally strong.
We consider the linear triangular model
where $Y$ is the outcome, $X$ is a scalar endogenous regressor, $Z_1$ is a vector of exogenous covariates included in the outcome equation, and $Z = (Z_1, Z_2)$ collects $Z_1$ and a vector of excluded instruments $Z_2$.
In the classical IV framework, instrument strength is measured by variation in the conditional mean of $X$ given $Z$: \[ m(Z) := \mathbb{E}[X\mid Z], \qquad \operatorname{Var}(m(Z)) > 0. \] When $\operatorname{Var}(m(Z))$ is close to zero, instruments are labelled “weak” and standard 2SLS inference is unreliable.
Our aim is to replace this mean-based notion of strength with a weaker and more general concept based on the full conditional distribution of $X$.
Let \[ F_{X\mid Z}(x\mid Z) := \mathbb{P}(X \le x \mid Z), \qquad \bar F_X(x) := \mathbb{E}\big[ F_{X\mid Z}(x\mid Z)\big] \] denote the conditional and unconditional distribution functions of $X$. We measure variation in $F_{X\mid Z}(\cdot\mid Z)$ using an $L^2$ norm over $x$. For concreteness, let $\lambda$ be Lebesgue measure on a compact interval containing the support of $X$, and define \[ \| F_{X\mid Z}(\cdot\mid Z) \|_{L^2(\lambda)}^2 := \int \big( F_{X\mid Z}(x\mid Z) \big)^2\,d\lambda(x). \]
Any change in $\mathbb{E}[X\mid Z]$ across $Z$ induces a difference in $F_{X\mid Z}(\cdot\mid Z)$, so mean relevance implies distributional relevance. The converse need not hold: $Z$ can be distributionally relevant even when $\mathbb{E}[X\mid Z]$ is constant in $Z$. This motivates the following concept.
Classical IV analysis treats mean relevance---variation in $\mathbb{E}[X\mid Z]$---as the benchmark notion of instrument strength. Distributional relevance weakens this requirement: it allows $\mathbb{E}[X\mid Z]$ to be completely flat in $Z$, as long as the conditional distribution $F_{X\mid Z}(\cdot\mid Z)$ varies nontrivially in $L^2$. In this sense, an instrument can be “strong” for the distribution of $X$ even when it is “weak” for its mean. We refer to instruments that satisfy distributional relevance but have $\operatorname{Var}(\mathbb{E}[X\mid Z])=0$ as purely distributional: they convey no information about $\mathbb{E}[X\mid Z]$ but still transmit rich information about the risk environment through higher moments and tail behavior.
In the linear IV analysis that follows, distributional relevance and the linear quantile representation for $X\mid Z$ guarantee the existence of an optimal Q--LS instrument that is relevant in the classical sense unless the projection degenerates. In the nonseparable triangular model in Appendix (ref), the condition that $Z$ is purely distributional is enough for the identification of the average structural function via a control function based on $F_{X\mid Z}(X\mid Z)$.
In what follows, distributional relevance will be our primitive assumption on the instrument. We show that, under a linear quantile representation for $X\mid Z$, it is enough to construct a strong instrument as a linear functional of the conditional quantile process, and to obtain consistent and asymptotically normal estimates of $(\alpha,\beta,\gamma)$ in the linear model (ref).
We now define the Quantile Least Squares IV (Q--LS) estimator. The key object is an optimal instrument constructed as a linear functional of the conditional quantile process of $X\mid Z$.
We work with the conditional quantile function of $X$ given $Z$. For each $\tau\in(0,1)$, assume
where $\pi_0(\tau)\in\mathbb{R}^{\dim(Z)}$ is a measurable function of $\tau$. We write \[ g(\tau,Z) := Q_{X\mid Z}(\tau\mid Z) = Z'\pi_0(\tau). \] Under mild regularity, distributional relevance is equivalent to the statement that the conditional quantile process $\{g(\tau,Z):\tau\in(0,1)\}$ is non-constant in $Z$ in $L^2$.\footnote{For concreteness, we implement Q--LS using linear quantile regressions of $X$ on $Z$ at a finite grid of quantile indices. This choice is purely for illustration: any alternative specification for the conditional quantiles (e.g., nonlinear, spline or series-based, or machine–learning first stages) could be used to construct the dictionary $G(Z)$ and hence the Q--LS instrument.}
We restrict attention to instruments that are linear functionals of this quantile process. Let $\omega(\cdot)$ be a square-integrable weight function on $(0,1)$. We define the quantile-aggregated instrument
The corresponding moment conditions for the structural parameter $\theta := (\alpha,\beta,\gamma')'$ are
Let \[ \mathcal{H} := \Big\{ h_\omega(\cdot): h_\omega(Z) = \int_0^1 \omega(\tau)\,g(\tau,Z)\,d\tau, \ \omega\in L^2(0,1) \Big\} \] be the linear span (in $L^2$) of the conditional quantile process.
Following the classical optimal IV logic for a single endogenous regressor, we consider the element of $\mathcal{H}$ that best predicts $X$ in mean square error.
By construction, $h_{\mathrm{opt}}(Z)$ is the $L^2$-projection of $X$ onto the closed linear span $\mathcal{H}$ of quantile-generated instruments. The next lemma shows that, whenever this projection is nonzero, the resulting instrument is automatically relevant for $X$ in the classical sense.
Lemma (ref) delivers a clean dichotomy: the optimal Q--LS instrument is either identically zero or, if nonzero, it is automatically relevant and its “strength” is measured by $\mathbb{E}[h_{\mathrm{opt}}(Z)^2]$. Distributional relevance guarantees that $\mathcal{H}$ contains more than just constants, but it does not by itself rule out the degenerate case $h_{\mathrm{opt}}\equiv 0$. A simple example is a pure variance-shift design, \[ X = \sigma(Z)\,U, \qquad \mathbb{E}[U]=0,\quad U \perp \!\!\! \perp Z, \] with $\sigma(Z)$ non-constant. In this case the conditional distributions $F_{X\mid Z}(\cdot\mid Z)$ vary with $Z$ (distributional relevance), yet $\operatorname{cov}(X,h(Z))=0$ for every $h\in\mathcal{H}$ so that $h_{\mathrm{opt}}(Z)\equiv 0$. Our Design B in Section (ref) provides a Monte Carlo illustration of this phenomenon.
For identification, we therefore add an explicit relevance condition that is the exact analogue of the usual mean-relevance condition in classical IV.
Assumption (ref) is the distributional analogue of the usual “strong IV” condition in linear models, where one assumes that the first-stage coefficient on the excluded instruments does not drift to zero as $n$ grows (e.g.\ staiger1997instrumental, stock2005asymptotic. Here we impose a lower bound on the variance of the optimal Q--LS instrument $h_{\mathrm{opt}}(Z)$, ruling out sequences of designs in which instruments become weak either in the mean or in the distribution. Throughout our asymptotic analysis we work under Assumption (ref), so our results should be interpreted as strong-distributional-IV asymptotics.
Under Assumption (ref), Lemma (ref) implies that $h_{\mathrm{opt}}$ is a non-degenerate, relevant instrument constructed purely from the conditional distribution of $X\mid Z$, even when $\mathbb{E}[X\mid Z]$ is constant in $Z$.
Throughout the linear IV analysis, we assume that distributional relevance holds and that $X\mid Z$ admits the linear quantile representation in (ref). Under these conditions, Lemma (ref) shows that the optimal Q--LS instrument $h_{\mathrm{opt}}(Z)$ is either identically zero or a classically relevant instrument for $X$, and Assumption (ref) rules out the degenerate case.
In practice, we approximate the integral in (ref) using a finite grid of quantiles. Let \[ 0 < \tau_1 < \cdots < \tau_K < 1 \] be a set of quantile indexes with gap of order $1/K$, and let $\Delta\tau_k$ denote the associated quadrature weights (for simplicity, one can take $\Delta\tau_k = 1/K$).
For each $\tau_k$, we estimate the first-stage quantile regression
and define the estimated quantile-based instruments \[ \widehat{g}_k(Z_i) := Z_i'\widehat{\pi}_n(\tau_k), \qquad i=1,\dots,n,\quad k=1,\dots,K. \] Stacking them, let \[ \widehat{G}_i := \big( \widehat{g}_1(Z_i), \dots, \widehat{g}_K(Z_i) \big)', \qquad \widehat{G} :=
\in \mathbb{R}^{n\times K}. \]
\paragraph{Discrete Q--LS weights.} A natural discrete analogue of (ref) chooses weights $w\in\mathbb{R}^K$ to best predict $X$ from the $K$ quantile instruments:
so that, when $\widehat{G}'\widehat{G}$ is invertible,
The corresponding finite-sample Q--LS first-stage prediction is
This $\widehat{X}_i^{\mathrm{Q\text{--}LS}}$ is the empirical best linear predictor of $X_i$ in the span of the $K$ quantile-based instruments.
\paragraph{Second stage.} Let $y = (Y_1,\dots,Y_n)'$ be the outcome vector, and define the regressor and instrument matrices \[ S = [\mathbf{1},\, X,\, Z_1], \qquad M = [\widehat{X}^{\mathrm{Q\text{--}LS}},\, Z_1], \] where $\widehat{X}^{\mathrm{Q\text{--}LS}} = (\widehat{X}_1^{\mathrm{Q\text{--}LS}}, \dots,\widehat{X}_n^{\mathrm{Q\text{--}LS}})'$ and $\mathbf{1}$ denotes the $n$-vector of ones. The Q--LS estimator of $\theta = (\alpha,\beta,\gamma')'$ is the usual 2SLS estimator
We now state conditions under which the Q--LS estimator is consistent and asymptotically normal. The key identification condition is distributional relevance, which, via Lemma (ref), implies relevance of the optimal Q--LS instrument.
For asymptotic normality, let \[ S_i := \big(1, X_i, Z_{1i}'\big)', \qquad M_i := \big(\widehat{X}_i^{\mathrm{Q\text{--}LS}}, Z_{1i}'\big)' \] denote the stacked regressor and instrument vectors used in the second stage, and let \[ \widehat{\varepsilon}_i := Y_i - S_i'\widehat{\theta}_n^{\mathrm{Q\text{--}LS}}. \] Define the sample matrices \[ \widehat{A}_n := \frac{1}{n}\sum_{i=1}^n S_i M_i', \qquad \widehat{Q}_n := \frac{1}{n}\sum_{i=1}^n M_i M_i', \] and the heteroskedasticity-robust “meat” matrix \[ \widehat{\Omega}_n := \frac{1}{n}\sum_{i=1}^n \widehat{\varepsilon}_i^{\,2}\, M_i M_i'. \]
Let \[ A := \mathbb{E}[S_i M_i^{0\prime}], \qquad Q := \mathbb{E}[M_i^0 M_i^{0\prime}], \qquad \Omega := \mathbb{E}\big[\varepsilon_i^{2} M_i^0 M_i^{0\prime}\big]. \] The population asymptotic covariance matrix of the Q--LS estimator can then be written in the usual IV “sandwich” form
In applications, we report standard errors as the square roots of the diagonal elements of $\widehat{V}_n^{\mathrm{Q\text{--}LS}}$ in (ref), which coincide with the usual heteroskedasticity-robust 2SLS standard errors computed using the generated Q--LS instrument and the exogenous regressors $Z_1$ as instruments.
Proposition (ref), is derived under Assumption (ref), which enforces a strong form of distributional relevance for the generated instrument $h(Z)$. As in the classical linear-IV literature, this rules out weak-IV sequences in which the covariance between the endogenous regressor and the instrument shrinks at a $\sqrt{n}$--local rate. In particular, our asymptotic normality result for Q--LS should be interpreted as a strong-identification limit theory: it guarantees standard Gaussian inference when the optimal distributional instrument $h_{\mathrm{opt}}(Z)$ has nonvanishing variance, but it does not characterize the behavior of Q--LS under general weak-IV sequences. Extending weak-identification robust inference to settings with distributional instruments would require combining our construction with the robust testing approaches of, for example, andrews2019weak or olea2013robust, which we leave for future work.
In this section we compare the quantile least squares (Q--LS) estimator to classical two-stage least squares (2SLS) in settings where both procedures are valid.
In the classical IV framework with a single endogenous regressor, relevance is formulated in terms of the conditional mean \[ m(Z) := \mathbb{E}[X\mid Z]. \] Under homoskedasticity, the “best” instrument for a single endogenous regressor is the efficient instrument $h^*(Z)$, which is proportional to the best $L^2$ predictor of $X$ given $Z$, i.e.\ to $m(Z)$. The optimal 2SLS estimator uses instruments in the linear span of $\{1,Z_1',m(Z)\}$ and is semiparametrically efficient for $\theta_0$ within this mean-based class.
By contrast, Q--LS is built from the entire conditional distribution of $X\mid Z$ via its conditional quantiles. Under (ref), we construct quantile-aggregated instruments of the form \[ h_\omega(Z) = \int_0^1 \omega(\tau)\,Q_{X\mid Z}(\tau\mid Z)\,d\tau, \] and define the optimal Q--LS instrument $h_{\mathrm{opt}}$ as in Definition (ref), i.e.\ as the best $L^2$ predictor of $X$ within the quantile-generated class $\mathcal{H}$. The Q--LS estimator $\widehat{\theta}_n^{\mathrm{Q\text{--}LS}}$ is the 2SLS estimator that uses $(h_{\mathrm{opt}}(Z),Z_1)$ as instruments.
The following proposition shows that, in a Gaussian/location setting with mean relevance, Q--LS and optimal 2SLS are asymptotically equivalent.
Proposition (ref) shows that, in a conventional Gaussian or location setting with mean relevance, Q--LS and optimal 2SLS are asymptotically equivalent: they use instruments that differ only by an irrelevant scale factor and attain the same efficiency bound.
Outside of the pure location case, Q--LS remains consistent under distributional relevance even when mean relevance fails or the conditional mean is misspecified, because it exploits shifts in the entire distribution of $X\mid Z$. In designs where both mean relevance and distributional relevance hold and the efficient IV $h^*(Z)$ lies in the closure of the quantile-generated class $\mathcal{H}$, Q--LS again attains the same efficiency bound as optimal 2SLS. When $h^*(Z)\notin\mathcal{H}$, Q--LS is efficient within $\mathcal{H}$ but may be slightly less efficient than an oracle 2SLS that can use arbitrary functions of $Z$ as instruments. The main gain of Q--LS is robustness: it can remain strongly identified in settings where the classical first-stage mean shift is close to zero but the distribution of $X\mid Z$ varies substantially with $Z$.
When the endogenous regressor is (or can be reduced to) a binary treatment $X=D\in\{0,1\}$, Q--LS also admits a familiar LATE interpretation under additional structure. In the positive-normalized version, the Q--LS index $H = h(Z)$ is a convex combination of conditional quantiles of $D\mid Z$, which collapses possibly high-dimensional instruments into a single scalar encouragement variable. Under standard IV validity and a no-defiers monotonicity condition with respect to $H$, the 2SLS coefficient using $H$ as the sole instrument can be written as a positive-weighted average of local average treatment effects for units whose treatment status changes as $H$ increases. Thus, in binary-treatment applications, Q--LS can be viewed not only as a strength/efficiency device but also as a way to construct a scalar monotone instrument that delivers a conventional LATE interpretation (see Appendix (ref) for details).
In the binary-treatment case, the scalar Q--LS index also admits a LATE representation with nonnegative weights on adjacent “index-complier” groups (Appendix (ref)), and these weights can be estimated in practice, so the researcher can diagnose which parts of the instrument support contribute most to the Q--LS effect. In addition, because the Q--LS index is built as a convex combination of conditional quantiles, the estimated weight vector \(\hat{w}\) reveals which parts of the \(D\mid Z\) distribution (e.g.\ central vs.\ tail quantiles) contribute most to instrument strength.
The population optimal Q--LS instrument $h_{\mathrm{opt}}$ solves the projection problem in Definition (ref). In the finite-grid implementation, this leads to the least-squares problem (ref) for the discrete weights $w\in\mathbb{R}^K$. When $K$ is moderate and the quantile-based instruments are highly collinear, the Gram matrix $\widehat{G}'\widehat{G}$ can be ill-conditioned, and the unregularized solution $\widehat{w}_n$ becomes unstable. This is the familiar ill-posedness of inverse problems and many-instrument IV.
We discuss two regularized versions of Q--LS that address this issue: ridge-regularized Q--LS and sparse (LASSO) Q--LS.
A simple and effective approach is Tikhonov (ridge) regularization. Instead of (ref), we solve
which yields the closed-form solution
The corresponding regularized Q--LS instrument is
Finite-sample stability is improved because the ridge term $n\lambda_n I_K$ stabilizes the inversion of $\widehat{G}'\widehat{G}$ and shrinks the weights toward zero. Under the same conditions as in Proposition (ref), and provided $\lambda_n\to 0$ at a suitable rate (for example, $\lambda_n\downarrow 0$ and $K_n^2/(n\lambda_n)\to 0$), the regularized weights $\widehat{w}_n^{\lambda}$ remain consistent for the population optimal weights and $\widehat{X}_i^{\mathrm{Q\text{--}LS},\lambda}$ converges in $L^2$ to $h_{\mathrm{opt}}(Z_i)$. The asymptotic distribution in Proposition (ref) therefore continues to hold with $\widehat{X}_i^{\mathrm{Q\text{--}LS},\lambda}$ replacing $\widehat{X}_i^{\mathrm{Q\text{--}LS}}$.
Two additional practical devices can be used in tandem with the ridge penalty. First, one may limit the number of quantile instruments $K$ to a moderate value (e.g.\ $K=5$ or $K=10$), or work with a low-dimensional principal-component representation of the columns of $\widehat{G}$. Second, instead of solving (ref) explicitly, one can treat each $\widehat{g}_k(Z_i)$ as a separate instrument, run 2SLS or LIML with the full instrument vector, and let the IV estimator choose an implicit linear combination of quantile-based instruments. These strategies reduce the effective instrument dimension and mitigate many-instrument bias while preserving the distributional information exploited by Q--LS.
The ill-posedness of the optimal weight problem arises because the finite-grid instrument dictionary $\widehat{G}$ can be high-dimensional and highly collinear. An alternative to ridge regularization is to enforce sparsity in the weights and let the data select a subset of quantile-based instruments. This leads naturally to an $\ell_1$-penalized (LASSO) version of Q--LS.
Given the $K$ quantile-based instruments $\widehat{g}_k(Z_i) = Z_i'\widehat{\pi}_n(\tau_k)$, stacked in $\widehat{G}_i$ and $\widehat{G}$ as in (ref), we define the LASSO-regularized weights as
where $\lambda_n>0$ is a tuning parameter. The corresponding sparse Q--LS instrument is
By construction, only a subset of the quantile indexes has $\widehat{w}_{n,k}^{\mathrm{LASSO}}\neq 0$, so the effective instrument dimension is reduced and the many-instrument problem is mitigated. In practice, we also consider a post-LASSO refinement: letting $\widehat{S}_n = \{k: \widehat{w}_{n,k}^{\mathrm{LASSO}}\neq 0\}$ denote the selected support, we re-estimate the weights by least squares restricted to $\widehat{S}_n$ and use the resulting fitted values as $\widehat{X}_i^{\mathrm{Q\text{--}LS,L}}$.
The second-stage estimator is then defined as in (ref), with $\widehat{X}^{\mathrm{Q\text{--}LS}}$ replaced by $\widehat{X}^{\mathrm{Q\text{--}LS,L}}$. From the perspective of the structural parameter $\theta_0$, the first stage (quantile processes and weights) is a high-dimensional nuisance. The IV moment condition is Neyman-orthogonal with respect to this nuisance, so that, under standard sparsity and rate conditions for LASSO and a suitable choice of $\lambda_n$, the estimation error in $\widehat{X}_i^{\mathrm{Q\text{--}LS,L}}$ is $o_p(n^{-1/2})$ in the moment equations. As a result, the asymptotic normality result in Proposition (ref) continues to hold with $\widehat{X}_i^{\mathrm{Q\text{--}LS,L}}$ in place of $\widehat{X}_i^{\mathrm{Q\text{--}LS}}$, and the usual heteroskedasticity-robust 2SLS covariance matrix computed with the sparse Q--LS instrument remains valid for inference.
In empirical work, the tuning parameter $\lambda_n$ can be chosen either by cross-validation (targeting prediction of $X$) or by more conservative theoretical rules designed for high-dimensional linear regression. A moderate level of penalization, combined with post-LASSO refitting, tends to yield stable first stages and good finite-sample size for Q--LS inference.
Our baseline Q--LS instrument is a linear functional of the conditional quantile process, \[ h_\omega(Z)=\int_0^1 \omega(\tau)\,Q_{X\mid Z}(\tau\mid Z)\,d\tau, \] with $\omega\in L^2(0,1)$. In finite samples, we approximate this object using a quantile grid $0<\tau_1<\cdots<\tau_K<1$ and fitted quantiles $\hat g_k(Z)\approx \widehat Q_{X\mid Z}(\tau_k\mid Z)$, so that $h(Z)\approx \hat G(Z)'w$ for $\hat G(Z)=(\hat g_1(Z),\ldots,\hat g_K(Z))'$.
For interpretability and stability, we also consider a constrained version that enforces positive, normalized weights:
Equivalently, in population form, this corresponds to restricting $\omega(\tau)\ge 0$ and $\int_0^1\omega(\tau)\,d\tau=1$.
\paragraph{Interpretation.} Under these constraints, $h_{\mathrm{PN}}(Z)=\hat G(Z)'w_{\mathrm{PN}}$ is an L-statistic: it is a convex combination of conditional quantiles, i.e.\ a “quantile average” index. Positivity also preserves distributional ordering: if the conditional distribution of $X$ under $Z=z_2$ first-order stochastically dominates that under $Z=z_1$ (so that $Q_{X\mid Z=z_2}(\tau)\ge Q_{X\mid Z=z_1}(\tau)$ for all $\tau$), then $h_{\mathrm{PN}}(z_2)\ge h_{\mathrm{PN}}(z_1)$ for any admissible weight function. This rules out sign-cancellation across quantiles and yields an index whose movements track distributional shifts in a single direction.
\paragraph{Finite-sample stability.} Because extreme quantiles are noisier to estimate, unconstrained Q--LS may assign large positive and negative weights that amplify estimation error. The simplex constraint $w\ge 0$, $\mathbf{1}'w=1$ regularizes the instrument, bounding its variability and often improving robustness when $\hat G$ is generated (e.g.\ via quantile regression). In our empirical implementation we compute $w_{\mathrm{PN}}$ by quadratic programming, and we can combine it with cross-fitting to further limit overfitting of $\hat G$.
We view Q--LS as a complement to, rather than a replacement for, the standard 2SLS toolkit with weak instruments. In practice, applied work can proceed in three steps: (i) run conventional mean-based diagnostics, (ii) assess distributional relevance, and (iii) choose between 2SLS and Q--LS (and its regularised variants) based on a small menu of test statistics and simple stability checks.
\paragraph{Step 1: Start from the usual 2SLS diagnostics.}
\paragraph{Step 2: Diagnose distributional relevance.}
\paragraph{Step 3: Choosing $K$, regularisation, and inference.}
\paragraph{Step 4: Stability checks and reporting.}
Table (ref) summarises these recommendations in a simple diagnosis chart.
This section will examine the finite-sample performance of Q--LS and its regularised variants in a set of Monte Carlo designs. We focus on three types of designs that map directly to our theoretical discussion. (A) strong mean relevance (benchmark Gaussian case), (B) Trigonometric and distributional instruments (variance shifts) and (C) mixed mean and distributional relevance.
The following parts of the data-generating processes (DGPs) are the same in all three designs;
We compare 10 estimators:
The weights for the weighted average of the quantiles are generated by sample splitting. The first 50% of the data is used to estimate the Ridge/LASSO regression regularisation parameters with 10 fold cross-validation predication accuracy. The estimated hyperparamter and the other 50% of the data is then used to estimated the Ridge/LASSO coefficient estimate.
For all Q–LS variants and DR-CF we consider quantiles between $v=0.01$ and $1-v$ and set the gap between the quantiles to $c=\{0.10,0.05,0.01,0.001\}$ this corresponds to using $K=\{10,20,99,981\}$ quantiles to generate the average of the quantile regression fitted values of $X$ or approximate the CDF ($\hat{v}$). The trimming of the quantiles follows cfnsv20_cfa.
Simulation results will be reported in terms of bias, root mean squared error (RMSE), and empirical 95% coverage of the simulation estimate.
The first design is a “best-case” benchmark in which classical mean relevance holds and the conditional distribution of $X\mid Z$ belongs to a Gaussian location family. The data-generating process of the first stage is
with $\alpha_1=\alpha_2=1$.
In this design, both 2SLS and Q--LS are correctly specified and the conditions of Proposition (ref) hold. The simulation will document:
The second design studies a case where the excluded instrument $Z_2$ generates rich non-linear and distributional variation in $X$, while the linear first stage of $X$ on $Z_2$ can be arbitrarily weak. The data-generating process of the first stage is
with $\alpha_1 = 1$ and a variance parameter $\gamma = 0.5$.
Hence, conditional on $(Z_1,Z_2)$, \[ \mathbb{E}[X \mid Z_1,Z_2] = \alpha_1 Z_1 - \cos(\omega Z_2), \qquad \operatorname{Var}(X \mid Z_1,Z_2) = \exp(2\gamma Z_2). \] The excluded instrument $Z_2$ affects the conditional mean through the nonlinear term $\cos(\omega Z_2)$ and the conditional dispersion through $\exp(\gamma Z_2)$. The conditional distribution $F_{X\mid Z}(\cdot \mid Z_1,Z_2)$ is therefore strongly distributionally relevant in the sense of the main text.
Note that, since $Z_2 \sim \mathcal{N}(0,1)$ is symmetric around zero and $z \mapsto \cos(\omega z)$ is an even function, we have \[ Cov\big(Z_2, \cos(\omega Z_2)\big) = \mathbb{E}\big[-Z_2 \cos(\omega Z_2)\big] = 0 \] for all $\omega \ge 0$. Moreover, $\nu$ is independent of $Z_2$ and has mean zero, so $Cov(Z_2,\exp(\gamma Z_2)\nu)=0$. Thus \[ Cov(X,Z_2) = 0 \] in the population for every value of $\omega$. Classical linear IV based on the regression of $X$ on $Z_2$ therefore has an exactly zero linear first stage, even though $Z_2$ induces strong nonlinear and distributional variation in $X$.
\paragraph{Sub-designs B1--B3.} We consider three values of the frequency parameter $\omega$:
Overall, Designs B1--B3 trace a progression from a purely distributional instrument (B1, no mean effect of $Z_2$) through an “almost mean-irrelevant” case (B2, very flat mean effect) to a design with strong nonlinear mean and variance effects (B3). This family highlights that classical first-stage diagnostics based on the linear regression of $X$ on $Z_2$ can substantially understate instrument strength, whereas Q--LS remains well identified whenever $X$ has a nonzero projection in the quantile-generated space.
The third design introduces both mean and variance effects:
Here, both mean relevance and distributional relevance hold, but the contribution of each can be tuned by varying $(\alpha_2,\sigma(0),\sigma(1))$. The simulation will:
Note as $Z_2$ is binary in this design we set $\zeta_{\tau_{k,4}}=0$ in (ref) to avoid a singular QR design matrix, for all Q--LS variants and DR-CF.
Table (ref) presents the results for Design A and Design C for sample sizes of 500 and 1000.\footnote{Full results for Design A and Design C are available upon request.} It is important to note that classical IV is correctly specified, as $Z_2$ is strongly mean-relevant in both Design A and Design C. For Design A, we find that all Q-LS variants perform similarly to classical IV, regardless of the number of quantiles used. This is due to the quantile projection being similar to the mean projection. The analytical Q-LS variants provide better coverage (closer to 0.95) than the OLS plug-in standard errors, which tend to yield over-coverage, indicating conservative standard errors in this design. The control function version (DR-CF) performs worse in terms of bias, RMSE, and coverage than all Q-LS variants and classical IV, regardless of sample size. However, its performance does improve as the number of quantiles increases in this design.
For Design C, $Z_2$ is both strongly mean-relevant and variance-relevant. As mean relevance is less clean, the bias and RMSE of classical IV are ten and two times as large, respectively, as in Design A for a given sample size. Unlike Design A, the unregularized Q-LS variants exhibit substantial bias when 10 quantiles are used; however, this bias decreases as the number of quantiles and the sample size increase. When 981 quantiles are used, all Q-LS variants have the same or smaller bias than classical IV, and the regularized versions (plug-in and analytical) all yield good coverage. The performance of DR-CF relative to the other estimators and as the number of quantiles and the sample size increase, is similar to that in Design A.
Table (ref) shows the results for Design B with $\omega={0,0.5}$.\footnote{Full results for Design B are available upon request.} Interestingly, classical IV performs worse when $\omega$ is non-zero in terms of both bias and RMSE for a given sample size, whereas the bias and RMSE of the Q-LS variants decrease, except for the bias of Q–LS-L1, which marginally increases (in absolute value) as the number of quantiles increases. More quantiles are needed for better approximation because the distributional relevance of $Z_2$ becomes stronger as $\omega$ increases due to the $\cos$ transformation. Similar to Designs A and C, DR-CF performance improves as the number of quantiles increases. However, unlike Designs A and C, DR-CF has the smallest bias and RMSE when $\omega=0$ for a given sample size, and also the smallest bias and RMSE when $\omega=0.5$ and $n=1000$. Despite this, DR-CF coverage is the worst among all estimators considered for a given number of quantiles and sample size.
Overall, across all designs considered, we find that at least one variant of Q-LS with 99 or more quantiles performs as well as or better than classical IV, even in designs where the instrument $Z_2$ is strongly mean-relevant. When a large number of quantiles are used, analytical Q-LS empirical coverage tends to be close to, but slightly below, 0.95, while the plug-in variants tend to exhibit coverage above 0.95. When the instrument is not strongly mean-relevant, the plug-in point estimates tend to be more stable than their analytical counterparts. The DR-CF estimator performs better as the number of quantiles increases across all designs considered, but it requires a large number of quantiles and weak mean relevance to achieve lower bias and RMSE than all Q-LS variants.
In Design A and C (strong mean relevance), Q--LS essentially reproduces optimal 2SLS. This confirms the theoretical prediction that, in a pure location-shift setting with strong instruments, exploiting the full conditional quantile process does not alter first-order behavior. Design B1 illustrates the degenerate branch of Lemma (ref). Because the optimal Q--LS instrument $h_{\mathrm{opt}}(Z)$ is identically zero, the first stage for Q--LS collapses and the estimator fails in the usual weak-IV sense. We do not regard this as a pathology of Q--LS, but as a reminder that distributional relevance alone is not sufficient for a useful first stage; Assumption (ref) is needed to rule out such cases. The main action comes in Design B3, where instruments are weak in the mean but strong in the distribution. Here classical 2SLS exhibits substantial finite-sample distortions. In contrast, Q--LS estimators based on a moderate grid of quantiles (e.g.\ $K=10$) are well centered and display near-correct coverage across the same sample sizes. These patterns mirror the HRS application: when policy reforms primarily reshape tails and dispersion rather than means, exploiting distributional relevance through Q--LS restores a strong, stable first stage from the same underlying instrument.
We use the Health and Retirement Study (HRS) to study how financial exposure to out-of-pocket (OOP) medical spending risk affects mental health among older Americans. Our main outcome is an indicator for any depressive symptoms (CESD) among respondents aged 65 and older. The endogenous regressor is a scalar measure of annual OOP medical spending, and the key policy instrument is an indicator for the post--Medicare Part D period (year $\ge 2006$), during which subsidized prescription drug coverage expanded and catastrophic protection was introduced.
To mirror the simulation designs, we work with two complementary specifications. First, we treat OOP spending in nominal dollars and use a single post--Part D indicator as the excluded instrument. This “stress test” design generates a classical weak-IV environment when the risk index is summarized by mean OOP spending, and it provides a direct counterpart to the distributional simulation in Section (ref). Second, we express OOP spending in real 2015 dollars using the OECD CPI and show that the same policy indicator becomes a strong instrument in this scaling. The real-dollar specification serves as a well-identified benchmark in which Q–LS and conventional 2SLS can be expected to coincide when we omit a time trend. Once we add a linear trend in survey year, the point estimates remain very similar but Q–LS delivers somewhat tighter standard errors. Throughout, we trim the top 1% of the OOP distribution to reduce the influence of extreme outliers.
\paragraph{Descriptive patterns.} Figure (ref) plots average out-of-pocket (OOP) medical spending in thousands of dollars (left axis) and the share of respondents with any depressive symptoms (right axis) among HRS respondents aged 65 and older from 2000 to 2014. Mean OOP spending rises from about \$3{,}000 in 2000 to a peak of roughly \$5{,}200 in the mid-2000s, then falls back toward \$4{,}000–\$4{,}300 in the late 2000s and early 2010s. Over the same period, the prevalence of depression declines fairly steadily: the share with any depression falls from around 19–20 percent in 2000 to roughly 11–12 percent by the early 2010s.
These raw trends are consistent with the idea that older households faced substantial OOP risk in the early 2000s, followed by some compression of spending risk---coinciding with the implementation of Medicare Part D in 2006---while mental health outcomes improved over time. However, the figure is purely descriptive: it does not control for changes in the composition of the 65+ population, secular improvements in treatment, or other contemporaneous policy changes. In the regression analysis below, we therefore focus on isolating the effect of changes in financial risk exposure, as measured by OOP spending, using instrumental‐variables methods. In addition, we trim the regression sample at the 99th percentile of realized annual OOP spending to remove a small number of extreme outliers; this has negligible impact on the aggregate time-series patterns in Figure (ref) but stabilizes the covariate–outcome relationships.
Figure (ref) plots the mean, median, and upper quantiles (75th, 90th, and 95th percentiles) of annual OOP medical spending for HRS respondents aged 65 and older between 2000 and 2020. The median and 75th percentile are remarkably stable over time, indicating that typical OOP spending changes little over this period. By contrast, the mean closely tracks movements in the upper tail: the 90th and 95th percentiles rise sharply in the early 2000s and then flatten or decline following the introduction of Medicare Part D in 2006. The resulting compression of the gap between the upper quantiles and the median after 2006 is consistent with Part D primarily reducing exposure to extreme spending realizations rather than shifting the center of the distribution. From the perspective of our identification strategy, this pattern illustrates why standard first-stage diagnostics based on mean shifts label the post–Part D indicator as a weak instrument for mean OOP risk, even though it is strongly relevant for the distribution of spending, especially in the right tail.
We begin with a deliberately “stressful” specification that uses nominal OOP medical spending (in thousands of dollars) as the endogenous regressor and a single post–Part D indicator as the excluded instrument. This design closely mirrors the weak-IV, heavy-tailed environments in our Monte Carlo exercises and is useful for showing how Q--LS behaves when the instrument is weak for the mean but strongly relevant for the distribution of OOP spending.
Table (ref) reports linear probability models relating an indicator for any depressive symptoms to mean OOP medical spending. The sample consists of 16{,}313 person–year observations for respondents aged 65 and older in the RAND HRS panel between 2000 and 2020, after excluding person–years with OOP spending above the 99th percentile of the nominal OOP distribution.
In column (1), higher nominal OOP spending is positively and statistically significantly associated with depression: a \$1{,}000 increase in annual OOP spending is associated with a 0.2 percentage point increase in the probability of reporting depressive symptoms ($\hat\beta = 0.002$, s.e.\ 0.001). The 2SLS specification in column (2), which instruments mean OOP spending with post–Part D, yields a much larger point estimate ($\hat\beta = 0.442$) but an even larger standard error (s.e.\ 0.410). The IV estimate is therefore imprecise and not statistically different from zero at conventional levels.
Table (ref) shows that this lack of precision reflects the fact that the post–Part D indicator is a very weak instrument for mean nominal OOP spending in the trimmed sample: the cluster-robust partial $F$-statistic for the excluded instrument is 1.19, far below usual thresholds.
\paragraph{Q--LS vs 2SLS in the weak nominal design.} We now use this weak-IV nominal specification as a stress test to compare Q--LS and linear 2SLS. Our goal here is diagnostic: to document how the two estimators behave when the instrument is weak for the mean, yet strongly relevant for the distribution of OOP spending.
Table (ref) compares the Q--LS estimate based on the projected OOP risk index with the conventional linear 2SLS estimate that uses mean nominal OOP spending as the endogenous regressor.
The point estimates in columns (1) and (2) are numerically identical (0.442), but the standard errors differ dramatically. In the weak-IV nominal specification, Q--LS delivers a standard error of 0.071 and a precisely estimated positive effect, whereas linear 2SLS yields a standard error of 0.410 and a statistically insignificant estimate. This pattern mirrors our simulation results: when the policy instrument reshapes the distribution of OOP spending but produces only a noisy mean shift, Q--LS can extract a much stronger effective first stage from the same variation, tightening confidence intervals without changing the underlying signal.
In this nominal specification, the combination of weak mean relevance and heavy tails also makes the estimated magnitude difficult to interpret substantively. In the next subsection, we therefore turn to a more natural specification that expresses OOP spending in real terms.
For substantive interpretation, it is more natural to express OOP spending in real dollars. In the second part of the empirical section we therefore: (i) redefine OOP spending in thousands of 2015 U.S.\ dollars, and (ii) show how the Part D instrument looks first like a strong mean instrument, and then, after soaking up secular trends, like a primarily distributional instrument where Q--LS mainly improves precision.
We deflate nominal OOP using the OECD consumer price index obtained from FRED, constructing a monthly index with base year 2015 and aggregating to the relevant survey years. Real OOP spending in thousands of 2015 dollars is denoted $oop\_med\_2015\_k$.\footnote{Organization for Economic Co-operation and Development, Consumer Price Indices (CPIs, HICPs), COICOP 1999: Consumer Price Index: Total for United States, retrieved from FRED, Federal Reserve Bank of St. Louis; \url{https://fred.stlouisfed.org/series/USACPIALLMINMEI}, accessed January 14, 2026.}
\paragraph{Real OOP without time trend: a strong-IV benchmark.} In the first real-dollar specification, we repeat the nominal regressions replacing $oop\_med\_k$ with $oop\_med\_2015\_k$ and keeping the same controls. In this case, the post–Part D indicator is a strong instrument for real OOP: the first-stage partial $F$-statistic is around 50 in the trimmed sample. Table (ref) compares Q--LS and linear 2SLS for the depression outcome in this strong-IV environment.
In this specification, Q--LS and 2SLS are virtually indistinguishable: both yield an effect of about 0.045 per \$1{,}000 of real OOP, with very similar standard errors. This is exactly the “well behaved” case in which our theory predicts that Q--LS collapses to optimal 2SLS when the instrument is strong and the linear mean representation is adequate. Table (ref) , in Appendix, confirms that three different ways of evaluating uncertainty for the Q--LS coefficient---cluster-robust OLS, clustered bootstrap, and the analytical Q--LS sandwich formula---deliver nearly identical standard errors. \paragraph{Real OOP with time trend: a distributional relevance design.} Finally, we add a flexible linear time trend to the real-OOP specification. This absorbs much of the secular drift in spending and returns us to a situation in which post--Part D induces rich distributional changes in real OOP but only a small additional mean shift conditional on time trend: with first-stage partial $F$-statistic is around 2. From the perspective of classical diagnostics, the instrument again looks fragile; from the perspective of distributional relevance, it remains informative.
Table (ref) reports Q--LS and 2SLS estimates for this trend-adjusted real specification.
In this trend-adjusted real specification, Q--LS and 2SLS again deliver the same point estimate (0.07 per \$1{,}000 of real OOP). The main difference is precision: the Q--LS standard error is 0.05, whereas the 2SLS standard error is 0.07. Thus Q--LS reduces the confidence interval width by roughly one-third, but the estimate remains only marginally significant at best (p-value around 0.14). Economically, the estimates across the real-dollar specifications suggest a modest positive effect of OOP risk on depression, with the strongest evidence in the specification that treats Part D as a strong mean instrument and somewhat weaker evidence once secular trends are absorbed.
Taken together, the three designs -- nominal weak-IV, real strong-IV, and real--with-trend -illustrate the two key empirical messages of the paper. First, the way in which a policy reform is summarized in the first stage (nominal vs real, with or without trend) can dramatically change whether the design looks weak or strong in mean terms, even when the underlying distributional shifts are similar. Second, when mean relevance is fragile but distributional relevance can be strong, Q--LS behaves like a natural complement to 2SLS: it recovers the same underlying signal where the design is well identified, and it delivers substantial precision gains relative to linear IV in settings where standard weak-IV diagnostics based on mean shifts understate the information content of the instrument.
Our results can be summarized along two dimensions. On the identification side, we formalize a notion of distributional relevance and show that even purely distributional instruments---for which $\operatorname{Var}(\mathbb{E}[X\mid Z]) = 0$ while $F_{X\mid Z}(\cdot\mid Z)$ varies nontrivially in $Z$---are sufficient to identify average structural functions in a nonseparable triangular model via the control function $V = F_{X\mid Z}(X\mid Z)$. This makes explicit how instruments that look weak in the mean can still be strong in the distribution, and connects the control–function identification argument to weak-IV concerns in linear models.
On the estimation side, we propose a simple, implementable estimator, Quantile Least Squares (Q--LS), which constructs an optimal distributional instrument from conditional quantiles and then uses it in a conventional linear IV step. In designs where policy reforms are primarily meant to reshape the distribution of risk rather than to move its mean, Q--LS complements the existing 2SLS and weak-IV toolkits by extracting a stronger effective first stage from the same policy variation, while leaving classical strong-IV applications essentially unchanged.
Q--LS exploits this structure by choosing, within a given dictionary of quantile-based functions, the instrument that best predicts $X$ in mean square error and then using it in a standard 2SLS procedure. We showed that Q--LS is consistent and asymptotically normal under distributional relevance, that it coincides with optimal 2SLS in Gaussian/location settings where mean relevance holds, and that it remains reliable when classical first-stage diagnostics indicate weak instruments from the perspective of the conditional mean. We also developed ridge and LASSO regularisation schemes that stabilise the first stage and mitigate many-instrument bias, while preserving the first-order asymptotics of Q--LS and providing practically useful variants (Q–LS--R, Q–LS--L1, Q–LS--L2) with distinct robustness properties.
Our Monte Carlo evidence highlights three main messages. First, when the instrument primarily shifts the tails of the $X$-distribution, Q--LS can recover the same limiting parameter as optimal 2SLS but with substantially smaller finite-sample variance, even in designs where the mean first stage is formally weak. Second, the regularised Q--LS variants are effective at controlling the dispersion of the estimator in many-instrument designs and under moderate misspecification of the quantile representation. Third, conventional IV estimators can behave erratically in these designs, with large RMSE and unstable confidence intervals, even when coverage remains nominal, whereas Q--LS delivers more concentrated sampling distributions.
The empirical application to Medicare Part D illustrates these features in a setting where policy reforms plausibly operate through the distribution of medical spending risk rather than its mean. In the nominal-dollar specifications, standard first-stage diagnostics classify the post--Part D indicator as a very weak instrument for mean OOP spending, and conventional 2SLS delivers large but imprecise estimates of the effect of OOP risk on depression. By contrast, Q--LS exploits the pronounced compression of the upper tail of the OOP distribution to construct a strong distributional first stage from the same policy variation, yielding substantially tighter confidence intervals while preserving the economic interpretation of the effect. When we deflate spending to 2015 dollars, the mean first stage becomes much stronger and conventional 2SLS performs well; in that case Q--LS and 2SLS again agree closely on the point estimates, but Q--LS offers at most modest efficiency gains. Taken together, the simulations and the Medicare Part D evidence suggest viewing Q--LS as a complement to standard IV: it nests the optimal mean-based instrument in classical designs, but remains informative in applications where instruments are weak in the mean and strong in the distribution.
The distributional relevance framework extends beyond linear IV. In an appendix, we show that purely distributional instruments can identify average structural functions in nonseparable triangular models via a control-function representation based on the conditional distribution $F_X(X\mid Z)$. This connects our approach to the literature on nonparametric IV and completeness, and suggests several directions for future work: deriving more primitive conditions that guarantee nondegenerate optimal instruments; developing formal diagnostics for the linear quantile representation and for instrument-instability; designing data-driven procedures for selecting quantile grids and penalties; and applying Q--LS in empirical settings where policy or institutional variation plausibly operates through changes in risk exposure rather than simple mean shifts.