EconBase
← Back to paper

Distributional Instruments: Identification and Estimation with Quantile Least Squares

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

129,665 characters · 32 sections · 52 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Distributional Instruments: Identification and Estimation with Quantile Least Squares

abstractWe study instrumental-variable designs where policy reforms strongly shift the distribution of an endogenous variable but only weakly move its mean. We formalize this by introducing distributional relevance: instruments may be “purely distributional" with $\operatorname{Var}(\mathbb{E}[X\mid Z])=0$ while $F_{X\mid Z}(\cdot\mid Z)$ varies nontrivially. Within a triangular model, distributional relevance suffices for nonparametric identification of average structural effects via a control function constructed from $F_{X\mid Z}$. We then propose Quantile Least Squares (Q–LS), which aggregates conditional quantiles of $X$ given $Z$ into an optimal mean-square predictor and uses this projection as an instrument in a linear IV estimator. We establish consistency, asymptotic normality, and the validity of standard 2SLS variance formulas, and we discuss regularization across quantiles. Monte Carlo designs show that Q–LS delivers well-centered estimates and near-correct size when mean-based 2SLS suffers from weak instruments. In Health and Retirement Study data, Q–LS exploits Medicare Part D–induced distributional shifts in out-of-pocket risk to sharpen estimates of its effects on depression.\\ \noindentKeywords: Distributional instruments; Quantile least squares; Instrumental variables; Medicare Part D; Out-of-pocket medical spending risk; Depression. \noindentJEL codes: C26; C14; C36; I13; D14.

Introduction

A central promise of public health insurance is insurance, not just subsidies: it should protect households against the risk of large, unpredictable medical expenses. In the United States this promise is especially salient for older adults. Nearly all Americans over age 65 are enrolled in Medicare, yet many still face substantial exposure to out-of-pocket (OOP) medical spending, particularly for prescription drugs.\footnote{See, for example, finkelstein2008did on the initial impact of Medicare on OOP spending and jones2018lifetime on lifetime medical spending of retirees.} The introduction of Medicare Part D in 2006 was explicitly designed to reduce this financial risk by expanding subsidized prescription drug coverage and introducing catastrophic protection. A large empirical literature has since documented how Part D and related coverage changes affect drug utilization, adherence, and OOP spending ketcham2008medicare,engelhardt2011medicare,park2017medicare,carvalho2019impact, and a parallel literature shows that Medicare eligibility and supplemental coverage sharply reduce the probability of catastrophic medical expenditures and financial strain among the elderly barcellos2015effects,scott2021assessing,jones2019predicting.

Much of this work treats financial exposure itself---for example, the level, variance, or tail probability of OOP spending---as an outcome of policy. From a policy design perspective, however, we often want the reverse object: the causal effect of financial risk exposure on behavior and welfare. Does facing a higher risk of catastrophic prescription drug expenses change adherence, portfolio choice, or retirement decisions? Does improved financial risk protection from Part D or supplemental policies translate into better mental health or reduced financial distress?\footnote{See ayyagari2015does and ayyagari2016medicare for related evidence on mental health and portfolio choice.} Answering such questions requires treating a measure of risk exposure---for example, a catastrophic-spending indicator or a predicted OOP risk index---as an endogenous regressor and using policy variation as an instrument.

In practice, this approach runs into a familiar econometric obstacle. The most natural instruments for a risk index are policy indicators (e.g.\ Part D eligibility or plan generosity) and features of the plan menu. These instruments primarily shift the distribution of OOP spending---compressing right tails, changing dispersion, and altering the frequency of catastrophic events---while often having relatively modest effects on the mean of the scalar risk measure itself. Empirically, first-stage regressions of a single risk index on these policy variables can have very low $F$-statistics, even when the underlying policy clearly reshapes the full distribution of spending engelhardt2011medicare,barcellos2015effects,scott2021assessing. \footnote{ Table (ref) summarizes some representative studies showing mean effect of insurance on mean OOP spending. }Standard IV diagnostics based on mean shifts therefore label the design as “weak” and push researchers toward reduced-form analyses, regression discontinuity designs at age 65, or structural models of plan choice abaluck2011choice, rather than causal IV estimates for the effect of financial risk on outcomes.

This paper starts from the observation that, in such settings, the instruments are not weak in any substantive sense: they are strongly relevant for the distribution of the endogenous variable, even if they appear weak for its mean. We formalize this idea by replacing the classical notion of mean relevance---variation in $\mathbb{E}[X\mid Z]$---with a weaker and more general notion of distributional relevance: the conditional distribution $F_{X\mid Z}(\cdot\mid Z)$ of the endogenous variable $X$ shifts with the instrument $Z$ in an $L^2$ sense, even when $\mathbb{E}[X\mid Z]$ is constant. In the Medicare example, $X$ may be a scalar index of financial exposure constructed from the OOP distribution, while $Z$ captures policy-induced variation in coverage generosity. Part D and related reforms satisfy distributional relevance whenever they alter the shape (tails, variance, skewness) of the spending distribution, regardless of whether they strongly move its mean.

Our first contribution is to show that distributional relevance is a meaningful notion of instrument strength, not just a descriptive property. Within a non-separable triangular model, we characterize a class of purely distributional instruments: variables $Z$ for which $\operatorname{Var}(\mathbb{E}[X\mid Z])=0$ but $F_{X\mid Z}(\cdot\mid Z)$ is non-degenerate in $L^2$. Building on the control-function framework of cfnsv20_cfa, we prove that such instruments still identify average structural effects via the control variable $V=F_X(X\mid Z)$, and we relate this notion to existing concepts of instrument strength based on heteroskedasticity lewbel1997constructing. Thus, even when policy variables are “purely distributional” for a risk index, they can support point identification of its causal effect on outcomes.

Our second, and main, contribution is to develop a practical IV estimator tailored to these designs, which we call Quantile Least Squares (Q--LS).\footnote{Our notion of “distributional instruments" is orthogonal to the shift-share (Bartik) literature (see adao2019shift, borusyak2025practical, goldsmith2020bartik for a detailed review). There, the focus is on constructing a composite instrument from many shocks and exposure shares. Here, we take the instrument $Z$ as given and show how to optimally exploit its distributional impact on $X$ via conditional quantiles. In particular, Q–LS does not impose a shift–share structure on $Z$; it can be applied whether $Z$ is a simple policy indicator or a more complex shock.} Instead of summarizing the first stage with a conditional mean, Q--LS builds instruments from the entire conditional quantile process of $X$ given $Z$. Concretely, we: (i) estimate a series of first-stage quantile regressions of $X$ on $Z$; (ii) form a data-driven linear combination of these conditional quantiles that best predicts $X$ in mean square error; and (iii) use this optimal quantile-aggregated prediction as the instrument in an otherwise standard linear IV estimator. Under distributional relevance and a mild non-degeneracy condition, the resulting optimal Q--LS instrument is always relevant in the classical sense, even when the conditional mean of $X$ is flat in $Z$. In a “best-case” world where $X\mid Z$ is a pure location shift, we show that Q--LS collapses to optimal 2SLS and attains the same efficiency bound. In more general cases, Q--LS remains reliable precisely when mean-based instruments become weak.

Methodologically, we characterize the asymptotic properties of Q--LS in a linear triangular model with a single endogenous regressor. We show that (i) under distributional relevance and mild regularity conditions, the Q--LS estimator is consistent and asymptotically normal; (ii) standard heteroskedasticity-robust 2SLS standard errors computed with the generated Q--LS instrument are valid for inference; and (iii) ill-posedness in the optimal-weight problem can be handled with ridge or LASSO regularization across quantiles, yielding stable finite-sample performance. We propose simple diagnostics and a decision chart that place Q--LS alongside conventional weak-IV tools: researchers can test for mean relevance, test for distributional relevance, and then choose between classical 2SLS and Q--LS (and between its regularized variants) based on these diagnostics. In this sense, Q--LS is designed to complement, not replace, the existing 2SLS toolkit.

Our Monte Carlo experiments are calibrated to the Medicare context and compare Q--LS to conventional 2SLS and a control-function estimator in designs where instruments primarily shift the distribution of risk rather than its mean. In designs with strong mean relevance, Q--LS and 2SLS deliver nearly identical point estimates and standard errors, confirming that Q--LS does not sacrifice efficiency when classical IV works well. In designs where traditional first-stage $F$-statistics indicate severe weak-instrument problems, Q--LS produces well-centered estimates and near-correct coverage, while 2SLS exhibits substantial bias and undercoverage. Regularized Q--LS variants further stabilise performance when the quantile dictionary is rich relative to sample size.

Finally, we illustrate the usefulness of distributional instruments in the Medicare Part D setting. Using the Health and Retirement Study (HRS), we construct individual-level measures of OOP medical spending risk in real (2015) dollars and exploit Part D as a source of exogenous variation in risk exposure. We show, first, that standard scalar risk measures, such as mean OOP spending or catastrophic-spending indicators, can render an otherwise informative policy design “weak” by classical first-stage diagnostics, even when the upper tail of the OOP distribution clearly shifts. Second, we demonstrate how Q--LS, built from a finite dictionary of conditional quantiles of OOP spending, recovers the same causal effects as linear 2SLS in specifications where the real OOP first stage is strong, while delivering much tighter confidence intervals and more stable estimates in specifications where mean relevance is weak. Substantively, our estimates indicate that higher OOP risk is associated with worse mental health among older Americans. To our knowledge, this is the first empirical illustration of how distributional instruments can materially improve the precision and credibility of IV estimates in applied work on medical spending risk.

\paragraph{Contributions.} Our findings contribute to two main parts of the instrumental‐variables literature. First, on the identification side, building on the control–function results of cfnsv20_cfa, we highlight that their argument continues to apply even when instruments are purely distributional: even when $\operatorname{Var}(\mathbb{E}[X\mid Z]) = 0$ and the conditional mean of the endogenous variable is completely flat in the instrument, nontrivial variation in the conditional distribution $F_{X\mid Z}(\cdot\mid Z)$ is sufficient to identify average structural effects in a nonseparable triangular model via a control function based on $F_{X\mid Z}(X\mid Z)$. Our contribution is to make this “purely distributional’’ case explicit and to connect it to weak–IV diagnostics in linear models, where instruments may look weak in the mean yet be strong in the distribution. This formalizes a notion of distributional relevance that extends classical mean relevance in a direction that is natural for risk and tail–outcome applications.

Second, on the linear IV side, we develop a simple, implementable estimator, Quantile Least Squares (Q--LS), which constructs an optimal distributional instrument from conditional quantiles and then uses it in a conventional linear IV step. Under a strong-distributional-relevance condition, Q--LS shares the same large-sample behavior as 2SLS when instruments are strong in the mean, but in designs where policy reforms mainly reshape the distribution of risk it can extract a much stronger effective first stage from the same variation, complementing---rather than replacing---existing weak-IV diagnostics and robust inference tools. We characterize its finite–sample behavior through Monte Carlo designs.

Our empirical application illustrates how Q--LS can extract a strong first stage from policy variation that looks weak in classical mean-based diagnostics, while collapsing back to conventional IV when the instrument is strong in the mean.

The remainder of the paper is organized as follows. Section (ref) discusses weak identification, quantile regression, and some existing ways to circumvent weak identification. Section (ref) introduces the linear IV setup and formalizes distributional relevance. Section (ref) defines the Q--LS estimator, presents its asymptotic properties, and compares Q--LS to classical 2SLS and shows that they coincide in a Gaussian/location setting. Section (ref) discusses regularization, diagnostics, and practical implementation. Section (ref) outlines the simulation designs and results. Section (ref) presents the Medicare Part D application. Section (ref) concludes.

Literature Review

The aim of this research is to develop a general framework that circumvents the classical notion of weak identification of average structural functions in an instrumental variables design. In this section, we first describe and characterize weak identification in triangular models. Second, we summarize quantile regression and discuss how it has been used to address endogeneity in structural functions. Finally, we discuss how our proposed method relates to the existing literature that imposes strong functional form restrictions on the underlying model to circumvent the problem of weak identification.

Weak identification in triangular and linear IV models

Let $(Y,X,Z)$ be random variables that follow a triangular system \[ X = g(Z,\eta), \qquad Y = f(X,\varepsilon), \qquad (\eta,\varepsilon)\perp Z, \] where $f$ and $g$ are unknown measurable structural functions, and $\varepsilon,\eta$ are unobserved disturbances. The vector $Z$ collects the instruments. The average structural function (ASF) is \[ \mu(x) := \mathbb{E}\big[f(x,\varepsilon)\big], \] so identification of $\mu(\cdot)$ hinges on how variation in $Z$ propagates to variation in $X$ through $g(\cdot)$ and, ultimately, to variation in $Y$.

In nonseparable triangular models, identification of $\mu(\cdot)$ often proceeds via control-function or distribution-regression arguments that construct a control variable $V=V(X,Z)$ (for example $V=F_{X\mid Z}(X\mid Z)$) such that $Y \perp Z \mid (X,V)$ and suitable completeness conditions hold; see, among others, newey2003instrumental,newey2021control, florens2008identification, imbens2009identification, cfnsv20_cfa, and dunker2023nonparametric. In this framework, weak identification arises when the mapping from $Z$ to the distribution of $X$ is too “flat’’ to generate informative variation in the control function. Intuitively, if changes in $Z$ induce only very small or nearly collinear shifts in $F_{X\mid Z}(\cdot\mid z)$, then the associated conditional moment conditions are nearly redundant and the inverse problem that recovers $\mu(\cdot)$ becomes ill–posed or nearly unidentified. This is the nonlinear analogue of weak instruments: the conditional distribution of $X$ given $Z$ is only weakly sensitive to $Z$, so the control function has little effective variation.

Closest in spirit to our work is the recent “Distributional Instrumental Variable (DIV)” method of holovchak2025div. They also exploit distributional shifts in $X\mid Z$ and show how to identify and estimate the entire interventional distribution of the outcome using flexible generative models in a nonlinear IV setting. DIV provides conditions under which the interventional distribution is identified, and illustrates designs where 2SLS fails but distributional information suffices for identification and accurate estimation of mean and quantile treatment effects. Relative to this approach, we focus on average structural effects in a triangular model and on the construction of an optimal scalar instrument for linear IV estimands. Our Q--LS procedure can be viewed as a low-dimensional, analytically tractable counterpart to DIV: it concentrates the distributional information in a finite-dimensional quantile dictionary and delivers an optimal $L^2$-projection that can be used within the classical IV toolkit. Moreover, when the endogenous regressor is (or can be reduced to) a binary treatment, the Q--LS index provides a scalar monotone instrument, so that the resulting 2SLS coefficient admits a standard LATE interpretation as a positive-weighted average of complier effects across index increments.

To connect with the more familiar linear IV setting, consider the linear special case \[ Y = X\beta + W'\gamma + u, \qquad X = Z'\pi + W'\delta + v, \] where $Z$ are excluded instruments, $W$ are included exogenous regressors, and $\mathbb{E}[u\mid Z,W]\neq 0$ while $\mathbb{E}[v\mid Z,W]=0$. Classical weak-instrument problems arise when the first-stage coefficients $\pi$ are small, so that the concentration parameter is close to zero and the instruments explain only a tiny fraction of the variation in $X$; see staiger1997instrumental, stock2005asymptotic, and dufour1997some for canonical analyses. In this case, 2SLS and related estimators exhibit large finite-sample bias toward the OLS estimand, heavy-tailed or multimodal sampling distributions, and nonstandard asymptotic behavior andrews2014weak,andrews2019weak. These issues are amplified in many-instrument settings bekker1994alternative,hansen2008estimation, where the effective first-stage signal per instrument becomes small even if the joint $F$-statistic is sizable.

The weak-identification phenomena in triangular models are conceptually similar but operate at the level of distributions rather than means. Instead of small first-stage coefficients on $Z$ in a linear projection of $X$ on $Z$, the problem is that the instrument-induced shifts in $F_{X\mid Z}(\cdot\mid z)$ are too limited or nearly fail completeness, so the mapping from structural objects to observables is nearly singular. As emphasized by l19_id_zoo, both linear and nonlinear models admit a continuum of identification strengths, ranging from strong point identification with regular asymptotics, through weak or set identification with nonstandard limits, to complete lack of identification.

Our Q--LS estimator can be viewed as a two-step device that first maps the original instrument vector $Z$ into a scalar distributional instrument \[ h(Z) \in \mathcal{H} = \Big\{ \int_0^1 \omega(\tau) Q_{X\mid Z}(\tau\mid Z)\,d\tau : \omega \in L^2(0,1)\Big\}, \] and then applies conventional linear IV using $h(Z)$ in place of $Z$. From this perspective, Assumption (ref) is the analogue of a strong-IV condition on the generated instrument $h(Z)$: it rules out sequences of designs in which the covariance between $X$ and $h(Z)$ drifts to zero at a $\sqrt{n}$–local rate.

We therefore do not claim that Q--LS is generally robust in the weak-IV sense of staiger1997instrumental or that it uniformly dominates Anderson--Rubin or conditional tests andrews2019weak,olea2013robust when instruments are arbitrarily weak. Rather, the contribution of our distributional relevance framework is to highlight a different margin of strength: designs in which $\operatorname{Var}(\mathbb{E}[X\mid Z])$ is small or zero, so that instruments appear weak for the mean, may still be strong for the distribution, with $\mathbb{E}[h_{\mathrm{opt}}(Z)^2]$ bounded away from zero. In such cases, Q--LS leverages distributional variation in $X\mid Z$ to construct an effective instrument $h(Z)$ that satisfies a strong-IV condition even when the original mean-based instrument fails conventional first-stage diagnostics. Our notion of distributional relevance thus specifies a “strong-instrument’’ regime at the distributional level and complements both the generative DIV approach and the classical weak-IV toolkit. This allows us to focus on designs where instruments may be weak for linear first stages but remain informative about the distribution of $X$, and hence can still support identification and efficient estimation once the first stage is recast in distributional terms.

Quantile regression and endogeneity

Quantile regression was first formalised as an optimization problem minimizing asymmetrically weighted absolute residuals by kb78_qr, to estimate any chosen quantile $\tau\in(0,1)$ of the distribution of an outcome variable $X$ conditional on a set of exogenous regressors $Z$,

equation*[equation* omitted — 80 chars of source]

where $F_{X|Z}(x|z)$ is the conditional cumulative distribution function (CDF) of $X$ given $Z=z$.

The linear quantile regression model specifies

equation*[equation* omitted — 55 chars of source]

where $\pi(\tau)$ is estimated by solving the expected check-loss function

equation*[equation* omitted — 120 chars of source]

This objective function is convex but non-differentiable, yielding a linear programming problem that can be solved efficiently in moderately large samples.

Identification of quantile regression relies on mild regularity conditions, including correct specification of the conditional quantile function, sufficient variation in the regressors, and a rank condition ensuring uniqueness of the conditional quantile. These conditions do not require parametric assumptions on the error distribution

aai02_qiv were the first to combine quantile methods with instrumental variables to estimate distributional treatment effects—that is, how a binary treatment affects quantiles of the outcome distribution—under endogeneity. ch05_qiv generalized this approach to estimate treatment effects at different points of the outcome distribution (quantile treatment effects) under endogeneity and introduced the notion of the structural quantile function. See w20 for a detailed comparison of ch05_qiv and aai02_qiv. Since these seminal contributions, many variants of quantile instrumental variable methods have been proposed. For example, p20_uncon_qiv develop an instrumental-variable approach to estimate unconditional quantile treatment effects, while clp16_gqiv propose methods for group-level quantile treatment effects.

There is also a growing literature that integrates over a grid of quantiles to estimate endogenous structural functions. cfk15_qiv propose a Censored Quantile Instrumental Variables (CQIV) estimator that uses a control-function approach to estimate structural quantiles of a censored outcome with endogenous regressors. In their framework, the control function is estimated by integrating over a grid of quantiles. cfnsv20_cfa extend the control-function framework of cfk15_qiv and develop a unified approach for estimating average, quantile, and distributional structural functions. The control-function estimator we propose for nonseparable triangular models estimates the control function using the same strategy as cfnsv20_cfa. See Appendix (ref) for further details.

Existing ways to circumvent weak identification

There is a substantial literature that seeks to circumvent the problem of weak identification by imposing specific functional form restrictions. In this section, we discuss several key strands of this literature and how it relates to our proposed notion of distributional relevance.

There is a growing literature that imposes functional form restrictions on higher moments. lewbel1997constructing show that, in linear models with measurement error and endogeneity, instruments can be constructed from second-moment variation: \[ Z_i^{\text{Lev}} = (Z_i - \bar Z)(X_i - \bar X), \] provided that the variation of $X$ conditional on $Z$ changes with $Z$ even when $\mathbb{E}[X\mid Z]$ is constant. This identification strategy exploits variance shifts. Several studies have extended lewbel1997constructing and used second-moment restrictions to achieve identification under weaker functional form restrictions; see, for example, kv10_iv_cfa_hetro,l12_hetro. These restrictions can be viewed as a specific form of distributional relevance (variance shifts) and are therefore encompassed by our framework.

t23_iv_icm also impose a functional form restriction that ensures identification in models of the form $\mathbb{E}[Y - X\beta \mid Z]=0$ without excluded instruments. Specifically, t23_iv_icm assume a nonlinear completeness condition requiring injectivity of the conditional mean operator:

\[ \mathbb{E}[X\mid Z] \tau = 0 \quad\Rightarrow\quad \tau = 0. \]

Completeness relies on sufficiently rich nonlinear mean dependence of $X$ conditional on $Z$, that is, on the mapping $X \mapsto \mathbb{E}[X\mid Z]$.

Our approach is fundamentally different from t23_iv_icm. Identification does not rely on the conditional mean operator or its injectivity. Instead, distributional variation in $X$ suffices to identify the average structural function $\mu(x)$, even when $\mathbb{E}[X\mid Z]$ is constant and completeness fails. Consequently, our results do not require completeness, mean dependence, or any functional-form restrictions of the type imposed by t23_iv_icm.

babii2020completeness study estimation in nonidentified linear inverse problems of the from $K\varphi = r$, where $K$ is a compact operator between Hilbert spaces. In nonparametric IV models, $K$ corresponds to the conditional expectation operator $(K\varphi)(z)=\mathbb{E}[\varphi(X)\mid Z=z]$. When completeness (injectivity of $K$) fails, $\varphi$ is not point-identified, but spectral regularization methods converge to the “best approximation” $\varphi_1$ in the orthogonal complement of the null space of $K$. Their work characterizes risk bounds and asymptotic distributions under varying degrees of identification.

Our framework differs in two central respects. First, identification is not formulated as an inverse problem with operator $K\varphi=r$. Instead, we use a structural triangular model with distributional relevance, which yields full nonparametric identification of the average structural function $\mu(x)$ when the usual nonparametric IV completeness condition fails. Second, whereas babii2020completeness obtain convergence to a best approximation under nonidentification, our results deliver point identification of $\mu(x)$ by exploiting the rank structure of the first stage and the distributional variation of $X$. Consequently, our contribution is complementary: we provide an identification and estimation framework that does not rely on operator injectivity or completeness, and remains applicable with instruments that are mean-irrelevant but distributionally strong.

Linear IV setup and distributional relevance

Linear IV model

We consider the linear triangular model

equation[equation omitted — 136 chars of source]

where $Y$ is the outcome, $X$ is a scalar endogenous regressor, $Z_1$ is a vector of exogenous covariates included in the outcome equation, and $Z = (Z_1, Z_2)$ collects $Z_1$ and a vector of excluded instruments $Z_2$.

In the classical IV framework, instrument strength is measured by variation in the conditional mean of $X$ given $Z$: \[ m(Z) := \mathbb{E}[X\mid Z], \qquad \operatorname{Var}(m(Z)) > 0. \] When $\operatorname{Var}(m(Z))$ is close to zero, instruments are labelled “weak” and standard 2SLS inference is unreliable.

Our aim is to replace this mean-based notion of strength with a weaker and more general concept based on the full conditional distribution of $X$.

Distributional relevance

Let \[ F_{X\mid Z}(x\mid Z) := \mathbb{P}(X \le x \mid Z), \qquad \bar F_X(x) := \mathbb{E}\big[ F_{X\mid Z}(x\mid Z)\big] \] denote the conditional and unconditional distribution functions of $X$. We measure variation in $F_{X\mid Z}(\cdot\mid Z)$ using an $L^2$ norm over $x$. For concreteness, let $\lambda$ be Lebesgue measure on a compact interval containing the support of $X$, and define \[ \| F_{X\mid Z}(\cdot\mid Z) \|_{L^2(\lambda)}^2 := \int \big( F_{X\mid Z}(x\mid Z) \big)^2\,d\lambda(x). \]

definition[Mean relevance] We say that the instrument $Z$ is mean-relevant for $X$ if \[ \operatorname{Var}\big( \mathbb{E}[X \mid Z] \big) = \mathbb{E}\Big[\big(\mathbb{E}[X\mid Z] - \mathbb{E}[X]\big)^2\Big] > 0. \]
definition[Distributional relevance] We say that the instrument $Z$ is distributionally relevant for $X$ if \begin{equation} \mathbb{E}\Big[ \big\| F_{X\mid Z}(\cdot \mid Z) - \bar F_X(\cdot) \big\|_{L^2(\lambda)}^2 \Big] > 0. \end{equation} Equivalently, there exist $z_p,z_q$ in the support of $Z$ such that \[ \big\| F_{X\mid Z}(\cdot \mid z_p) - F_{X\mid Z}(\cdot \mid z_q) \big\|_{L^2(\lambda)} > 0. \]

Any change in $\mathbb{E}[X\mid Z]$ across $Z$ induces a difference in $F_{X\mid Z}(\cdot\mid Z)$, so mean relevance implies distributional relevance. The converse need not hold: $Z$ can be distributionally relevant even when $\mathbb{E}[X\mid Z]$ is constant in $Z$. This motivates the following concept.

definition[Purely distributional instruments] We say that $Z$ is a purely distributional instrument for $X$ if \begin{enumerate}[label=(\roman*)] • $Z$ is distributionally relevant for $X$ in the sense of Definition (ref), and • $Z$ is mean-irrelevant for $X$, i.e. \[ \operatorname{Var}(\mathbb{E}[X\mid Z]) = 0 \quad\Longleftrightarrow\quad \mathbb{E}[X\mid Z=z] = \mathbb{E}[X] \quad \text{for all } z. \] \end{enumerate}

Classical IV analysis treats mean relevance---variation in $\mathbb{E}[X\mid Z]$---as the benchmark notion of instrument strength. Distributional relevance weakens this requirement: it allows $\mathbb{E}[X\mid Z]$ to be completely flat in $Z$, as long as the conditional distribution $F_{X\mid Z}(\cdot\mid Z)$ varies nontrivially in $L^2$. In this sense, an instrument can be “strong” for the distribution of $X$ even when it is “weak” for its mean. We refer to instruments that satisfy distributional relevance but have $\operatorname{Var}(\mathbb{E}[X\mid Z])=0$ as purely distributional: they convey no information about $\mathbb{E}[X\mid Z]$ but still transmit rich information about the risk environment through higher moments and tail behavior.

In the linear IV analysis that follows, distributional relevance and the linear quantile representation for $X\mid Z$ guarantee the existence of an optimal Q--LS instrument that is relevant in the classical sense unless the projection degenerates. In the nonseparable triangular model in Appendix (ref), the condition that $Z$ is purely distributional is enough for the identification of the average structural function via a control function based on $F_{X\mid Z}(X\mid Z)$.

remark[Illustrative example] Let $Z_2\in\{0,1\}$ be a binary instrument with $\mathbb{P}(Z_2=0)=\mathbb{P}(Z_2=1)=1/2$, and let $U\sim\mathcal{N}(0,1)$ be independent of $Z_2$. Define \[ X = \sigma(Z_2)\,U, \qquad \sigma(0)=1,\ \sigma(1)=2. \] Then $\mathbb{E}[X\mid Z_2=z_2]=0$ for all $z_2$, so $\operatorname{Var}(\mathbb{E}[X\mid Z_2])=0$ and $Z_2$ is mean-irrelevant. However, the conditional distributions differ: $X\mid Z_2=0\sim\mathcal{N}(0,1)$ and $X\mid Z_2=1\sim\mathcal{N}(0,4)$, so $F_{X\mid Z_2}(\cdot\mid 0)\neq F_{X\mid Z_2}(\cdot\mid 1)$ in $L^2$. Thus $Z_2$ is a purely distributional instrument. In a classical first-stage regression of $X$ on $Z_2$, the coefficient on $Z_2$ would be zero, but the instrument carries substantial information about the dispersion of $X$.

In what follows, distributional relevance will be our primitive assumption on the instrument. We show that, under a linear quantile representation for $X\mid Z$, it is enough to construct a strong instrument as a linear functional of the conditional quantile process, and to obtain consistent and asymptotically normal estimates of $(\alpha,\beta,\gamma)$ in the linear model (ref).

The Q--LS estimator in the linear model

We now define the Quantile Least Squares IV (Q--LS) estimator. The key object is an optimal instrument constructed as a linear functional of the conditional quantile process of $X\mid Z$.

Conditional quantiles and quantile-aggregated instruments

We work with the conditional quantile function of $X$ given $Z$. For each $\tau\in(0,1)$, assume

equation[equation omitted — 117 chars of source]

where $\pi_0(\tau)\in\mathbb{R}^{\dim(Z)}$ is a measurable function of $\tau$. We write \[ g(\tau,Z) := Q_{X\mid Z}(\tau\mid Z) = Z'\pi_0(\tau). \] Under mild regularity, distributional relevance is equivalent to the statement that the conditional quantile process $\{g(\tau,Z):\tau\in(0,1)\}$ is non-constant in $Z$ in $L^2$.\footnote{For concreteness, we implement Q--LS using linear quantile regressions of $X$ on $Z$ at a finite grid of quantile indices. This choice is purely for illustration: any alternative specification for the conditional quantiles (e.g., nonlinear, spline or series-based, or machine–learning first stages) could be used to construct the dictionary $G(Z)$ and hence the Q--LS instrument.}

We restrict attention to instruments that are linear functionals of this quantile process. Let $\omega(\cdot)$ be a square-integrable weight function on $(0,1)$. We define the quantile-aggregated instrument

equation[equation omitted — 167 chars of source]

The corresponding moment conditions for the structural parameter $\theta := (\alpha,\beta,\gamma')'$ are

equation[equation omitted — 203 chars of source]

Let \[ \mathcal{H} := \Big\{ h_\omega(\cdot): h_\omega(Z) = \int_0^1 \omega(\tau)\,g(\tau,Z)\,d\tau, \ \omega\in L^2(0,1) \Big\} \] be the linear span (in $L^2$) of the conditional quantile process.

Optimal Q--LS and relevance under distributional relevance

Following the classical optimal IV logic for a single endogenous regressor, we consider the element of $\mathcal{H}$ that best predicts $X$ in mean square error.

definition[Optimal Q--LS] The optimal Q--LS weight function $\omega_{\mathrm{opt}}(\cdot)$ is any solution to \begin{equation} \omega_{\mathrm{opt}} \in \arg\min_{\omega\in L^2(0,1)} \mathbb{E}\!\left[ \big( X - h_\omega(Z) \big)^2 \right], \end{equation} and the associated optimal Q--LS instrument is \begin{equation} h_{\mathrm{opt}}(Z) := h_{\omega_{\mathrm{opt}}}(Z) = \int_0^1 \omega_{\mathrm{opt}}(\tau)\, Q_{X\mid Z}(\tau\mid Z)\,d\tau. \end{equation}

By construction, $h_{\mathrm{opt}}(Z)$ is the $L^2$-projection of $X$ onto the closed linear span $\mathcal{H}$ of quantile-generated instruments. The next lemma shows that, whenever this projection is nonzero, the resulting instrument is automatically relevant for $X$ in the classical sense.

lemma[Distributional relevance and optimal Q--LS] Suppose Assumption (ref) and (ref) hold, and let $h_{\mathrm{opt}}$ be defined as in Definition (ref). Then \[ X = h_{\mathrm{opt}}(Z) + r(Z), \qquad r \perp \mathcal{H}, \] and either \[ h_{\mathrm{opt}}(Z) = 0 \quad \text{a.s.}, \] or \[ \operatorname{cov}\big(h_{\mathrm{opt}}(Z), X\big) = \mathbb{E}\big[h_{\mathrm{opt}}(Z)^2\big] > 0. \] In particular, whenever $h_{\mathrm{opt}}$ is non-degenerate it is a relevant instrument for $X$.
proofBy definition of $h_{\mathrm{opt}}$ as the $L^2$-projection of $X$ onto the closed linear subspace $\mathcal{H}$, we can write \[ X = h_{\mathrm{opt}}(Z) + r(Z), \] where $r(Z)$ is the projection residual and satisfies $\mathbb{E}[h(Z)\,r(Z)] = 0$ for all $h\in\mathcal{H}$. In particular, $\mathbb{E}[h_{\mathrm{opt}}(Z)\,r(Z)] = 0$, so \[ \mathbb{E}[X\,h_{\mathrm{opt}}(Z)] = \mathbb{E}\big[(h_{\mathrm{opt}}(Z)+r(Z))\,h_{\mathrm{opt}}(Z)\big] = \mathbb{E}\big[h_{\mathrm{opt}}(Z)^2\big] + \mathbb{E}\big[h_{\mathrm{opt}}(Z)\,r(Z)\big] = \mathbb{E}\big[h_{\mathrm{opt}}(Z)^2\big]. \] Thus either $h_{\mathrm{opt}}(Z)=0$ a.s., in which case both sides vanish, or $\mathbb{E}[h_{\mathrm{opt}}(Z)^2]>0$ and $h_{\mathrm{opt}}$ is correlated with $X$.

Lemma (ref) delivers a clean dichotomy: the optimal Q--LS instrument is either identically zero or, if nonzero, it is automatically relevant and its “strength” is measured by $\mathbb{E}[h_{\mathrm{opt}}(Z)^2]$. Distributional relevance guarantees that $\mathcal{H}$ contains more than just constants, but it does not by itself rule out the degenerate case $h_{\mathrm{opt}}\equiv 0$. A simple example is a pure variance-shift design, \[ X = \sigma(Z)\,U, \qquad \mathbb{E}[U]=0,\quad U \perp \!\!\! \perp Z, \] with $\sigma(Z)$ non-constant. In this case the conditional distributions $F_{X\mid Z}(\cdot\mid Z)$ vary with $Z$ (distributional relevance), yet $\operatorname{cov}(X,h(Z))=0$ for every $h\in\mathcal{H}$ so that $h_{\mathrm{opt}}(Z)\equiv 0$. Our Design B in Section (ref) provides a Monte Carlo illustration of this phenomenon.

For identification, we therefore add an explicit relevance condition that is the exact analogue of the usual mean-relevance condition in classical IV.

assumption[Strong distributional relevance] Under the linear quantile representation (ref), the optimal Q--LS instrument $h_{\mathrm{opt}}(Z)$ defined in (ref) satisfies \[ \mathbb{E}\big[h_{\mathrm{opt}}(Z)^2\big] \ge c > 0 \] for some constant $c$ that does not depend on the sample size $n$.

Assumption (ref) is the distributional analogue of the usual “strong IV” condition in linear models, where one assumes that the first-stage coefficient on the excluded instruments does not drift to zero as $n$ grows (e.g.\ staiger1997instrumental, stock2005asymptotic. Here we impose a lower bound on the variance of the optimal Q--LS instrument $h_{\mathrm{opt}}(Z)$, ruling out sequences of designs in which instruments become weak either in the mean or in the distribution. Throughout our asymptotic analysis we work under Assumption (ref), so our results should be interpreted as strong-distributional-IV asymptotics.

Under Assumption (ref), Lemma (ref) implies that $h_{\mathrm{opt}}$ is a non-degenerate, relevant instrument constructed purely from the conditional distribution of $X\mid Z$, even when $\mathbb{E}[X\mid Z]$ is constant in $Z$.

Throughout the linear IV analysis, we assume that distributional relevance holds and that $X\mid Z$ admits the linear quantile representation in (ref). Under these conditions, Lemma (ref) shows that the optimal Q--LS instrument $h_{\mathrm{opt}}(Z)$ is either identically zero or a classically relevant instrument for $X$, and Assumption (ref) rules out the degenerate case.

Finite-grid implementation

In practice, we approximate the integral in (ref) using a finite grid of quantiles. Let \[ 0 < \tau_1 < \cdots < \tau_K < 1 \] be a set of quantile indexes with gap of order $1/K$, and let $\Delta\tau_k$ denote the associated quadrature weights (for simplicity, one can take $\Delta\tau_k = 1/K$).

For each $\tau_k$, we estimate the first-stage quantile regression

equation[equation omitted — 173 chars of source]

and define the estimated quantile-based instruments \[ \widehat{g}_k(Z_i) := Z_i'\widehat{\pi}_n(\tau_k), \qquad i=1,\dots,n,\quad k=1,\dots,K. \] Stacking them, let \[ \widehat{G}_i := \big( \widehat{g}_1(Z_i), \dots, \widehat{g}_K(Z_i) \big)', \qquad \widehat{G} :=

pmatrix[pmatrix omitted — 69 chars of source]

\in \mathbb{R}^{n\times K}. \]

\paragraph{Discrete Q--LS weights.} A natural discrete analogue of (ref) chooses weights $w\in\mathbb{R}^K$ to best predict $X$ from the $K$ quantile instruments:

equation[equation omitted — 160 chars of source]

so that, when $\widehat{G}'\widehat{G}$ is invertible,

equation[equation omitted — 146 chars of source]

The corresponding finite-sample Q--LS first-stage prediction is

equation[equation omitted — 192 chars of source]

This $\widehat{X}_i^{\mathrm{Q\text{--}LS}}$ is the empirical best linear predictor of $X_i$ in the span of the $K$ quantile-based instruments.

\paragraph{Second stage.} Let $y = (Y_1,\dots,Y_n)'$ be the outcome vector, and define the regressor and instrument matrices \[ S = [\mathbf{1},\, X,\, Z_1], \qquad M = [\widehat{X}^{\mathrm{Q\text{--}LS}},\, Z_1], \] where $\widehat{X}^{\mathrm{Q\text{--}LS}} = (\widehat{X}_1^{\mathrm{Q\text{--}LS}}, \dots,\widehat{X}_n^{\mathrm{Q\text{--}LS}})'$ and $\mathbf{1}$ denotes the $n$-vector of ones. The Q--LS estimator of $\theta = (\alpha,\beta,\gamma')'$ is the usual 2SLS estimator

equation[equation omitted — 154 chars of source]

Asymptotic properties

We now state conditions under which the Q--LS estimator is consistent and asymptotically normal. The key identification condition is distributional relevance, which, via Lemma (ref), implies relevance of the optimal Q--LS instrument.

assumption[Regularity and identification] \leavevmode \begin{enumerate}[label=(\roman*)] • The sample $\{(Y_i,X_i,Z_i)\}_{i=1}^n$ is i.i.d., and $\mathbb{E}\|Z\|^2 < \infty$, $\mathbb{E}|X|^2 < \infty$, $\mathbb{E}|Y|^2 < \infty$. • For each $\tau\in(0,1)$, the conditional quantile $Q_{X\mid Z}(\tau\mid Z)$ exists and satisfies (ref), with $\pi_0(\cdot)$ continuous on $(0,1)$ and $\int_0^1 \|\pi_0(\tau)\|^2 d\tau < \infty$. • The structural error satisfies $\mathbb{E}[\varepsilon\mid Z]=0$ and $\mathbb{E}[\varepsilon^2\mid Z] < \infty$. • Distributional relevance holds in the sense of Definition (ref), and the span $\mathcal{H}$ generated by $\{g(\tau,Z):\tau\in(0,1)\}$ is not reduced to constants. Let $h_{\mathrm{opt}}$ be defined as in Definition (ref), and suppose the population matrix \[ \mathbb{E}\big[ S_0(Z)' S_0(Z) \big], \qquad S_0(Z) = (1, X, Z_1') \] is nonsingular. • The quantile grid $\{\tau_k\}_{k=1}^K$ is dense in $(0,1)$ as $K=K_n\to\infty$, and the number of quantiles satisfies $K_n^2/n \to 0$ as $n\to\infty$. \end{enumerate}
prop[Consistency of the Q--LS estimator] Suppose Assumptions (ref) and (ref) hold, and the weights $\widehat{w}_n$ are defined by (ref) with $K=K_n\to\infty$ and $K_n^2/n\to 0$. Then \[ \widehat{\theta}_n^{\mathrm{Q\text{--}LS}} \xrightarrow{p} \theta \qquad\text{as } n\to\infty. \]
proofBy Assumption (ref)(ii) and standard quantile regression theory, we have, for each fixed $\tau$, $\widehat{\pi}_n(\tau) \xrightarrow{p} \pi_0(\tau)$, and, since the grid $\{\tau_k\}$ becomes dense and $K_n^2/n\to 0$, the array $\{\widehat{\pi}_n(\tau_k)\}$ uniformly approximates $\{\pi_0(\tau_k)\}$ in mean square. It follows that \[ \max_{1\le k\le K_n} \mathbb{E}\big[ \big( \widehat{g}_k(Z) - g(\tau_k,Z) \big)^2 \big] \to 0, \] and hence $\widehat{G}'\widehat{G}/n$ converges in probability to the corresponding population limit, and similarly for $\widehat{G}'X/n$. Under $K_n^2/n\to 0$, the estimation error from the first-stage quantile regressions is $o_p(n^{-1/2})$ in the moment conditions. Therefore the discrete weights $\widehat{w}_n$ converge in probability to the population weights $w_0$ that minimize $\mathbb{E}[(X-G_0(Z)'w)^2]$ in the corresponding finite-dimensional approximation space, where $G_0(Z)$ stacks the $g(\tau_k,Z)$'s. As $K_n\to\infty$ and the grid becomes dense, the discrete approximation $G_0(Z)'w_0$ converges in $L^2$ to the continuum optimal instrument $h_{\mathrm{opt}}(Z)$ defined in Definition (ref). In particular, \[ \mathbb{E}\big[ \big( \widehat{X}_i^{\mathrm{Q\text{--}LS}} - h_{\mathrm{opt}}(Z_i) \big)^2 \big] \to 0. \] Standard arguments for 2SLS with generated instruments then imply that $\widehat{\theta}_n^{\mathrm{Q\text{--}LS}}$ converges in probability to the unique solution of the population IV moment condition with instrument $h_{\mathrm{opt}}(Z)$, which is $\theta_0$ by Assumption (ref)(iv).

For asymptotic normality, let \[ S_i := \big(1, X_i, Z_{1i}'\big)', \qquad M_i := \big(\widehat{X}_i^{\mathrm{Q\text{--}LS}}, Z_{1i}'\big)' \] denote the stacked regressor and instrument vectors used in the second stage, and let \[ \widehat{\varepsilon}_i := Y_i - S_i'\widehat{\theta}_n^{\mathrm{Q\text{--}LS}}. \] Define the sample matrices \[ \widehat{A}_n := \frac{1}{n}\sum_{i=1}^n S_i M_i', \qquad \widehat{Q}_n := \frac{1}{n}\sum_{i=1}^n M_i M_i', \] and the heteroskedasticity-robust “meat” matrix \[ \widehat{\Omega}_n := \frac{1}{n}\sum_{i=1}^n \widehat{\varepsilon}_i^{\,2}\, M_i M_i'. \]

assumption[Additional conditions for asymptotic normality] In addition to Assumption (ref), suppose that: \begin{enumerate}[label=(\roman*)] • The fourth moments of $(Y,X,Z)$ are finite: $\mathbb{E}\|Z\|^4 < \infty$, $\mathbb{E}|X|^4 < \infty$, $\mathbb{E}|Y|^4 < \infty$. • The conditional fourth moment of the structural error is finite: $\mathbb{E}[\varepsilon^4\mid Z] < \infty$ a.s. • The eigenvalues of $\mathbb{E}[M_i^0 M_i^{0\prime}]$ are bounded away from zero and infinity, where $M_i^0 = (h_{\mathrm{opt}}(Z_i),Z_{1i}')'$ is the population instrument vector associated with the optimal Q--LS, and the same holds for $\mathbb{E}[S_i S_i']$. \end{enumerate}

Let \[ A := \mathbb{E}[S_i M_i^{0\prime}], \qquad Q := \mathbb{E}[M_i^0 M_i^{0\prime}], \qquad \Omega := \mathbb{E}\big[\varepsilon_i^{2} M_i^0 M_i^{0\prime}\big]. \] The population asymptotic covariance matrix of the Q--LS estimator can then be written in the usual IV “sandwich” form

equation[equation omitted — 166 chars of source]
prop[Asymptotic normality of the Q--LS estimator] Suppose Assumptions (ref), (ref) and (ref) hold. Then \[ \sqrt{n}\big( \widehat{\theta}_n^{\mathrm{Q\text{--}LS}} - \theta \big) \ \xrightarrow{d}\ \mathcal{N}\big(0, V_{\mathrm{Q\text{--}LS}}\big), \] where $V_{\mathrm{Q\text{--}LS}}$ is given in (ref). Moreover, the heteroskedasticity-robust 2SLS covariance matrix computed with the generated Q--LS instrument, \begin{equation} \widehat{V}_n^{\mathrm{Q--LS}} := \big(\widehat{A}_n \widehat{Q}_n^{-1} \widehat{A}_n'\big)^{-1} \big(\widehat{A}_n \widehat{Q}_n^{-1} \widehat{\Omega}_n \widehat{Q}_n^{-1} \widehat{A}_n'\big) \big(\widehat{A}_n \widehat{Q}_n^{-1} \widehat{A}_n'\big)^{-1}, \end{equation} is consistent for $V_{\mathrm{Q\text{--}LS}}$, in the sense that $\widehat{V}_n^{\mathrm{Q\text{--}LS}} \xrightarrow{p} V_{\mathrm{Q\text{--}LS}}$. Consequently, for any fixed vector $c\in\mathbb{R}^{\dim(\theta_0)}$, \[ \frac{ \sqrt{n}\,c'\big(\widehat{\theta}_n^{\mathrm{Q\text{--}LS}} - \theta\big) }{ \sqrt{c'\widehat{V}_n^{\mathrm{Q\text{--}LS}}c} } \ \xrightarrow{d}\ \mathcal{N}(0,1). \]
proofBy Proposition (ref), the generated instrument $\widehat{X}_i^{\mathrm{Q\text{--}LS}}$ converges in $L^2$ to the population optimal instrument $h_{\mathrm{opt}}(Z_i)$, and the Q--LS estimator solves the sample moment condition \[ \frac{1}{n}\sum_{i=1}^n M_i\big(Y_i - S_i'\theta\big) = 0, \] with $M_i = (\widehat{X}_i^{\mathrm{Q\text{--}LS}},Z_{1i}')'$. The estimation error in $\widehat{X}_i^{\mathrm{Q\text{--}LS}}$ is $o_p(n^{-1/2})$ in the associated moment conditions under $K_n^2/n\to 0$, so the first-order asymptotics coincide with those of a conventional 2SLS estimator that uses the population instrument $M_i^0 = (h_{\mathrm{opt}}(Z_i),Z_{1i}')'$. Standard arguments for IV/GMM with i.i.d.\ observations, finite fourth moments, and nonsingular Jacobian then yield the stated result, and the consistency of $\widehat{V}_n^{\mathrm{Q\text{--}LS}}$ follows by the law of large numbers and Slutsky's theorem.

In applications, we report standard errors as the square roots of the diagonal elements of $\widehat{V}_n^{\mathrm{Q\text{--}LS}}$ in (ref), which coincide with the usual heteroskedasticity-robust 2SLS standard errors computed using the generated Q--LS instrument and the exogenous regressors $Z_1$ as instruments.

Proposition (ref), is derived under Assumption (ref), which enforces a strong form of distributional relevance for the generated instrument $h(Z)$. As in the classical linear-IV literature, this rules out weak-IV sequences in which the covariance between the endogenous regressor and the instrument shrinks at a $\sqrt{n}$--local rate. In particular, our asymptotic normality result for Q--LS should be interpreted as a strong-identification limit theory: it guarantees standard Gaussian inference when the optimal distributional instrument $h_{\mathrm{opt}}(Z)$ has nonvanishing variance, but it does not characterize the behavior of Q--LS under general weak-IV sequences. Extending weak-identification robust inference to settings with distributional instruments would require combining our construction with the robust testing approaches of, for example, andrews2019weak or olea2013robust, which we leave for future work.

Relation to classical 2SLS

In this section we compare the quantile least squares (Q--LS) estimator to classical two-stage least squares (2SLS) in settings where both procedures are valid.

2SLS and optimal mean-based instruments

In the classical IV framework with a single endogenous regressor, relevance is formulated in terms of the conditional mean \[ m(Z) := \mathbb{E}[X\mid Z]. \] Under homoskedasticity, the “best” instrument for a single endogenous regressor is the efficient instrument $h^*(Z)$, which is proportional to the best $L^2$ predictor of $X$ given $Z$, i.e.\ to $m(Z)$. The optimal 2SLS estimator uses instruments in the linear span of $\{1,Z_1',m(Z)\}$ and is semiparametrically efficient for $\theta_0$ within this mean-based class.

Q--LS versus 2SLS in a location model

By contrast, Q--LS is built from the entire conditional distribution of $X\mid Z$ via its conditional quantiles. Under (ref), we construct quantile-aggregated instruments of the form \[ h_\omega(Z) = \int_0^1 \omega(\tau)\,Q_{X\mid Z}(\tau\mid Z)\,d\tau, \] and define the optimal Q--LS instrument $h_{\mathrm{opt}}$ as in Definition (ref), i.e.\ as the best $L^2$ predictor of $X$ within the quantile-generated class $\mathcal{H}$. The Q--LS estimator $\widehat{\theta}_n^{\mathrm{Q\text{--}LS}}$ is the 2SLS estimator that uses $(h_{\mathrm{opt}}(Z),Z_1)$ as instruments.

The following proposition shows that, in a Gaussian/location setting with mean relevance, Q--LS and optimal 2SLS are asymptotically equivalent.

prop[Q--LS versus 2SLS in a location model] Suppose that, in addition to Assumptions (ref)--(ref), the conditional distribution of $X\mid Z$ belongs to a location family, i.e.\ there exists a random variable $U$ with $\mathbb{E}[U]=0$ such that \[ X \mid Z \ \overset{d}{=}\ m(Z) + U, \] and the distribution of $U$ does not depend on $Z$. Then: \begin{enumerate}[label=(\roman*)] • For each $\tau\in(0,1)$ we have \[ Q_{X\mid Z}(\tau\mid Z) = m(Z) + c(\tau), \] for some scalar function $c(\cdot)$ independent of $Z$. • Any quantile-aggregated instrument $h_\omega(Z)$ can be written as \[ h_\omega(Z) = a_\omega\,m(Z) + b_\omega, \qquad a_\omega = \int_0^1 \omega(\tau)\,d\tau, \ \ b_\omega = \int_0^1 \omega(\tau)\,c(\tau)\,d\tau. \] • The optimal Q--LS instrument $h_{\mathrm{opt}}(Z)$ is proportional to $m(Z)$ and, up to an irrelevant rescaling of the instrument, the Q--LS estimator coincides with the optimal 2SLS estimator: \[ \widehat{\theta}_n^{\mathrm{Q\text{--}LS}} \ \equiv\ \widehat{\theta}_n^{\mathrm{2SLS}} + o_p(n^{-1/2}), \] so that $V_{\mathrm{Q\text{--}LS}} = V_{\mathrm{2SLS}}$. \end{enumerate}
proofUnder the location assumption, $X\mid Z$ has distribution equal to $U$ shifted by $m(Z)$, so its $\tau$-quantile satisfies $Q_{X\mid Z}(\tau\mid Z) = m(Z) + c(\tau)$, where $c(\tau)$ is the $\tau$-quantile of $U$ and does not depend on $Z$. Any $h_\omega\in\mathcal{H}$ is therefore of the form $a_\omega m(Z) + b_\omega$ as stated. Minimizing $\mathbb{E}[(X-h_\omega(Z))^2]$ over $\omega$ then reduces to minimizing $\mathbb{E}[(X-a m(Z)-b)^2]$ over $(a,b)$, whose unique solution is $a=1$ and $b=0$ (up to the usual scale normalization of instruments), so $h_{\mathrm{opt}}(Z)$ is proportional to $m(Z)$. The Q--LS estimator thus corresponds to 2SLS with the optimal mean-based instrument and has the same asymptotic distribution.

Proposition (ref) shows that, in a conventional Gaussian or location setting with mean relevance, Q--LS and optimal 2SLS are asymptotically equivalent: they use instruments that differ only by an irrelevant scale factor and attain the same efficiency bound.

lemma[Equivalence between 2SLS and plug-in OLS on the projected regressor] Let $\{(Y_i,X_i,Z_{1i},Z_{2i})\}_{i=1}^n$ be a sample and consider the linear structural model \[ Y_i = \alpha + \beta X_i + Z_{1i}'\gamma + \varepsilon_i, \qquad \mathbb{E}[\varepsilon_i \mid Z_i] = 0, \] with $Z_i = (Z_{1i}',Z_{2i}')'$. Let \[ S := \begin{bmatrix} \mathbf{1} & X & Z_1 \end{bmatrix} \in\mathbb{R}^{n\times (1+1+d_1)}, \qquad M := \begin{bmatrix} Z_1 & G \end{bmatrix} \in\mathbb{R}^{n\times (d_1+K)}, \] where $G$ collects a set of $K$ instrument functions (e.g.\ quantile-based instruments) and $Z_1$ denotes the included exogenous regressors. Let \[ P_M := M(M'M)^{-1}M' \] be the projection matrix onto the column space of $M$ and define the projected regressor \[ \tilde X := P_M X. \] Consider the following two estimators for $\theta = (\alpha,\beta,\gamma')'$: \begin{enumerate}[label=(\roman*)] • The 2SLS estimator with instrument matrix $M$: \[ \widehat{\theta}^{\mathrm{2SLS}} := (S'P_M S)^{-1} S'P_M Y. \] • The plug-in OLS estimator based on the projected regressor $\tilde X$: \[ \widehat{\theta}^{\mathrm{OLS-proj}} := \big(\tilde S'\tilde S\big)^{-1}\tilde S'Y, \qquad \tilde S := \begin{bmatrix} \mathbf{1} & \tilde X & Z_1 \end{bmatrix}. \] \end{enumerate} Then, in finite samples, \[ \widehat{\theta}^{\mathrm{2SLS}} = \widehat{\theta}^{\mathrm{OLS-proj}}. \] In particular, the 2SLS coefficient on $X$ coincides with the OLS coefficient on $\tilde X$ in the regression of $Y$ on $(\tilde X,Z_1)$.
proofBy definition of $P_M$, we have \[ \tilde S := P_M S = \begin{bmatrix} P_M \mathbf{1} & P_M X & P_M Z_1 \end{bmatrix}. \] Because each column of $Z_1$ is in the column space of $M$, $P_M Z_1 = Z_1$. Similarly, if the constant is included in $Z_1$ (or in $M$), then $P_M\mathbf{1} = \mathbf{1}$. By definition, $P_M X = \tilde X$. Hence $\tilde S = [\mathbf{1},\,\tilde X,\,Z_1]$ as defined above. The 2SLS estimator can be written as \[ \widehat{\theta}^{\mathrm{2SLS}} = (S'P_M S)^{-1} S'P_M Y. \] Using $\tilde S = P_M S$, we have \[ S'P_M S = S'\tilde S = \tilde S'\tilde S, \qquad S'P_M Y = S'\tilde Y = \tilde S'Y, \] since $\tilde Y := P_M Y$ and $P_M$ is symmetric and idempotent. Therefore \[ \widehat{\theta}^{\mathrm{2SLS}} = (\tilde S'\tilde S)^{-1}\tilde S'Y = \widehat{\theta}^{\mathrm{OLS-proj}}, \] which proves the claim.
remark[Application to Q--LS] In the Q--LS setting, let $G$ collect the $K$ quantile-based instruments $\widehat g_k(Z_i) = Z_i'\widehat\pi(\tau_k)$ and let $M=[Z_1,G]$. The projected regressor $\tilde X = P_M X$ is exactly the best linear predictor of $X$ in the span of $(Z_1,G)$: \begin{equation} \tilde X_i = Z_{1i}'\hat\delta + G_i'\hat w, \end{equation} for some least-squares coefficients $(\hat\delta,\hat w)$. By Lemma (ref), the Q--LS 2SLS estimator that uses $(Z_1,G)$ as instruments is numerically equivalent to OLS of $Y$ on $(\tilde X,Z_1)$. In particular, if we define the Q--LS plug-in regressor as \begin{equation} \widehat X_i^{\mathrm{QLS}} := \tilde X_i = P_M X_i, \end{equation} then the plug-in estimator that regresses $Y$ on $(\widehat X^{\mathrm{QLS}},Z_1)$ is identical (in finite samples) to the corresponding Q--LS IV estimator.

General comparison and LATE interpretation

Outside of the pure location case, Q--LS remains consistent under distributional relevance even when mean relevance fails or the conditional mean is misspecified, because it exploits shifts in the entire distribution of $X\mid Z$. In designs where both mean relevance and distributional relevance hold and the efficient IV $h^*(Z)$ lies in the closure of the quantile-generated class $\mathcal{H}$, Q--LS again attains the same efficiency bound as optimal 2SLS. When $h^*(Z)\notin\mathcal{H}$, Q--LS is efficient within $\mathcal{H}$ but may be slightly less efficient than an oracle 2SLS that can use arbitrary functions of $Z$ as instruments. The main gain of Q--LS is robustness: it can remain strongly identified in settings where the classical first-stage mean shift is close to zero but the distribution of $X\mid Z$ varies substantially with $Z$.

When the endogenous regressor is (or can be reduced to) a binary treatment $X=D\in\{0,1\}$, Q--LS also admits a familiar LATE interpretation under additional structure. In the positive-normalized version, the Q--LS index $H = h(Z)$ is a convex combination of conditional quantiles of $D\mid Z$, which collapses possibly high-dimensional instruments into a single scalar encouragement variable. Under standard IV validity and a no-defiers monotonicity condition with respect to $H$, the 2SLS coefficient using $H$ as the sole instrument can be written as a positive-weighted average of local average treatment effects for units whose treatment status changes as $H$ increases. Thus, in binary-treatment applications, Q--LS can be viewed not only as a strength/efficiency device but also as a way to construct a scalar monotone instrument that delivers a conventional LATE interpretation (see Appendix (ref) for details).

In the binary-treatment case, the scalar Q--LS index also admits a LATE representation with nonnegative weights on adjacent “index-complier” groups (Appendix (ref)), and these weights can be estimated in practice, so the researcher can diagnose which parts of the instrument support contribute most to the Q--LS effect. In addition, because the Q--LS index is built as a convex combination of conditional quantiles, the estimated weight vector \(\hat{w}\) reveals which parts of the \(D\mid Z\) distribution (e.g.\ central vs.\ tail quantiles) contribute most to instrument strength.

Implementation and regularization

The population optimal Q--LS instrument $h_{\mathrm{opt}}$ solves the projection problem in Definition (ref). In the finite-grid implementation, this leads to the least-squares problem (ref) for the discrete weights $w\in\mathbb{R}^K$. When $K$ is moderate and the quantile-based instruments are highly collinear, the Gram matrix $\widehat{G}'\widehat{G}$ can be ill-conditioned, and the unregularized solution $\widehat{w}_n$ becomes unstable. This is the familiar ill-posedness of inverse problems and many-instrument IV.

We discuss two regularized versions of Q--LS that address this issue: ridge-regularized Q--LS and sparse (LASSO) Q--LS.

Ridge-regularized Q--LS

A simple and effective approach is Tikhonov (ridge) regularization. Instead of (ref), we solve

equation[equation omitted — 219 chars of source]

which yields the closed-form solution

equation[equation omitted — 152 chars of source]

The corresponding regularized Q--LS instrument is

equation[equation omitted — 160 chars of source]

Finite-sample stability is improved because the ridge term $n\lambda_n I_K$ stabilizes the inversion of $\widehat{G}'\widehat{G}$ and shrinks the weights toward zero. Under the same conditions as in Proposition (ref), and provided $\lambda_n\to 0$ at a suitable rate (for example, $\lambda_n\downarrow 0$ and $K_n^2/(n\lambda_n)\to 0$), the regularized weights $\widehat{w}_n^{\lambda}$ remain consistent for the population optimal weights and $\widehat{X}_i^{\mathrm{Q\text{--}LS},\lambda}$ converges in $L^2$ to $h_{\mathrm{opt}}(Z_i)$. The asymptotic distribution in Proposition (ref) therefore continues to hold with $\widehat{X}_i^{\mathrm{Q\text{--}LS},\lambda}$ replacing $\widehat{X}_i^{\mathrm{Q\text{--}LS}}$.

Two additional practical devices can be used in tandem with the ridge penalty. First, one may limit the number of quantile instruments $K$ to a moderate value (e.g.\ $K=5$ or $K=10$), or work with a low-dimensional principal-component representation of the columns of $\widehat{G}$. Second, instead of solving (ref) explicitly, one can treat each $\widehat{g}_k(Z_i)$ as a separate instrument, run 2SLS or LIML with the full instrument vector, and let the IV estimator choose an implicit linear combination of quantile-based instruments. These strategies reduce the effective instrument dimension and mitigate many-instrument bias while preserving the distributional information exploited by Q--LS.

Sparse Q--LS via LASSO

The ill-posedness of the optimal weight problem arises because the finite-grid instrument dictionary $\widehat{G}$ can be high-dimensional and highly collinear. An alternative to ridge regularization is to enforce sparsity in the weights and let the data select a subset of quantile-based instruments. This leads naturally to an $\ell_1$-penalized (LASSO) version of Q--LS.

Given the $K$ quantile-based instruments $\widehat{g}_k(Z_i) = Z_i'\widehat{\pi}_n(\tau_k)$, stacked in $\widehat{G}_i$ and $\widehat{G}$ as in (ref), we define the LASSO-regularized weights as

equation[equation omitted — 226 chars of source]

where $\lambda_n>0$ is a tuning parameter. The corresponding sparse Q--LS instrument is

equation[equation omitted — 234 chars of source]

By construction, only a subset of the quantile indexes has $\widehat{w}_{n,k}^{\mathrm{LASSO}}\neq 0$, so the effective instrument dimension is reduced and the many-instrument problem is mitigated. In practice, we also consider a post-LASSO refinement: letting $\widehat{S}_n = \{k: \widehat{w}_{n,k}^{\mathrm{LASSO}}\neq 0\}$ denote the selected support, we re-estimate the weights by least squares restricted to $\widehat{S}_n$ and use the resulting fitted values as $\widehat{X}_i^{\mathrm{Q\text{--}LS,L}}$.

The second-stage estimator is then defined as in (ref), with $\widehat{X}^{\mathrm{Q\text{--}LS}}$ replaced by $\widehat{X}^{\mathrm{Q\text{--}LS,L}}$. From the perspective of the structural parameter $\theta_0$, the first stage (quantile processes and weights) is a high-dimensional nuisance. The IV moment condition is Neyman-orthogonal with respect to this nuisance, so that, under standard sparsity and rate conditions for LASSO and a suitable choice of $\lambda_n$, the estimation error in $\widehat{X}_i^{\mathrm{Q\text{--}LS,L}}$ is $o_p(n^{-1/2})$ in the moment equations. As a result, the asymptotic normality result in Proposition (ref) continues to hold with $\widehat{X}_i^{\mathrm{Q\text{--}LS,L}}$ in place of $\widehat{X}_i^{\mathrm{Q\text{--}LS}}$, and the usual heteroskedasticity-robust 2SLS covariance matrix computed with the sparse Q--LS instrument remains valid for inference.

In empirical work, the tuning parameter $\lambda_n$ can be chosen either by cross-validation (targeting prediction of $X$) or by more conservative theoretical rules designed for high-dimensional linear regression. A moderate level of penalization, combined with post-LASSO refitting, tends to yield stable first stages and good finite-sample size for Q--LS inference.

Positive normalized Q--LS weights

Our baseline Q--LS instrument is a linear functional of the conditional quantile process, \[ h_\omega(Z)=\int_0^1 \omega(\tau)\,Q_{X\mid Z}(\tau\mid Z)\,d\tau, \] with $\omega\in L^2(0,1)$. In finite samples, we approximate this object using a quantile grid $0<\tau_1<\cdots<\tau_K<1$ and fitted quantiles $\hat g_k(Z)\approx \widehat Q_{X\mid Z}(\tau_k\mid Z)$, so that $h(Z)\approx \hat G(Z)'w$ for $\hat G(Z)=(\hat g_1(Z),\ldots,\hat g_K(Z))'$.

For interpretability and stability, we also consider a constrained version that enforces positive, normalized weights:

equation[equation omitted — 196 chars of source]

Equivalently, in population form, this corresponds to restricting $\omega(\tau)\ge 0$ and $\int_0^1\omega(\tau)\,d\tau=1$.

\paragraph{Interpretation.} Under these constraints, $h_{\mathrm{PN}}(Z)=\hat G(Z)'w_{\mathrm{PN}}$ is an L-statistic: it is a convex combination of conditional quantiles, i.e.\ a “quantile average” index. Positivity also preserves distributional ordering: if the conditional distribution of $X$ under $Z=z_2$ first-order stochastically dominates that under $Z=z_1$ (so that $Q_{X\mid Z=z_2}(\tau)\ge Q_{X\mid Z=z_1}(\tau)$ for all $\tau$), then $h_{\mathrm{PN}}(z_2)\ge h_{\mathrm{PN}}(z_1)$ for any admissible weight function. This rules out sign-cancellation across quantiles and yields an index whose movements track distributional shifts in a single direction.

\paragraph{Finite-sample stability.} Because extreme quantiles are noisier to estimate, unconstrained Q--LS may assign large positive and negative weights that amplify estimation error. The simplex constraint $w\ge 0$, $\mathbf{1}'w=1$ regularizes the instrument, bounding its variability and often improving robustness when $\hat G$ is generated (e.g.\ via quantile regression). In our empirical implementation we compute $w_{\mathrm{PN}}$ by quadratic programming, and we can combine it with cross-fitting to further limit overfitting of $\hat G$.

Practical recommendations

We view Q--LS as a complement to, rather than a replacement for, the standard 2SLS toolkit with weak instruments. In practice, applied work can proceed in three steps: (i) run conventional mean-based diagnostics, (ii) assess distributional relevance, and (iii) choose between 2SLS and Q--LS (and its regularised variants) based on a small menu of test statistics and simple stability checks.

\paragraph{Step 1: Start from the usual 2SLS diagnostics.}

itemize• Estimate the conventional first stage $X$ (endogenous variable) on $Z$ (instrument) and controls and report the usual weak-IV statistics (e.g. the cluster-robust Kleibergen--Paap $F$-statistic or the Stock--Yogo $F$ for the mean first stage). • If the mean first stage is clearly strong (e.g. robust $F \gtrsim 10$), treat linear 2SLS with the conventional instrument as your baseline and view Q--LS mainly as a robustness check. • If the mean first stage is weak (e.g. robust $F \ll 10$), proceed to distributional diagnostics rather than discarding the design.

\paragraph{Step 2: Diagnose distributional relevance.}

itemize• Run a small set of conditional quantile regressions of $X$ on $Z$ and controls (e.g. $\tau \in \{0.10,0.25,0.50,0.75,0.90\}$) and test the joint significance of $Z$ in each quantile regression. Large $F$-statistics at upper quantiles with small or insignificant mean effects are indicative of purely distributional relevance. • Graphically, compare the estimated conditional quantiles or empirical CDFs of $X$ across instrument values (e.g. pre/post policy). Pronounced tail shifts with little change in the median or mean point towards a distributional design where Q--LS is useful.

\paragraph{Step 3: Choosing $K$, regularisation, and inference.}

itemize• Quantile grid. Use a moderate grid of conditional quantiles, for example $K \in \{5,10,20\}$ equally spaced over $(0,1)$, possibly excluding extreme tails (e.g.\ $[\tau_{\min},1-\tau_{\min}]$ with $\tau_{\min} \in [0.02,0.05]$). In most empirical samples $K=10$ provides a good compromise between flexibility and stability. • Regularisation. If $K$ is large relative to $n$ or the quantile regressions are noisy, use ridge- or LASSO-regularised Q--LS (Q–LS--R, Q–LS--L1, Q–LS--L2) rather than the unpenalised version. Our simulations suggest that ridge (Q–LS--R) is a conservative default, while LASSO variants are helpful in very many-instrument designs. • Inference. Always report heteroskedasticity-robust (and, in panels, cluster-robust) standard errors based on the Q--LS sandwich variance (or its penalised analogue). Since Q--LS produces a conventional instrument $\hat h(Z)$, practitioners can also apply standard weak-IV-robust tools (e.g.\ Anderson--Rubin or conditional likelihood ratio tests) using $\hat h(Z)$ as the excluded instrument.

\paragraph{Step 4: Stability checks and reporting.}

itemize• Compare Q--LS and 2SLS. Report side-by-side Q--LS and classical 2SLS estimates and standard errors. In designs with strong mean relevance, point estimates and standard errors should be similar. In designs with weak mean relevance but strong distributional relevance, Q--LS is expected to deliver comparable point estimates with much tighter confidence intervals. • Check robustness to $(K,\lambda)$. Re-estimate Q--LS on a few nearby quantile grids (e.g.\ $K=5,10,20$) and, for regularised variants, a small set of tuning parameters $\lambda$. Large swings in the coefficient across neighbouring grids or penalties are a warning sign that the design may be too fragile for sharp causal interpretation. • Document instrument strength. Alongside the conventional first-stage $F$, report a “distributional $F$”: the joint significance of the quantile-generated instruments in the Q--LS first stage. This makes transparent what Q--LS is exploiting that 2SLS misses.

Table (ref) summarises these recommendations in a simple diagnosis chart.

table[table omitted — 1,621 chars of source]

Simulation study

This section will examine the finite-sample performance of Q--LS and its regularised variants in a set of Monte Carlo designs. We focus on three types of designs that map directly to our theoretical discussion. (A) strong mean relevance (benchmark Gaussian case), (B) Trigonometric and distributional instruments (variance shifts) and (C) mixed mean and distributional relevance.

The following parts of the data-generating processes (DGPs) are the same in all three designs;

itemize• The structural equation \begin{equation*} Y = \beta_0 + Z_1\beta_1 + X \beta_x + \varepsilon \end{equation*} where the parameter on the endogenous variable $X$, $\beta_x=1$ is the parameter of interest, and $\beta_0=1, \beta_1=1$ with $Z_1\sim N(0,1)$. • The errors between the structural equation ($\varepsilon$) and first-stage ($\nu$) follow the following process, \begin{align*} (\varepsilon,\nu)&\sim N\left(0, \begin{pmatrix} 1 & 0.6 \\ 0.6 & 1 \end{pmatrix}\right). \\ \end{align*} • We will consider sample sizes representative of typical applied work $n=500,1000$ and we run 1000 replications.

We compare 10 estimators:

enumerate• Classical IV where the first stage is a least squares projection with just first order terms i.e. \begin{align} X &= \zeta_0 + \zeta_1Z_1 +\zeta_2Z_2 + e_x \end{align} • Our Q–LS with all quantiles equally weighted, first stage is estimated as a second order polynomial, \begin{align} X &= \zeta_{\tau_{k,0}} + \zeta_{\tau_{k,1}}Z_1 +\zeta_{\tau_{k,2}}Z_2+ \zeta_{\tau_{k,3}}Z_1^2 +\zeta_{\tau_{k,4}}Z_2^2+ \zeta_{\tau_{k,5}}Z_1Z_2 + e_{x,\tau_k} \end{align} for a grid of quantiles $\{\tau_k\}$. The resulting fitted quantile regression functions are used to generate a prediction of the endogenous variable by an equally weighted average of the quantile regression fitted values, \begin{enumerate} • Estimate of $\beta_x$ is from analytical 2SLS analogue formulas (ref) denoted Q–LS-a. • Estimate of $\beta_x$ is from a plug-in OLS regression of the structural equation where the equally weighted average of quantile fitted values is used in place of $X$, denoted Q–LS. \end{enumerate} • Distributionally relevant control function, first stage is also estimated as a second order polynomial (ref) and the control function is estimates as in cfnsv20_cfa. \begin{align} y &= \beta_0 +\beta_1z_1+\beta_xx + \beta_v\hat{v} + e_y \end{align} where $\hat{v}\approx F_X(X|Z)$ is a control function and is defined in (ref). Estimate of $\beta_x$ is from (ref), denoted DR-IV. • Our Q–LS with all quantiles weighted by ridge regression coefficients, Same as estimator (2) but where the weights are ridge regression coefficients. \begin{enumerate} • Estimate of $\beta_x$ is from analytical 2--SLS analogue formulas (ref) denoted Q–LS-R-a. • Estimate of $\beta_x$ is from a plug-in OLS regression of the structural equation, where the ridge weighted average of quantile fitted values is used in place of $X$, denoted Q–LS-R. \end{enumerate} • Our Q–LS with all LASSO selected quantiles are weighted equally weighted (all other given zero weights), Same as estimator (2) but only LASSO selected quantile fitted values given non-zero weights. \begin{enumerate} • Estimate of $\beta_x$ is from analytical 2--SLS analogue formulas (ref) denoted Q–LS-L1-a. • Estimate of $\beta_x$ is from a plug-in OLS regression of the structural equation where the LASSO selected weighted average of quantile fitted values is used in place of $X$ denoted Q–LS-L1. \end{enumerate} • Our Q–LS with all quantiles weighted by LASSO regression coefficients, Same as estimator (2) but where the weights are LASSO regression coefficients. \begin{enumerate} • Estimate of $\beta_x$ is from analytical 2--SLS formulas in the main paper denoted Q–LS-L2-a. • Estimate of $\beta_x$ is from a plug-in OLS regression of the structural equation where the LASSO weighted average of quantile fitted values is used in place of $X$ denoted Q–LS-L2. \end{enumerate}

The weights for the weighted average of the quantiles are generated by sample splitting. The first 50% of the data is used to estimate the Ridge/LASSO regression regularisation parameters with 10 fold cross-validation predication accuracy. The estimated hyperparamter and the other 50% of the data is then used to estimated the Ridge/LASSO coefficient estimate.

For all Q–LS variants and DR-CF we consider quantiles between $v=0.01$ and $1-v$ and set the gap between the quantiles to $c=\{0.10,0.05,0.01,0.001\}$ this corresponds to using $K=\{10,20,99,981\}$ quantiles to generate the average of the quantile regression fitted values of $X$ or approximate the CDF ($\hat{v}$). The trimming of the quantiles follows cfnsv20_cfa.

Simulation results will be reported in terms of bias, root mean squared error (RMSE), and empirical 95% coverage of the simulation estimate.

Design A: Strong mean relevance (benchmark Gaussian case)

The first design is a “best-case” benchmark in which classical mean relevance holds and the conditional distribution of $X\mid Z$ belongs to a Gaussian location family. The data-generating process of the first stage is

align*[align* omitted — 84 chars of source]

with $\alpha_1=\alpha_2=1$.

In this design, both 2SLS and Q--LS are correctly specified and the conditions of Proposition (ref) hold. The simulation will document:

itemize• similarity of bias, RMSE and empirical coverage for Q--LS and 2--SLS; • sensitivity of Q--LS to the choice of quantile grid and regularisation in a setting where both estimators are efficient.

Design B: Trigonometric and distributional instruments

The second design studies a case where the excluded instrument $Z_2$ generates rich non-linear and distributional variation in $X$, while the linear first stage of $X$ on $Z_2$ can be arbitrarily weak. The data-generating process of the first stage is

align[align omitted — 144 chars of source]

with $\alpha_1 = 1$ and a variance parameter $\gamma = 0.5$.

Hence, conditional on $(Z_1,Z_2)$, \[ \mathbb{E}[X \mid Z_1,Z_2] = \alpha_1 Z_1 - \cos(\omega Z_2), \qquad \operatorname{Var}(X \mid Z_1,Z_2) = \exp(2\gamma Z_2). \] The excluded instrument $Z_2$ affects the conditional mean through the nonlinear term $\cos(\omega Z_2)$ and the conditional dispersion through $\exp(\gamma Z_2)$. The conditional distribution $F_{X\mid Z}(\cdot \mid Z_1,Z_2)$ is therefore strongly distributionally relevant in the sense of the main text.

Note that, since $Z_2 \sim \mathcal{N}(0,1)$ is symmetric around zero and $z \mapsto \cos(\omega z)$ is an even function, we have \[ Cov\big(Z_2, \cos(\omega Z_2)\big) = \mathbb{E}\big[-Z_2 \cos(\omega Z_2)\big] = 0 \] for all $\omega \ge 0$. Moreover, $\nu$ is independent of $Z_2$ and has mean zero, so $Cov(Z_2,\exp(\gamma Z_2)\nu)=0$. Thus \[ Cov(X,Z_2) = 0 \] in the population for every value of $\omega$. Classical linear IV based on the regression of $X$ on $Z_2$ therefore has an exactly zero linear first stage, even though $Z_2$ induces strong nonlinear and distributional variation in $X$.

\paragraph{Sub-designs B1--B3.} We consider three values of the frequency parameter $\omega$:

itemize• Design B1 ($\omega = 0$). \[ X = \alpha_1 Z_1 + \exp(\gamma Z_2)\,\nu. \] The conditional mean is linear in $Z_1$ and does not depend on $Z_2$, \[ \mathbb{E}[X\mid Z_1,Z_2] = \alpha_1 Z_1, \] so the excluded instrument $Z_2$ is purely distributional: it affects only the variance of $X$ through $\exp(\gamma Z_2)$, and $\mathbb{E}[X\mid Z_2]$ is constant. Classical 2SLS that treats $Z_2$ as a mean shifter is therefore weak, while Q--LS can exploit the variance shifts. • Design B2 ($\omega = 0.01$). \[ X = \alpha_1 Z_1 - \cos(0.01 Z_2) + \exp(\gamma Z_2)\,\nu. \] Here \[ \mathbb{E}[X\mid Z_2] = -\cos(0.01 Z_2) \] is non-constant but very flat. The linear projection of $X$ on $Z_2$ has a small slope and a low first-stage $F$-statistic, so classical 2SLS using $Z_2$ appears weak. Nevertheless, the true mean and the conditional distribution depend nontrivially on $Z_2$, so the optimal Q--LS instrument $h_{\mathrm{opt}}(Z)$ is non-degenerate and remains strongly relevant. • Design B3 ($\omega = 0.5$). \[ X = \alpha_1 Z_1 - \cos(0.5 Z_2) + \exp(\gamma Z_2)\,\nu. \] The conditional mean \[ \mathbb{E}[X\mid Z_2] = -\cos(0.5 Z_2) \] is highly nonlinear and oscillatory. As $\omega$ increases, the covariance $\mathrm{Cov}(X,Z_2)$, and hence the linear first-stage coefficient of $X$ on $Z_2$, becomes small even though the amplitude of $\cos(\omega Z_2)$ remains large. Thus the linear first stage for 2SLS can look weak or misleading, while the quantile-generated space used by Q--LS can approximate $\cos(0.5 Z_2)$ well, delivering a strong and stable first stage in the Q--LS sense.

Overall, Designs B1--B3 trace a progression from a purely distributional instrument (B1, no mean effect of $Z_2$) through an “almost mean-irrelevant” case (B2, very flat mean effect) to a design with strong nonlinear mean and variance effects (B3). This family highlights that classical first-stage diagnostics based on the linear regression of $X$ on $Z_2$ can substantially understate instrument strength, whereas Q--LS remains well identified whenever $X$ has a nonzero projection in the quantile-generated space.

Design C: Mixed mean and distributional relevance

The third design introduces both mean and variance effects:

itemize$Z_2\in\{0,1\}$ with $\mathbb{P}(Z_2=0)=\mathbb{P}(Z_2=1)=1/2$. • $X = \alpha_1 Z_1 + \alpha_2 Z_2 + \sigma(Z_2) \nu$, with $\sigma(0)=1$, $\sigma(1)=3$, so that $\mathbb{E}[X\mid Z_2]\neq 0$, and $\alpha_1=\alpha_2=1$.

Here, both mean relevance and distributional relevance hold, but the contribution of each can be tuned by varying $(\alpha_2,\sigma(0),\sigma(1))$. The simulation will:

itemize• study the performance of Q--LS and 2SLS as the mean effect $\alpha_2$ becomes small while variance shifts remain large; • assess whether Q--LS remains competitive with 2SLS when mean effects are strong, and dominates when mean effects are weak but distributional shifts are large.

Note as $Z_2$ is binary in this design we set $\zeta_{\tau_{k,4}}=0$ in (ref) to avoid a singular QR design matrix, for all Q--LS variants and DR-CF.

Simulation Results

table[table omitted — 3,711 chars of source]
table[table omitted — 3,712 chars of source]

Table (ref) presents the results for Design A and Design C for sample sizes of 500 and 1000.\footnote{Full results for Design A and Design C are available upon request.} It is important to note that classical IV is correctly specified, as $Z_2$ is strongly mean-relevant in both Design A and Design C. For Design A, we find that all Q-LS variants perform similarly to classical IV, regardless of the number of quantiles used. This is due to the quantile projection being similar to the mean projection. The analytical Q-LS variants provide better coverage (closer to 0.95) than the OLS plug-in standard errors, which tend to yield over-coverage, indicating conservative standard errors in this design. The control function version (DR-CF) performs worse in terms of bias, RMSE, and coverage than all Q-LS variants and classical IV, regardless of sample size. However, its performance does improve as the number of quantiles increases in this design.

For Design C, $Z_2$ is both strongly mean-relevant and variance-relevant. As mean relevance is less clean, the bias and RMSE of classical IV are ten and two times as large, respectively, as in Design A for a given sample size. Unlike Design A, the unregularized Q-LS variants exhibit substantial bias when 10 quantiles are used; however, this bias decreases as the number of quantiles and the sample size increase. When 981 quantiles are used, all Q-LS variants have the same or smaller bias than classical IV, and the regularized versions (plug-in and analytical) all yield good coverage. The performance of DR-CF relative to the other estimators and as the number of quantiles and the sample size increase, is similar to that in Design A.

Table (ref) shows the results for Design B with $\omega={0,0.5}$.\footnote{Full results for Design B are available upon request.} Interestingly, classical IV performs worse when $\omega$ is non-zero in terms of both bias and RMSE for a given sample size, whereas the bias and RMSE of the Q-LS variants decrease, except for the bias of Q–LS-L1, which marginally increases (in absolute value) as the number of quantiles increases. More quantiles are needed for better approximation because the distributional relevance of $Z_2$ becomes stronger as $\omega$ increases due to the $\cos$ transformation. Similar to Designs A and C, DR-CF performance improves as the number of quantiles increases. However, unlike Designs A and C, DR-CF has the smallest bias and RMSE when $\omega=0$ for a given sample size, and also the smallest bias and RMSE when $\omega=0.5$ and $n=1000$. Despite this, DR-CF coverage is the worst among all estimators considered for a given number of quantiles and sample size.

Overall, across all designs considered, we find that at least one variant of Q-LS with 99 or more quantiles performs as well as or better than classical IV, even in designs where the instrument $Z_2$ is strongly mean-relevant. When a large number of quantiles are used, analytical Q-LS empirical coverage tends to be close to, but slightly below, 0.95, while the plug-in variants tend to exhibit coverage above 0.95. When the instrument is not strongly mean-relevant, the plug-in point estimates tend to be more stable than their analytical counterparts. The DR-CF estimator performs better as the number of quantiles increases across all designs considered, but it requires a large number of quantiles and weak mean relevance to achieve lower bias and RMSE than all Q-LS variants.

In Design A and C (strong mean relevance), Q--LS essentially reproduces optimal 2SLS. This confirms the theoretical prediction that, in a pure location-shift setting with strong instruments, exploiting the full conditional quantile process does not alter first-order behavior. Design B1 illustrates the degenerate branch of Lemma (ref). Because the optimal Q--LS instrument $h_{\mathrm{opt}}(Z)$ is identically zero, the first stage for Q--LS collapses and the estimator fails in the usual weak-IV sense. We do not regard this as a pathology of Q--LS, but as a reminder that distributional relevance alone is not sufficient for a useful first stage; Assumption (ref) is needed to rule out such cases. The main action comes in Design B3, where instruments are weak in the mean but strong in the distribution. Here classical 2SLS exhibits substantial finite-sample distortions. In contrast, Q--LS estimators based on a moderate grid of quantiles (e.g.\ $K=10$) are well centered and display near-correct coverage across the same sample sizes. These patterns mirror the HRS application: when policy reforms primarily reshape tails and dispersion rather than means, exploiting distributional relevance through Q--LS restores a strong, stable first stage from the same underlying instrument.

Empirical application

We use the Health and Retirement Study (HRS) to study how financial exposure to out-of-pocket (OOP) medical spending risk affects mental health among older Americans. Our main outcome is an indicator for any depressive symptoms (CESD) among respondents aged 65 and older. The endogenous regressor is a scalar measure of annual OOP medical spending, and the key policy instrument is an indicator for the post--Medicare Part D period (year $\ge 2006$), during which subsidized prescription drug coverage expanded and catastrophic protection was introduced.

To mirror the simulation designs, we work with two complementary specifications. First, we treat OOP spending in nominal dollars and use a single post--Part D indicator as the excluded instrument. This “stress test” design generates a classical weak-IV environment when the risk index is summarized by mean OOP spending, and it provides a direct counterpart to the distributional simulation in Section (ref). Second, we express OOP spending in real 2015 dollars using the OECD CPI and show that the same policy indicator becomes a strong instrument in this scaling. The real-dollar specification serves as a well-identified benchmark in which Q–LS and conventional 2SLS can be expected to coincide when we omit a time trend. Once we add a linear trend in survey year, the point estimates remain very similar but Q–LS delivers somewhat tighter standard errors. Throughout, we trim the top 1% of the OOP distribution to reduce the influence of extreme outliers.

Descriptive evidence on depression and medical spending risk

figure[figure omitted — 565 chars of source]

\paragraph{Descriptive patterns.} Figure (ref) plots average out-of-pocket (OOP) medical spending in thousands of dollars (left axis) and the share of respondents with any depressive symptoms (right axis) among HRS respondents aged 65 and older from 2000 to 2014. Mean OOP spending rises from about \$3{,}000 in 2000 to a peak of roughly \$5{,}200 in the mid-2000s, then falls back toward \$4{,}000–\$4{,}300 in the late 2000s and early 2010s. Over the same period, the prevalence of depression declines fairly steadily: the share with any depression falls from around 19–20 percent in 2000 to roughly 11–12 percent by the early 2010s.

These raw trends are consistent with the idea that older households faced substantial OOP risk in the early 2000s, followed by some compression of spending risk---coinciding with the implementation of Medicare Part D in 2006---while mental health outcomes improved over time. However, the figure is purely descriptive: it does not control for changes in the composition of the 65+ population, secular improvements in treatment, or other contemporaneous policy changes. In the regression analysis below, we therefore focus on isolating the effect of changes in financial risk exposure, as measured by OOP spending, using instrumental‐variables methods. In addition, we trim the regression sample at the 99th percentile of realized annual OOP spending to remove a small number of extreme outliers; this has negligible impact on the aggregate time-series patterns in Figure (ref) but stabilizes the covariate–outcome relationships.

figure[figure omitted — 714 chars of source]

Figure (ref) plots the mean, median, and upper quantiles (75th, 90th, and 95th percentiles) of annual OOP medical spending for HRS respondents aged 65 and older between 2000 and 2020. The median and 75th percentile are remarkably stable over time, indicating that typical OOP spending changes little over this period. By contrast, the mean closely tracks movements in the upper tail: the 90th and 95th percentiles rise sharply in the early 2000s and then flatten or decline following the introduction of Medicare Part D in 2006. The resulting compression of the gap between the upper quantiles and the median after 2006 is consistent with Part D primarily reducing exposure to extreme spending realizations rather than shifting the center of the distribution. From the perspective of our identification strategy, this pattern illustrates why standard first-stage diagnostics based on mean shifts label the post–Part D indicator as a weak instrument for mean OOP risk, even though it is strongly relevant for the distribution of spending, especially in the right tail.

Nominal OOP and a weak-IV design

We begin with a deliberately “stressful” specification that uses nominal OOP medical spending (in thousands of dollars) as the endogenous regressor and a single post–Part D indicator as the excluded instrument. This design closely mirrors the weak-IV, heavy-tailed environments in our Monte Carlo exercises and is useful for showing how Q--LS behaves when the instrument is weak for the mean but strongly relevant for the distribution of OOP spending.

Table (ref) reports linear probability models relating an indicator for any depressive symptoms to mean OOP medical spending. The sample consists of 16{,}313 person–year observations for respondents aged 65 and older in the RAND HRS panel between 2000 and 2020, after excluding person–years with OOP spending above the 99th percentile of the nominal OOP distribution.

table[table omitted — 2,189 chars of source]

In column (1), higher nominal OOP spending is positively and statistically significantly associated with depression: a \$1{,}000 increase in annual OOP spending is associated with a 0.2 percentage point increase in the probability of reporting depressive symptoms ($\hat\beta = 0.002$, s.e.\ 0.001). The 2SLS specification in column (2), which instruments mean OOP spending with post–Part D, yields a much larger point estimate ($\hat\beta = 0.442$) but an even larger standard error (s.e.\ 0.410). The IV estimate is therefore imprecise and not statistically different from zero at conventional levels.

Table (ref) shows that this lack of precision reflects the fact that the post–Part D indicator is a very weak instrument for mean nominal OOP spending in the trimmed sample: the cluster-robust partial $F$-statistic for the excluded instrument is 1.19, far below usual thresholds.

table[table omitted — 1,922 chars of source]

\paragraph{Q--LS vs 2SLS in the weak nominal design.} We now use this weak-IV nominal specification as a stress test to compare Q--LS and linear 2SLS. Our goal here is diagnostic: to document how the two estimators behave when the instrument is weak for the mean, yet strongly relevant for the distribution of OOP spending.

Table (ref) compares the Q--LS estimate based on the projected OOP risk index with the conventional linear 2SLS estimate that uses mean nominal OOP spending as the endogenous regressor.

table[table omitted — 1,868 chars of source]

The point estimates in columns (1) and (2) are numerically identical (0.442), but the standard errors differ dramatically. In the weak-IV nominal specification, Q--LS delivers a standard error of 0.071 and a precisely estimated positive effect, whereas linear 2SLS yields a standard error of 0.410 and a statistically insignificant estimate. This pattern mirrors our simulation results: when the policy instrument reshapes the distribution of OOP spending but produces only a noisy mean shift, Q--LS can extract a much stronger effective first stage from the same variation, tightening confidence intervals without changing the underlying signal.

In this nominal specification, the combination of weak mean relevance and heavy tails also makes the estimated magnitude difficult to interpret substantively. In the next subsection, we therefore turn to a more natural specification that expresses OOP spending in real terms.

Real OOP: from strong-IV benchmark to trend-adjusted design

For substantive interpretation, it is more natural to express OOP spending in real dollars. In the second part of the empirical section we therefore: (i) redefine OOP spending in thousands of 2015 U.S.\ dollars, and (ii) show how the Part D instrument looks first like a strong mean instrument, and then, after soaking up secular trends, like a primarily distributional instrument where Q--LS mainly improves precision.

We deflate nominal OOP using the OECD consumer price index obtained from FRED, constructing a monthly index with base year 2015 and aggregating to the relevant survey years. Real OOP spending in thousands of 2015 dollars is denoted $oop\_med\_2015\_k$.\footnote{Organization for Economic Co-operation and Development, Consumer Price Indices (CPIs, HICPs), COICOP 1999: Consumer Price Index: Total for United States, retrieved from FRED, Federal Reserve Bank of St. Louis; \url{https://fred.stlouisfed.org/series/USACPIALLMINMEI}, accessed January 14, 2026.}

\paragraph{Real OOP without time trend: a strong-IV benchmark.} In the first real-dollar specification, we repeat the nominal regressions replacing $oop\_med\_k$ with $oop\_med\_2015\_k$ and keeping the same controls. In this case, the post–Part D indicator is a strong instrument for real OOP: the first-stage partial $F$-statistic is around 50 in the trimmed sample. Table (ref) compares Q--LS and linear 2SLS for the depression outcome in this strong-IV environment.

table[table omitted — 1,821 chars of source]

In this specification, Q--LS and 2SLS are virtually indistinguishable: both yield an effect of about 0.045 per \$1{,}000 of real OOP, with very similar standard errors. This is exactly the “well behaved” case in which our theory predicts that Q--LS collapses to optimal 2SLS when the instrument is strong and the linear mean representation is adequate. Table (ref) , in Appendix, confirms that three different ways of evaluating uncertainty for the Q--LS coefficient---cluster-robust OLS, clustered bootstrap, and the analytical Q--LS sandwich formula---deliver nearly identical standard errors. \paragraph{Real OOP with time trend: a distributional relevance design.} Finally, we add a flexible linear time trend to the real-OOP specification. This absorbs much of the secular drift in spending and returns us to a situation in which post--Part D induces rich distributional changes in real OOP but only a small additional mean shift conditional on time trend: with first-stage partial $F$-statistic is around 2. From the perspective of classical diagnostics, the instrument again looks fragile; from the perspective of distributional relevance, it remains informative.

Table (ref) reports Q--LS and 2SLS estimates for this trend-adjusted real specification.

table[table omitted — 2,027 chars of source]

In this trend-adjusted real specification, Q--LS and 2SLS again deliver the same point estimate (0.07 per \$1{,}000 of real OOP). The main difference is precision: the Q--LS standard error is 0.05, whereas the 2SLS standard error is 0.07. Thus Q--LS reduces the confidence interval width by roughly one-third, but the estimate remains only marginally significant at best (p-value around 0.14). Economically, the estimates across the real-dollar specifications suggest a modest positive effect of OOP risk on depression, with the strongest evidence in the specification that treats Part D as a strong mean instrument and somewhat weaker evidence once secular trends are absorbed.

Taken together, the three designs -- nominal weak-IV, real strong-IV, and real--with-trend -illustrate the two key empirical messages of the paper. First, the way in which a policy reform is summarized in the first stage (nominal vs real, with or without trend) can dramatically change whether the design looks weak or strong in mean terms, even when the underlying distributional shifts are similar. Second, when mean relevance is fragile but distributional relevance can be strong, Q--LS behaves like a natural complement to 2SLS: it recovers the same underlying signal where the design is well identified, and it delivers substantial precision gains relative to linear IV in settings where standard weak-IV diagnostics based on mean shifts understate the information content of the instrument.

Concluding remarks

Our results can be summarized along two dimensions. On the identification side, we formalize a notion of distributional relevance and show that even purely distributional instruments---for which $\operatorname{Var}(\mathbb{E}[X\mid Z]) = 0$ while $F_{X\mid Z}(\cdot\mid Z)$ varies nontrivially in $Z$---are sufficient to identify average structural functions in a nonseparable triangular model via the control function $V = F_{X\mid Z}(X\mid Z)$. This makes explicit how instruments that look weak in the mean can still be strong in the distribution, and connects the control–function identification argument to weak-IV concerns in linear models.

On the estimation side, we propose a simple, implementable estimator, Quantile Least Squares (Q--LS), which constructs an optimal distributional instrument from conditional quantiles and then uses it in a conventional linear IV step. In designs where policy reforms are primarily meant to reshape the distribution of risk rather than to move its mean, Q--LS complements the existing 2SLS and weak-IV toolkits by extracting a stronger effective first stage from the same policy variation, while leaving classical strong-IV applications essentially unchanged.

Q--LS exploits this structure by choosing, within a given dictionary of quantile-based functions, the instrument that best predicts $X$ in mean square error and then using it in a standard 2SLS procedure. We showed that Q--LS is consistent and asymptotically normal under distributional relevance, that it coincides with optimal 2SLS in Gaussian/location settings where mean relevance holds, and that it remains reliable when classical first-stage diagnostics indicate weak instruments from the perspective of the conditional mean. We also developed ridge and LASSO regularisation schemes that stabilise the first stage and mitigate many-instrument bias, while preserving the first-order asymptotics of Q--LS and providing practically useful variants (Q–LS--R, Q–LS--L1, Q–LS--L2) with distinct robustness properties.

Our Monte Carlo evidence highlights three main messages. First, when the instrument primarily shifts the tails of the $X$-distribution, Q--LS can recover the same limiting parameter as optimal 2SLS but with substantially smaller finite-sample variance, even in designs where the mean first stage is formally weak. Second, the regularised Q--LS variants are effective at controlling the dispersion of the estimator in many-instrument designs and under moderate misspecification of the quantile representation. Third, conventional IV estimators can behave erratically in these designs, with large RMSE and unstable confidence intervals, even when coverage remains nominal, whereas Q--LS delivers more concentrated sampling distributions.

The empirical application to Medicare Part D illustrates these features in a setting where policy reforms plausibly operate through the distribution of medical spending risk rather than its mean. In the nominal-dollar specifications, standard first-stage diagnostics classify the post--Part D indicator as a very weak instrument for mean OOP spending, and conventional 2SLS delivers large but imprecise estimates of the effect of OOP risk on depression. By contrast, Q--LS exploits the pronounced compression of the upper tail of the OOP distribution to construct a strong distributional first stage from the same policy variation, yielding substantially tighter confidence intervals while preserving the economic interpretation of the effect. When we deflate spending to 2015 dollars, the mean first stage becomes much stronger and conventional 2SLS performs well; in that case Q--LS and 2SLS again agree closely on the point estimates, but Q--LS offers at most modest efficiency gains. Taken together, the simulations and the Medicare Part D evidence suggest viewing Q--LS as a complement to standard IV: it nests the optimal mean-based instrument in classical designs, but remains informative in applications where instruments are weak in the mean and strong in the distribution.

The distributional relevance framework extends beyond linear IV. In an appendix, we show that purely distributional instruments can identify average structural functions in nonseparable triangular models via a control-function representation based on the conditional distribution $F_X(X\mid Z)$. This connects our approach to the literature on nonparametric IV and completeness, and suggests several directions for future work: deriving more primitive conditions that guarantee nondegenerate optimal instruments; developing formal diagnostics for the linear quantile representation and for instrument-instability; designing data-driven procedures for selecting quantile grids and penalties; and applying Q--LS in empirical settings where policy or institutional variation plausibly operates through changes in risk exposure rather than simple mean shifts.