Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
86,677 characters · 23 sections · 83 citation commands
Representativeness and Efficiency in Overidentified IV
{ \singlespacing
\noindentKeywords: Overidentified IV, heterogeneous treatment effects, GMM, LATE, semiparametric efficiency
\noindentJEL Classification: C14, C26, C36 }
In the classical linear model, the estimand remains the same whether the estimator is efficient or not; the Gauss-Markov theorem guarantees that ordinary least squares minimizes variance without altering the underlying parameter. Under heterogeneous treatment effects, however, this logic breaks down for instrumental variables. When a researcher has multiple instruments for a binary treatment, the Generalized Method of Moments (GMM) weighting matrix determines whose treatment effects the estimand represents, not merely how precisely they are estimated. Every researcher running two-stage least squares with multiple instruments is choosing a specific GMM weighting matrix, a choice that dictates exactly which subpopulations' treatment effects the estimate recovers.
In this paper, we give concrete causal content to the “estimator determines estimand” phenomenon HallInoue2003, AndrewsChen2025 and develop a constructive solution to the trade-off between statistical efficiency and causal interpretability. We first characterize the implicit weights assigned to Wald estimands by any GMM weighting matrix. We demonstrate that Efficient GMM (EGMM) embeds a heterogeneity penalty that actively downweights instruments with high treatment effect dispersion, exacerbating the negative-weight problem of Two-Stage Least Squares (2SLS) identified by MogstadTorgovitskyWalters2021. Furthermore, we prove an impossibility result: unless all instrument-specific Wald estimands perfectly coincide, no weighting matrix can simultaneously achieve the semiparametric efficiency bound and deliver researcher-specified weights. To resolve this tension, we introduce the Representative Targeting (RT) estimator. By computing each Wald ratio separately and averaging them with researcher-specified weights, RT sidesteps GMM's common-residual architecture. Relying on the observable property of Positive Regression Dependence (PRD) to guarantee non-negative weights, RT restores causal interpretability and achieves the semiparametric efficiency bound for its specific target at a known variance cost.
Our analysis rests on compliance types, the natural generalization of AngristImbensRubin1996 to $L$ binary instruments, where each compliance type records how an individual would respond to every possible instrument configuration. The framework of ImbensAngrist1994, applied one instrument at a time, gives $L$ Wald estimands, each a ratio of reduced-form to first-stage effects. Under joint independence and monotonicity, each Wald estimand is a positive weighted sum of compliance-type-specific average treatment effects. However, if the instruments are correlated, these weights can become negative. This happens because the weights depend on the responsiveness of a type's treatment take-up to instrument $\ell$, and shifting the value of one instrument inherently alters the distribution of the others.
We show that PRD, a condition on the joint distribution of the instruments introduced by Lehmann1966, eliminates negative weights: each Wald estimand becomes a genuine convex combination of type-specific treatment effects (Proposition (ref)). PRD is a condition on the instruments, not on potential outcomes, and it holds for independent instruments (multi-site randomized trials), cumulative threshold instruments (examiner and judge leniency designs), and more generally instruments that are nondecreasing functions of a common source EsaryProschanWalkup1967. The question is how GMM combines these $L$ causally interpretable building blocks under PRD.
EGMM inverts the moment-condition second moment matrix Hansen1982, embedding a heterogeneity penalty that downweights instruments with high residual variance in the common-residual fit (Proposition (ref)). Crucially, this penalty amplifies the negative-weight problem documented for 2SLS by MogstadTorgovitskyWalters2021. This heterogeneity also provides a natural compliance-type interpretation for the Hansen1982 $J$-statistic. Under maintained instrument validity, rejection of the overidentifying restrictions signals treatment effect heterogeneity across compliance types, not instrument invalidity (Proposition (ref)). As AndrewsChen2025 demonstrate, the $J$-statistic asymptotically characterizes the range of estimates achievable across weighting matrices at a given standard error relative to the efficient estimator.
The Tennessee STAR class-size experiment makes these consequences concrete, with Figure (ref) illustrating the presence of substantial treatment effect heterogeneity. The instruments are individual schools with independent within-school randomization, yet the J-statistic rejects decisively ($p<0.001$; Table (ref)). Because instrument validity is guaranteed by design, this rejection points to differing Wald estimands across schools rather than invalid instruments. Each school shifts a distinct, non-overlapping set of compliers into small classes; thus, the rejection implies that these school-specific subpopulations experience varying average treatment effects, preventing any single parameter from satisfying all moment conditions simultaneously.
This heterogeneity makes the choice of weighting matrix consequential for the estimand itself, not just for precision. EGMM and 2SLS yield estimates of 6.55 and 8.84, respectively, for the effect of small classes (Table (ref)). This gap reflects fundamentally different estimands: 2SLS weights compliance types by their first-stage contributions, while EGMM embeds a heterogeneity penalty that downweights schools whose Wald estimands diverge from the common-residual fit (Proposition (ref)); the two estimators target distinct points along the interval of achievable estimates characterized by AndrewsChen2025, recovering treatment effects for different weighted averages of compliance-type-specific effects.
Researchers interested in a specific estimand (e.g., an equally weighted average) can choose a weighting matrix that delivers the desired weights directly. GMM accommodates this: for any set of target weights, there exists a weighting matrix that delivers them (Lemma (ref)), and every such matrix produces the same asymptotic variance (Proposition (ref)). However, this variance relies on a misspecified common residual. Consequently, unless all Wald estimands coincide, no weighting matrix can simultaneously deliver the target weights and achieve the efficiency bound (Proposition (ref)). In short, GMM can target any parameter the researcher desires; it just cannot do so efficiently.
We propose Representative Targeting (RT), a semiparametrically efficient estimator for a weighted average of Wald estimands. Rather than relying on a shared common residual, RT computes each instrument-specific Wald ratio separately and averages them using researcher-specified weights. Under PRD, any such average is guaranteed to be a convex combination of type-specific treatment effects (Proposition (ref)). Crucially, the RT estimator achieves the semiparametric efficiency bound for its targeted estimand (Proposition (ref)): no estimator, GMM or otherwise, can do better. Furthermore, its variance is a closed-form quadratic in the target weights (Proposition (ref)), computable from pilot estimates before committing to a specification. Two targets have natural economic content: the complier-share-weighted ATE (CSW-ATE), which weights each Wald estimand by the size of its first stage, and the equal-weight ATE (EW-ATE), which gives each instrument equal influence.
Under the latent index model (Section (ref)), compliance types map to intervals of latent resistance, and the weights for estimators like 2SLS, EGMM, CSW-ATE, and EW-ATE become integrals over the Marginal Treatment Effect (MTE) curve HeckmanVytlacil2005 (Proposition (ref)). By leveraging this continuous representation, RT can also target the Policy-Relevant Treatment Effect (PRTE),\footnote{The PRTE captures the average treatment effect of a counterfactual policy that shifts the distribution of the instrument from $F_0(\mathbf{z})$ to $F_1(\mathbf{z})$ HeckmanVytlacil2005, CarneiroHeckmanVytlacil2011. In the patent application below, the counterfactual is a uniform relaxation of examiner scrutiny in which each leniency group adopts the approval rate of the next-higher group; see Section (ref).} or its closest feasible approximation, at a known cost (Proposition (ref)).\footnote{While the latent index model is necessary to connect RT to continuous MTE curves and target the PRTE, foundational results, such as recovering a convex combination of treatment effects without negative weights, rely on primitive, observable features of the instruments such as PRD.}
Figure (ref) illustrates this representation using a patent examiner leniency design FarreMensaHegdeLjungqvist2020. Plotting the composite MTE weight function $\bar{h}(u;\omega)$ for each estimator reveals how the choice of weighting matrix dictates which regions of latent resistance receive weight and thereby determines the estimand. EGMM is visibly attenuated at the high-resistance margin, where the heterogeneity penalty hollows out weight; 2SLS spreads weight more broadly but without researcher control; and the PRTE-targeted RT nearly replicates the policy weight function (a relative $L^2$ error $=0.12\%$). Under treatment effect heterogeneity, the estimator is the estimand, and RT provides a practical, variance-optimal tool for actively choosing among them.
Hansen1982 derives the asymptotic properties of GMM estimators and the overidentification test. HallInoue2003 establish the theory of GMM under misspecification, characterizing the pseudo-true value when moment conditions are inconsistent. Under treatment effect heterogeneity, the moment conditions from different instruments target different Wald estimands; the “misspecification” is not a failure of identification but a consequence of heterogeneous effects. Building on this, AndrewsChen2025 show that under misspecification the choice of estimator implicitly determines the estimand. This paper advances their framework in three ways: the heterogeneity penalty (Proposition (ref)) pins down the precise mechanism driving this phenomenon; the impossibility result (Proposition (ref)) proves that estimand distortion is generically unavoidable within the GMM class; because of this, Representative Targeting (RT) provides a constructive solution by leaving the GMM class entirely, a constructive remedy not developed within their framework.
ImbensAngrist1994 establish LATE identification for a single instrument. Extending this to multiple instruments, MogstadTorgovitskyWalters2021 show 2SLS weights can be negative under partial monotonicity when instruments correlate; PRD generalizes their two-instrument sufficient condition for positive weights to an arbitrary number of instruments (Proposition (ref)). BlandholBonneyMogstadTorgovitsky2022 establish that covariates and a monotonicity-correct first stage are jointly necessary for 2SLS to remain weakly causal. Relatedly, GoldsmithPinkhamSorkinSwift2020 decompose the Bartik 2SLS estimator into Rotemberg-weighted just-identified IV estimates, providing diagnostics for assessing sensitivity and detecting problematic negative weights. While Kolesar2013 shows LIML can escape the convex hull of LATEs, our analysis characterizes the full GMM class, and like PoirierSloczynski2025, we ask what alignment costs when interpreting weighted estimands.
Finally, our work builds on the marginal treatment effect (MTE) framework MogstadTorgovitsky2024. HeckmanVytlacil2005 formalize the MTE, PRTE, and the IV weight representation, which HeckmanUrzuaVytlacil2006 apply to compare estimators. The latent index model Vytlacil2002 connects our compliance types to intervals of latent resistance, while RT operationalizes the structure of discrete instruments identifying discrete MTE values BrinchMogstadWiswall2017. For evaluating policies, MogstadSantosTorgovitsky2018 provide partial identification for the exact PRTE by bounding it via linear programming over marginal treatment response functions. In contrast, our proposed RT estimator point-identifies a surrogate: the $L ^2$-closest approximation of the PRTE within the span of instrument-specific weight functions (Proposition (ref)). These approaches are natural complements: where our method delivers a minimum-variance point estimate for this optimal approximation, their bounds deliver partial identification of the exact target.
Section (ref) sets up the framework: compliance types, the Wald decomposition, and PRD. Sections (ref)--(ref) develop the GMM weight characterization, the impossibility result, and RT. Section (ref) develops the MTE representation and PRTE targeting. Applications are in Section (ref); proofs in Supplemental Appendix (ref).
We observe a random sample $\{(Y_i, D_i, Z_{1i}, \ldots, Z_{Li})\}_{i=1}^n$. The outcome $Y_i$ is scalar, the treatment decision $D_i \in \{0,1\}$ is binary, and $Z_{1i}, \ldots, Z_{Li}$ are $L \geq 2$ binary instruments with $Z_{\ell i} \in \{0,1\}$ for each $\ell = 1, \ldots, L$. Each individual has potential outcomes $\{Y_i(d): d \in \{0,1\}\}$ and a potential treatment function $D_i(\cdot) : \{0,1\}^L \to \{0,1\}$, called the individual's compliance type, that records which instrument configurations move that individual into treatment.\footnote{The compliance type generalizes the complier/always-taker/never-taker classification of AngristImbensRubin1996 to $L$ instruments and is a special case of the treatment response type in BaiHuangMoonSantosShaikhVytlacil2024, who develop inference for treatment effects conditional on generalized principal strata under weaker monotonicity conditions and for nonbinary treatments.} The set of all compliance types is $\mathcal{T}$, each occurring with probability $\theta_t = P(D_i(\cdot) = t)$ and carrying a type-specific average treatment effect $\mathrm{LATE}_t \equiv \mathbb{E}[Y_i(1) - Y_i(0) \mid D_i(\cdot) = t]$. The observed outcome is $Y_i = D_i \, Y_i(1) + (1 - D_i) \, Y_i(0)$. For each instrument $\ell$, define the instrument probability $p_\ell = P(Z_{\ell i} = 1)$, first-stage coefficient $\pi_\ell = \mathbb{E}[D_i \mid Z_{\ell i} = 1] - \mathbb{E}[D_i \mid Z_{\ell i} = 0]$, reduced-form coefficient $\rho_\ell = \mathbb{E}[Y_i \mid Z_{\ell i} = 1] - \mathbb{E}[Y_i \mid Z_{\ell i} = 0]$, and Wald estimand $\mathrm{Wald}_\ell \equiv \rho_\ell / \pi_\ell$. We also denote $\mathbf{Z}_i = (Z_{1i}, \ldots, Z_{Li})'$.
Under Assumption (ref), $\mathcal{T}$ consists of all nondecreasing functions $t : \{0,1\}^L \to \{0,1\}$.\footnote{For $L = 2$, there are six monotone types: never-takers, always-takers, and four complier types ($Z_1$-only, $Z_2$-only, joint, either). Joint compliers need both instruments “on”; either-compliers respond to whichever instrument is “on” first.} The complier group for instrument $\ell$, \[ \mathcal{C}_\ell = \{i : D_i(1, z_{-\ell}) > D_i(0, z_{-\ell}) \text{ for some } z_{-\ell}\}, \] can contain multiple types. The complier-group average treatment effect for instrument $\ell$ averages $\mathrm{LATE}_t$ over all types in $\mathcal{C}_\ell$, weighted by type probabilities. When compliance groups do not overlap ($\mathcal{C}_\ell$ contains a single type for each $\ell$), the complier-group average equals $\mathrm{LATE}_t$ for that type.\footnote{BaiHuangMoonSantosShaikhVytlacil2024 provide confidence regions for treatment effects conditional on individual strata with uniform size control.}
For each instrument $\ell$, the Wald estimand gives rise to the moment condition
The use of centered instruments $Z_{\ell i} - p_\ell$ is equivalent to including a constant in the instrument matrix and concentrating out the intercept.\footnote{Concentrating out the intercept induces a transformed weighting matrix for the remaining moments via a Schur complement, and we formulate the analysis directly in the centered moment system with the effective weighting matrix. } Setting $g_\ell(\beta) = 0$ yields $\beta = \operatorname{Cov}(Y_i, Z_{\ell i}) / \operatorname{Cov}(D_i, Z_{\ell i}) = \rho_\ell / \pi_\ell = \mathrm{Wald}_\ell$. The second moment matrix of the stacked moment vector $g_i(\beta) = ((Y_i - \beta D_i)(Z_{1i} - p_1), \ldots, (Y_i - \beta D_i)(Z_{Li} - p_L))'$ is $\Omega(\beta) = \mathbb{E}[g_i(\beta)\,g_i(\beta)']$, which depends on $\beta$ through the residuals $Y_i - \beta D_i$.\footnote{Under misspecification, $\mathbb{E}[g(\beta^*)] \neq 0$ and hence the second moment matrix $\Omega(\beta^*)$ differs from $\operatorname{Var}(g_i(\beta^*))$ by the rank-one term $\mathbb{E}[g_i(\beta^*)]\mathbb{E}[g_i(\beta^*)]' \neq 0$. However, the population first-order condition $\boldsymbol{\gamma}'W\,\mathbb{E}[g(\beta^*(W))] = 0$ annihilates this term and thus the weight formula and sandwich variance are identical under either definition.}
When all Wald estimands coincide ($\mathrm{Wald}_1 = \cdots = \mathrm{Wald}_L = \beta$), all $L$ moment conditions hold simultaneously. Under treatment effect heterogeneity, different instruments shift treatment for different compliance types, and the Wald estimands are distinct. The moment conditions are then misspecified in the sense of HallInoue2003: no single $\beta$ satisfies all of them.\footnote{AndrewsBarahonaGentzkow2025 develop a general framework for structural estimation under misspecification; the IV setting is a special case where the constant-treatment-effect restriction is the misspecified model.} The GMM estimand under weighting matrix $W$ is the pseudo-true value $\beta^*(W) = \arg\min_b \, \mathbb{E}[g(b)]' W \, \mathbb{E}[g(b)]$, and the choice of $W$ determines which weighted average of Wald estimands $\beta^*(W)$ represents AndrewsChen2025.
The weight $\alpha_t(\ell)$ on type $t$ is proportional to $\theta_t \cdot \varphi_t(\ell)$: the probability of being type $t$, multiplied by how much instrument $\ell$ shifts treatment for that type.\footnote{MogstadTorgovitskyWalters2021 decompose the combined 2SLS estimand into complier-group-specific treatment effects. Proposition (ref) operates at a finer level, decomposing each instrument-specific Wald estimand separately into compliance-type-specific contributions.} The type contribution $\varphi_t(\ell)$ compares type $t$'s expected treatment under $Z_\ell = 1$ versus $Z_\ell = 0$, averaging over the other instruments. The weights can be negative when instruments are not independent, as decomposed by Lemma (ref).
The direct component $\varphi_t^D(\ell) \geq 0$ by monotonicity: switching $Z_\ell$ from 0 to 1 can only increase treatment for type $t$, holding $Z_{-\ell}$ fixed. The indirect component $\varphi_t^I(\ell)$ depends on how conditioning on $Z_\ell = 1$ shifts the distribution of $Z_{-\ell}$. Under independent instruments, $q_\ell = q^0_\ell$, the indirect component vanishes, and all weights are non-negative. Under positive dependence, conditioning on $Z_\ell = 1$ shifts $Z_{-\ell}$ upward ($q_\ell(z_{-\ell}) \geq q^0_\ell(z_{-\ell})$ for large $z_{-\ell}$) and all weights remain non-negative. Under negative dependence, $Z_\ell = 1$ shifts $Z_{-\ell}$ downward, and the indirect component $\varphi_t^I(\ell) < 0$ can overwhelm the direct component, producing negative Wald weights. Whether the weights are non-negative is therefore a design question, determined entirely by the instrument dependence structure (Appendix (ref) provides a concrete example with negative weights under negatively correlated instruments).
PRD is the corresponding positive-weight condition for overidentified IV settings\footnote{For $L = 2$, PRD reduces to non-negative covariance, recovering the sufficient condition in MogstadTorgovitskyWalters2021; for $L \geq 3$, pairwise non-negative covariance no longer suffices. GoldsmithPinkhamSorkinSwift2020 require a same-sign first-stage condition and an exclusion restriction that jointly give each just-identified IV estimate a convex-combination interpretation in the shift-share setting; HahnKuersteinerSantosWilligrod2024 require share orthogonality or non-negatively correlated shocks.}, and is a primitive, observable features of the instruments themselves. The following designs imply PRD:
Conversely, designing an experiment with mutually exclusive treatment arms (e.g., receiving Subsidy A or B, but never both) negatively correlates the instruments, which violates PRD. This negative indirect effect can overwhelm the positive direct effect, mechanically producing negative weights.
PRD is a testable design feature that can be assessed directly from the observable instrument distribution. Researchers could verify whether the empirical conditional distribution of $Z_{-\ell}$ given $Z_\ell = 1$ first-order stochastically dominates the distribution given $Z_\ell = 0$, which for $L = 2$ reduces to $P(Z_k = 1 \mid Z_\ell = 1) \geq P(Z_k = 1 \mid Z_\ell = 0)$.
The GMM estimator with weighting matrix $W$ solves
where $g_n(\beta) = (g_{n,1}(\beta), \ldots, g_{n,L}(\beta))'$ with $g_{n,\ell}(\beta) = n^{-1} \sum_{i=1}^n (Y_i - \beta D_i)(Z_{\ell i} - \hat{p}_\ell)$ and $\hat{p}_\ell = n^{-1}\sum_{i=1}^n Z_{\ell i}$.
Since the model is linear in $\beta$, $g_n(\beta) = g_n(0) - \beta \hat{\boldsymbol{\gamma}}$, where $\hat{\boldsymbol{\gamma}} = (\hat{\gamma}_1, \ldots, \hat{\gamma}_L)'$ with $\hat{\gamma}_\ell = n^{-1} \sum_{i=1}^n D_i (Z_{\ell i} - \hat{p}_\ell)$. The first-order condition yields
Let $\gamma_\ell = \operatorname{Cov}(D_i, Z_{\ell i}) = \pi_\ell p_\ell(1-p_\ell)$ denote the population first-stage covariance and $\boldsymbol{\gamma} = (\gamma_1, \ldots, \gamma_L)'$. Under Assumptions (ref)--(ref), $\hat{\gamma}_\ell \xrightarrow{p} \gamma_\ell > 0$.
By the Wald decomposition (Proposition (ref)), the GMM estimand is also a weighted average of type-specific treatment effects: $\beta^*(W) = \sum_{t \in \mathcal{T}} \mathrm{LATE}_t \cdot \psi_t(W)$, where $\psi_t(W) = \sum_{\ell=1}^L \lambda_\ell(W) \, \alpha_t(\ell)$ and the composite type weights sum to one.\footnote{Proposition (ref) generalizes the Rotemberg decomposition of GoldsmithPinkhamSorkinSwift2020 from a single estimator to the full GMM class. Their decomposition applies to the specific weighting matrix $W = GG'$ in the Bartik setting; Proposition (ref) holds for arbitrary $W$. Both decompositions have a two-level weight structure: outer weights ($\lambda_\ell(W)$ here; Rotemberg weights there) on just-identified estimates, and inner weights ($\alpha_t(\ell)$ here; location weights there) on unit-specific treatment effects. In both cases, the inner weights are non-negative under a design condition (Assumption (ref) here; a same-sign and exogeneity condition there), but the outer weights can be negative. The instrument structures differ: they work with continuous industry shares that sum to one, while the present framework uses binary instruments with no summing constraint. Their framework has broader empirical scope across labor, trade, macro, and development economics.} Under PRD, each inner weight $\alpha_t(\ell) \geq 0$ (Proposition (ref)), but the outer weights $\lambda_\ell(W)$ can be negative, and $\psi_t(W)$ can be negative as a result. When $\psi_t(W) < 0$ for some type, the GMM estimand cannot be interpreted as a weighted average treatment effect for any subpopulation. A common treatment effect ($\mathrm{LATE}_t = \beta_0$ for all active types) makes $W$ irrelevant, since all weighting matrices recover $\beta_0$. Heterogeneity breaks this invariance: the choice of $W$ is a choice of estimand.
The weight $\lambda^{2SLS}_\ell$ is the partial first-stage contribution of instrument $\ell$ after partialling out the other instruments; with correlated instruments ($\Sigma_Z$ non-diagonal), this partial contribution can be negative even when every instrument has a positive marginal first stage, and $\psi_t^{2SLS}$ can be negative.
Within the HallInoue2003 framework, EGMM minimizes the asymptotic variance of $\hat{\beta}$ around its own pseudo-true value $\beta^*_{EGMM}$, but $\beta^*_{EGMM}$ is itself a weighted average of $\mathrm{LATE}_t$ determined by the variance structure. The 2SLS weighting matrix $\Sigma_Z$ reflects only the instrument covariance; $\Omega$ also absorbs the residual variance from fitting a single $\beta$ to $L$ distinct Wald estimands. When instrument $\ell$'s compliers have dispersed treatment effects, the common residual $Y_i - \beta^*_{EGMM} D_i$ fits that instrument's moment condition poorly, inflating $[\Omega]_{\ell\ell}$. EGMM downweights these instruments: $\partial \lambda^{EGMM}_\ell / \partial [\Omega]_{\ell\ell} < 0$ for any positive definite $\Omega$ when $\lambda^{EGMM}_\ell > 0$ (Appendix (ref)). This is the heterogeneity penalty: the resulting estimand is determined by the variance structure of the data, not by the researcher.
As with 2SLS, the weights $\lambda^{EGMM}_\ell$ can be negative when $[\Omega^{-1}\boldsymbol{\gamma}]_\ell < 0$ under dependent instruments. The heterogeneity penalty operates on the composite weights $\psi_t(\Omega^{-1})$: by concentrating the outer weights $\lambda^{EGMM}_\ell$ on instruments with low treatment effect dispersion, EGMM can push $\psi_t(\Omega^{-1}) < 0$ for compliance types that are weighted heavily by the penalized instruments.\footnote{MogstadTorgovitskyWalters2021 show that 2SLS weights on complier groups can be negative under partial monotonicity, weaker than Assumption (ref). The heterogeneity penalty amplifies this: the variance-minimizing objective makes EGMM weights more concentrated and more likely to fall outside the simplex. GoldsmithPinkhamSorkinSwift2020 observe that Rotemberg weights can be negative; the heterogeneity penalty identifies the mechanism for efficient GMM.}
Rejection does not necessarily imply instrument invalidity.\footnote{The diagnostic interpretation of $J$-test rejection as evidence of heterogeneity rather than invalidity is explicit in GoldsmithPinkhamSorkinSwift2020 and consistent with the compliance-type framework of MogstadTorgovitskyWalters2021. AndrewsChen2025 show that, asymptotically under local misspecification, the $J$-statistic characterizes the range of estimates achievable across weighting matrices at a given standard error relative to the efficient estimator.} Under maintained validity, rejection indicates that the type-specific treatment effects $\mathrm{LATE}_t$ are heterogeneous and that different instruments weight them differently. For instance, the 2SLS and EGMM estimands then converge to different weighted averages of $\mathrm{LATE}_t$, and the gap between them reflects different weighting of compliance types by $\Sigma_Z^{-1}$ and $\Omega^{-1}$ respectively.
Two assumptions isolate a setting in which every Wald estimand equals $\mathrm{LATE}_t$ for a single type, all weights are non-negative, $\Omega$ is diagonal, and all equations admit closed-form expressions.
Under Assumption (ref), the indirect effect in Lemma (ref) vanishes and all weights $\alpha_t(\ell)$ are non-negative. Combined with Assumption (ref), each complier group $\mathcal{C}_\ell$ contains a single compliance type, and $\mathrm{Wald}_\ell = \mathbb{E}[Y_i(1) - Y_i(0) \mid i \in \mathcal{C}_\ell]$: each Wald estimand is the complier-group average treatment effect in the sense of ImbensAngrist1994.
The ratio $\lambda^{EGMM}_\ell / \lambda^{2SLS}_\ell \propto \tfrac{1}{\sigma^2_{\epsilon,\ell}}$: instruments with larger residual variance receive less weight. Since $\sigma^2_{\epsilon,\ell}$ is increasing in $\sigma^2_{\tau,\ell} \equiv \operatorname{Var}(Y_i(1) - Y_i(0) \mid i \in \mathcal{C}_\ell)$, the within-complier treatment effect variance, EGMM downweights instruments whose complier groups exhibit high treatment effect dispersion. The 2SLS and EGMM weights coincide if and only if $\sigma^2_{\epsilon,\ell}$ is constant across instruments (Corollary (ref)). Under diagonality, the heterogeneity penalty derivative is $\partial \lambda^{EGMM}_\ell / \partial \sigma^2_{\tau,\ell} < 0$ with no ceteris paribus qualification, since $\sigma^2_{\tau,\ell}$ enters only $[\Omega]_{\ell\ell}$. The full diagonal specialization is in Appendix (ref).
A natural response is to choose a weighting matrix that delivers the desired weights directly: specify a causally interpretable target $\omega$, find the $W$ that delivers $\lambda(W) = \omega$, and run GMM. The first question is whether such a $W$ always exists.
Lemma (ref) confirms GMM can target any weighted average of Wald estimands. Under misspecification, the asymptotic variance for a fixed W is the HallInoue2003 sandwich: $V(W;\,\Omega) = \frac{\boldsymbol{\gamma}' W \,\Omega\, W \boldsymbol{\gamma}}{(\boldsymbol{\gamma}' W \boldsymbol{\gamma})^2},$ where $\Omega = \mathbb{E}[g_i(\beta^*)\,g_i(\beta^*)']$ is the second moment matrix at the pseudo-true value. Using this formula, we can show that any W delivering the same target produces the exact same asymptotic variance.
Therefore, there is no room to optimize within the constrained GMM class: fixing the target $\omega$ fixes the variance. The question remains whether the constrained GMM variance can reach the GMM efficiency floor.
The variance-minimizing matrix $W = \Omega_\omega^{-1}$ produces implied weights that drift away from the target $\omega$, simply because $\beta^*(\omega)$ is not a fixed point of the EGMM mapping. Forcing the estimator to hit $\omega$ requires a suboptimal $W$. This impossibility is a structural flaw of GMM with the $L$ moment conditions in (ref): any matrix that successfully delivers the researcher's target must rely on a misspecified common residual, pushing its variance above the Cauchy-Schwarz floor. Representative Targeting escapes this constraint entirely. By leveraging instrument-specific residuals, RT achieves a variance strictly below the constrained GMM variance whenever Wald estimands differ.
Within the GMM class, the researcher can choose any target, but every GMM estimator at that target fits a single common residual to $L$ distinct Wald estimands, and produces a variance that is pinned by this misspecified fit (Proposition (ref)). Pursuing efficiency makes things worse: the heterogeneity penalty distorts the efficient weights away from the target, and no weighting matrix can simultaneously achieve the efficiency bound and deliver researcher-specified weights (Proposition (ref)). MogstadTorgovitskyWalters2021 observe that the treatment effect parameter identified by 2SLS may not answer an economically relevant question even when the weights are non-negative. RT resolves all these tensions by leaving the GMM class entirely: it directly computes the weighted average of the instrument-specific Wald estimates, without fitting a common residual and without the GMM variance penalty.
Unlike 2SLS and EGMM, whose composite type weights $\psi_t(W)$ can be negative (Section (ref)), RT with $\omega \in \Delta^{L-1}$ guarantees $\psi_t(\omega) \geq 0$ under PRD: the estimand is a proper weighted average of type-specific treatment effects. All results extend to settings with covariates $X_i$ under conditional versions of Assumptions (ref)--(ref) and full first-stage saturation BlandholBonneyMogstadTorgovitsky2022; see Appendix (ref).
RT is also semiparametrically efficient: its variance $V_{RT}(\omega)$ equals the efficiency bound for $\beta^*(\omega)$.\footnote{The bound coincides with the Chamberlain1987 bound for the nonparametric model with $\mathbb{E}[Y_i^2] < \infty$, $\pi_\ell > 0$, and $\Sigma_Z \succ 0$. Imposing the LATE model does not lower it when $\theta_t > 0$ for all $t \in \mathcal{T}$: with $L$ binary instruments there are $|\mathcal{Z}| \leq 2^L$ conditional means (where $\mathcal{Z}=\operatorname{supp}(\mathbf{Z}_i)$)but far more monotone compliance types ($6$ for $L=2$, $20$ for $L=3$), each with its own $\mathrm{LATE}_t$. When all types are present, the structural parameters outnumber the moments they determine, and the model cannot restrict the joint distribution beyond what the data alone impose. When the number of “active types" $\mathcal{T}_+ = \{t \in \mathcal{T} : \theta_t > 0\}$ is very small, imposing the LATE model may lower the bound.}
For $L \ge 3$, multiple weighting schemes $\omega$ can yield the same scalar estimand $\beta^*$. For any target estimand $\beta^* \in [\min_\ell Wald_\ell, \max_\ell Wald_\ell]$, we define the Variance Frontier as the minimum semiparametric variance achievable across all non-negative weights delivering that estimand:
Because $V_{RT}(\omega)$ is the semiparametric bound for a specific set of weights $\omega$ (Proposition (ref)), $V_{frontier}(\beta^*)$ represents the absolute minimum semiparametric cost over all admissible $\omega$ that deliver the same target $\beta^*$.\footnote{For $L = 2$, each $\beta^*$ determines a unique $\omega \in \Delta^{L-1}$, every target sits on the frontier, and the weight-composition cost is zero. For $L \geq 3$, the frontier is the lower envelope of a quadratic over a polytope, computable by parametric quadratic programming in ((ref)).}
For any specific choice of $\omega \in \Delta^{L-1}$, the RT variance naturally decomposes into the frontier variance and a weight-composition cost:
The weight-composition cost in (ref) captures what the researcher pays for targeting a specific compliance-type composition $\psi_t(\omega)$, not the cheapest weights achieving the same scalar $\beta^*$. The Variance Frontier is computed entirely from $\Gamma^{Wald}$; EGMM does not lie on it.\footnote{Under heterogeneity, $\Gamma^{Wald} \neq D^{-1}\Omega D^{-1}$: each Wald estimator uses its own instrument-specific residual, while the GMM sandwich uses a common residual at $\beta^*_{EGMM}$. The two formulas coincide under homogeneity when $\Omega$ is computed with FWL-centered residuals (i.e., $\Omega^{FWL}_{\ell k} = E[\tilde{\epsilon}_i^2(Z_{\ell i} - p_\ell)(Z_{ki} - p_k)]$).}
Several choices of $\omega$ carry natural causal interpretations without the identification of $\alpha_t(\ell)$ in compliance-type composition $\psi_t(\omega)$.
Both 2SLS and EGMM are GMM estimators while RT is a different object. The two constructions share a probability limit when $\omega = \lambda(W)$ but have different influence functions and, under heterogeneous treatment effects, different asymptotic variances.\footnote{Both RT and constrained GMM are weighted averages of Wald estimators with the same weights $\omega$, but their influence functions differ in the first-stage component. RT's influence function linearizes each Wald ratio at its own population value $\mathrm{Wald}_\ell$, producing instrument-specific residuals $\tilde{\epsilon}_{i,\ell} = Y_i - \mathrm{Wald}_\ell\, D_i - c_\ell$. The GMM influence function linearizes around a common pseudo-true value $\beta^*(\omega)$, producing a common residual $Y_i - \beta^*(\omega)\, D_i$. Under homogeneity ($\mathrm{Wald}_1 = \cdots = \mathrm{Wald}_L$), the two coincide and the variance gap vanishes.}
Vytlacil2002 shows that Assumptions (ref)--(ref) are equivalent to Assumption (ref). Under the normalization $U_i \sim U(0,1)$, the propensity score $p(z) = P(D_i = 1 \mid \mathbf{Z}_i = z)$ equals $V(z)$, and the marginal treatment effect $\mathrm{MTE}(u) = \mathbb{E}[Y_i(1) - Y_i(0) \mid U_i = u]$ is the average treatment effect at latent resistance $u$, which represents an individual's unobserved reluctance to taking the treatment.
Each compliance type with $\theta_t > 0$ maps to a contiguous interval $\mathcal{R}_t \subset [0,1]$ that partitions the resistance space, converting the compliance-type decomposition into a continuous MTE integral with $\alpha_t(\ell) = \int_{\mathcal{R}_t} h_\ell(u)\,du$ (Lemma (ref), Appendix (ref)).
For each instrument $\ell$, the Wald estimand admits the MTE representation HeckmanVytlacil2005, HeckmanUrzuaVytlacil2006
where the weight function is
The weight function integrates to one but need not be non-negative.
The composite type weight $\psi_t(\omega)$ from Proposition (ref) gives one number per compliance type. The function $\bar{h}(u;\omega) = \sum_\ell \omega_\ell h_\ell(u)$ does the same thing continuously, assigning a weight to each value of latent resistance $u$. Under the latent index model, each compliance type $t$ corresponds to an interval of $u$, and $\psi_t(\omega) = \int_{\mathcal{R}_t} \bar{h}(u;\omega)\,du$.
The heterogeneity penalty (Section (ref)) has a direct MTE interpretation. The within-complier treatment effect variance $\sigma^2_{\tau,\ell}$ that inflates $[\Omega]_{\ell\ell}$ reflects variation in the MTE curve over the region of $u$ weighted by $h_\ell(u)$: where the MTE is steep, individual treatment effects within the complier group are dispersed, and the $\ell$-th moment condition is noisy. Under dependent instruments, $h_\ell(u)$ spreads across multiple compliance-type intervals through the indirect channel (Lemma (ref)), and the mechanism operates through the full weight function with cross-instrument interactions in $\Omega$.\footnote{Under independent instruments with non-overlapping compliance, each $h_\ell(u)$ is concentrated on a single interval of $u$ and the chain from MTE curvature to $[\Omega]_{\ell\ell}$ is direct (Appendix (ref)).} EGMM downweights instruments whose $h_\ell(u)$ span high-variation regions of the MTE, attenuating $\bar{h}(u;\lambda^{EGMM})$ at those margins. The composite weight function is hollowed out where treatment effects are most heterogeneous, and concentrated where the MTE is nearly constant. The EGMM estimand is pulled toward margins of low heterogeneity, regardless of whether those margins are economically relevant. RT reverses this: the researcher specifies $\omega$ to restore weight at the margins that matter for the economic question, and the composite $\bar{h}(u;\omega)$ fills in the regions that EGMM hollows out.
Proposition (ref) extends to MTE weight functions.
With $h_\ell(u) \geq 0$, each Wald estimand is a proper weighted average of marginal treatment effects, and any RT composite satisfies $\bar{h}(u; \omega) \geq 0$ for $\omega \in \Delta^{L-1}$.
A policy that changes the instrument distribution from $F_0(\mathbf{z})$ to $F_1(\mathbf{z})$ induces the policy-relevant treatment effect HeckmanVytlacil2005, CarneiroHeckmanVytlacil2011:
Under the latent index model, the PRTE admits the MTE representation
where $F_j(u) = P(p(\mathbf{Z}_i) \geq u \mid \text{policy } j)$ for $j = 0,1$, and $\int_0^1 w^P(u)\,du = 1$ by construction.
With discrete instruments, policy-relevant treatment effects are generally only partially identified, because the policy weight function $w^P$ rarely lies within the $(L-1)$-dimensional convex hull $\mathrm{conv}\{h_1, \ldots, h_L\}$ spanned by the observed instruments MogstadSantosTorgovitsky2018. MogstadSantosTorgovitsky2018 construct bounds on the PRTE using restrictions on marginal treatment response functions; their linear programming approach delivers finite nonparametric bounds that tighten substantially with shape restrictions.\footnote{The two frameworks differ in maintained assumptions: MogstadSantosTorgovitsky2018 require the latent index model but can impose shape restrictions (monotone treatment response, separability) that tighten bounds; the compliance-type results in Sections (ref)--(ref) require only Assumptions (ref)--(ref). MogstadSantosTorgovitsky2018 also accommodate continuous instruments through the propensity score, while the analysis here focuses on binary instruments.} When a point estimate is needed, a natural approach is to choose $\omega \in \Delta^{L-1}$ so that $\bar{h}(u;\omega)$ is as close to $w^P(u)$ as possible.
The RT estimand $\beta^*(\omega^{PRTE})$ point-identifies a variance-optimal, $L^2$-closest surrogate for the PRTE. This complements the partial identification approach of MogstadSantosTorgovitsky2018, who instead bound the exact PRTE.
The identification gap $\Delta \equiv \beta^*(\omega^{PRTE}) - \mathrm{PRTE}$ vanishes for any constant MTE and is bounded in Proposition (ref) (Appendix (ref)). Under a Lipschitz condition, $|\Delta| \leq M\,\|e\|_{L^2}/(2\sqrt{3})$, where $e(u) = \bar{h}(u;\omega^{PRTE}) - w^P(u)$ and the Wald range $\max_\ell \mathrm{Wald}_\ell - \min_\ell \mathrm{Wald}_\ell$ heuristically scales $M$. Alternatively, $\Delta$ can be bounded without smoothness assumptions via linear programming over shape-restricted marginal treatment response functions MogstadSantosTorgovitsky2018.
The Tennessee STAR experiment WordEtal1990, Krueger1999 and the patent examiner design FarreMensaHegdeLjungqvist2020 sit at opposite ends of the assumption hierarchy. In STAR (Section (ref)), instruments satisfy the diagonal specialization. The patent design (Section (ref)) has correlated cumulative leniency instruments and overlapping compliance groups.
Tennessee's STAR experiment randomly assigned kindergarteners to small, regular, or regular-with-aide classes within 79 schools WordEtal1990. Following the standard comparison of small versus regular classes, the aide arm is excluded; 78 schools have sufficient enrollment in both arms for the analysis.\footnote{Schools with fewer than 10 students or fewer than 3 per arm are dropped. Baseline covariates are balanced (Table (ref)); results are robust to varying these thresholds (Tables (ref)--(ref) in Appendix (ref)).} The outcome is the kindergarten math score (Stanford Achievement Test, scaled); sample restrictions and robustness checks are in Table (ref) and Appendix (ref). Each school operates an independent randomization and perfect within-school compliance, giving joint independence (Assumption (ref)) and monotonicity (Assumption (ref)) respectively. Each student attends exactly one school, giving non-overlapping compliance (Assumption (ref)). Treating each school as a separate instrument and conditioning on school fixed effects via Frisch-Waugh-Lovell gives $L=78$ mutually exclusive and mean-zero instruments.\footnote{The school fixed effects fully saturate the first stage, preserving the causal interpretation of each within-school Wald estimand through the FWL aggregation BlandholBonneyMogstadTorgovitsky2022.} Thus, $\Sigma_Z$ and $\Omega$ are diagonal (Footnote (ref)), and the diagonal specialization of Section (ref) applies.
Each school's Wald estimand is the average treatment effect for that school's compliers ImbensAngrist1994, and these effects span a wide range (Figure (ref)). The $J$-statistic decisively rejects equality across schools (Table (ref), Proposition (ref)). EGMM reduces the 2SLS estimate by a quarter (Table (ref)).\footnote{The gap is from the heterogeneity penalty, not many-instruments bias: the demeaned treatment lies in the column space of $\mathbf{Z}_i$, the first-stage projection is exact, and LIML, JIVE, and 2SLS coincide BoundJaegerBaker1995. Details in Appendix (ref).} Under the STAR instrument construction with perfect within-school compliance ($\pi_\ell$ constant across schools after FWL normalization), $\lambda^{2SLS}_\ell = \gamma_\ell / \sum_k \gamma_k = \omega^{CSW}_\ell$ exactly; 2SLS and CSW-ATE target the same estimand but, aligning with Proposition (ref), have different standard errors because 2SLS uses the GMM sandwich while RT uses $\Gamma^{Wald}$. EGMM weight decreases monotonically with the residual variance $\sigma^2_{\epsilon,\ell}$ (Figure (ref)), as Proposition (ref) predicts. Schools where small classes generate the largest gains also have the most within-school outcome dispersion; EGMM penalizes this dispersion and shifts the estimand toward schools with more moderate effects (Appendix (ref); Figure (ref)).
Within-school randomization identifies $\sigma^2_{\tau,\ell}$ and $\sigma^2_{Y(0),\ell}$ separately, allowing decomposition of $\sigma^2_{\epsilon,\ell}$ into baseline outcome variance, treatment effect heterogeneity, and LATE-deviation components (equation (ref); Figure (ref)). Baseline outcome variance dominates; treatment effect heterogeneity contributes a small share. The penalty operates on the cross-school variation in $\sigma^2_{\epsilon,\ell}$, not its level: schools with high $\sigma^2_{\tau,\ell}$ have differentially high residual variance, and EGMM penalizes this differential. A modest heterogeneity share generates the quarter estimand reduction because the penalty acts multiplicatively through $\lambda^{EGMM}_\ell / \lambda^{2SLS}_\ell \propto 1/\sigma^2_{\epsilon,\ell}$. The weight distortion map (Figure (ref)) confirms that the heterogeneity penalty is not collinear with instrument strength.
The variance frontier (Figure (ref); Proposition (ref); Corollary (ref)) confirms that representativeness is not expensive in this design. A calibrated simulation (2{,}000 replications; Table (ref) in Appendix (ref); DGP and calibration in Appendices (ref) and (ref)) confirms that these results survive sampling variability.
The analysis sample covers 34{,}434 first-time patent applications examined by 5{,}915 examiners at the United States Patent and Trademark Office, 2001--2009 (Table (ref)), constructed from the replication data of FarreMensaHegdeLjungqvist2020. The treatment $D_i$ is patent approval, and the outcome $Y_i$ is the 5-year forward citation count. Patent applications are quasi-randomly assigned to examiners within art unit $\times$ year cells. We estimate examiner leniency as the leave-one-out approval rate, residualized on art unit $\times$ year fixed effects. We group examiners into $Q = 7$ leniency quantile groups and construct $L = 6$ cumulative instruments $Z_k = \mathbf{1}\{G_i \geq k + 1\}$ for $k = 1, \ldots, 6$, where $G_i$ denotes the leniency group. The leniency distribution is in Figure (ref).
Assumptions (ref)--(ref) hold: joint independence follows because $\mathbf{Z}_i$ is a deterministic function of $G_i$ and quasi-random assignment gives $G_i \perp\!\!\!\perp (Y_i(0), Y_i(1), D_i(\cdot))$. PRD (Assumption (ref)) holds because the cumulative instruments are nondecreasing in the common source $G_i$. The cumulative instruments are positively correlated and compliance groups overlap; Assumptions (ref)--(ref) fail by construction. All results use the general weight formulas with the full (non-diagonal) $\hat{\Omega}$, clustered at the examiner level.\footnote{GoldsmithPinkhamHullKolesar2025 recommend UJIVE and non-clustered standard errors for many-examiner leniency designs. With $L = 6$ cumulative threshold instruments, many-instrument concerns are less acute; examiner-level clustering is conservative relative to their recommendation.} The first-stage $F$-statistic is 253.6. Pre-treatment characteristics are largely balanced across leniency groups, and estimates are stable under leave-one-group-out (Tables (ref) and (ref) in Appendix (ref)).
Monotonicity violations are a concern in leniency designs: examiners may apply heterogeneous scrutiny standards across technology classes, leaving a globally lenient examiner strict in certain fields.\footnote{The broader leniency-design literature has scrutinized monotonicity in judge and examiner settings ChynFrandsenLeslie2024. FrandsenLefgrenLeslie2023 develop a joint test in a judge leniency design that cannot distinguish exclusion from monotonicity violations, and propose a weaker average monotonicity condition. Sigstad2026 finds violations in up to 50% of non-unanimous judicial panel decisions, though these typically induce little bias; the patent setting uses individual examiners, not panels.} The cumulative threshold structure mitigates this: since $Z_k = \mathbf{1}\{G_i \geq k+1\}$ and $G_i$ is a scalar leniency index, component-wise monotonicity (Assumption (ref)) requires only that higher overall leniency weakly increases approval, which is the defining feature of the leniency instrument.
The $J$-statistic rejects Wald estimand equality ($J = 16.36$, $p = 0.006$; Table (ref), Proposition (ref)), with the threshold-specific Wald estimates shown in Figure (ref). EGMM cuts the 2SLS estimate nearly in half (Table (ref)); The mechanism is weight concentration (Figure (ref), Table (ref)): EGMM places $86\%$ of its weight on the lowest threshold and assigns negative weights to $G \geq 5$ and $G \geq 6$, pulling the estimand below every individual Wald estimate. Results are robust across $Q \in \{4, \ldots, 20\}$ (Appendix (ref)). The three RT targets redistribute weight across thresholds and produce estimands that exceed 2SLS.
Returning to the composite MTE weight functions previewed in Figure (ref), EGMM's composite is attenuated at the high-resistance end ($u$ near 1), where the Wald estimates are largest with the widest confidence intervals, reflecting the heterogeneity penalty and the hollowing-out effect made visible in the data.\footnote{With negative $\lambda^{EGMM}_\ell$, the EGMM composite $\bar{h}(u;\lambda^{EGMM})$ can go negative at some $u$. Here it remains positive because the $86\%$ weight on $h_1(u)$, which spans the full propensity-score range, dominates the small negative contributions from $h_5$ and $h_6$. RT composites $\bar{h}(u;\omega)$ with $\omega \in \Delta^{L-1}$ are non-negative by PRD (Proposition (ref)): each $h_\ell(u) \geq 0$ and $\omega_\ell \geq 0$. The six instrument-specific weight functions are in Figure (ref) (Appendix (ref)).} The PRTE-targeted composite places mass at both ends of the resistance distribution, mirroring the staircase shape of the policy target $w^P(u)$; the $L^2$ projection in Proposition (ref) achieves this fit with a relative error of $0.12\%$.
A policymaker considering a uniform relaxation of examiner scrutiny needs the PRTE. We define the PRTE policy as a one-step upward shift of approval rate in examiner leniency, while the most lenient group remains unchanged. The staircase policy shifts all intermediate margins uniformly, and $\omega^{PRTE}$ concentrated on $G \geq 2$ and $G \geq 7$ (Figure (ref)). The RT surrogate for the PRTE is $11.75$ citations per marginal approval (Table (ref)). The identification gap between this surrogate and the true PRTE is bounded at less than $0.03$ citations under non-negative MTE and below $0.001$ when monotone treatment selection is added (Proposition (ref); Table (ref) and Figure (ref) in Appendix (ref)).
The variance frontier (Figure (ref)) maps the minimum semiparametric variance $V_{frontier}(\beta^*)$ across all RT weights delivering a given estimand, which is an absolute semiparametric bound. RT targets (CSW-ATE, EW-ATE, PRTE) sit weakly above the frontier, with the vertical gap equal to the weight-composition cost in (ref). The PRTE cost is the largest as the policy weight function uniquely pins the target weights and leaves no room for variance optimization. For comparison, 2SLS is marked at its own GMM sandwich standard error which is close to the RT variance at the same weights. EGMM's implied weights have negative entries ($\lambda^{EGMM}_4 = -0.14$), placing its estimand below every individual Wald estimate and outside the simplex-feasible range.
GMM is the standard tool for combining moment conditions, but it may fail as a causal estimand under heterogeneous treatment effects due to negative weighting. Even under positive weights, it is fundamentally suboptimal for combining Wald estimands because it forces a single common residual onto multiple distinct estimands. As a result, its variance is driven by the misspecification structure rather than instrument-specific precision. Pursuing efficiency within GMM makes things worse: the heterogeneity penalty distorts the weights, and no GMM weighting matrix achieves the efficiency bound while delivering researcher-specified weights. The semiparametrically optimal estimator for a weighted average of Wald estimands is the weighted average itself. RT computes each Wald ratio separately, uses instrument-specific residuals, and achieves the semiparametric efficiency bound at a closed-form quadratic cost.
These methodological differences yield substantial empirical consequences. In the STAR experiment, GMM's common-residual architecture penalizes schools with the largest treatment effects, pulling the EGMM estimand (6.55) a quarter below the 2SLS estimand (8.84). In a patent examiner leniency design, the distortions are severe: GMM weights concentrate 86% on the lowest threshold and apply negative weights to $G \geq 5$ and $G \geq 6$, cutting the EGMM estimate ($5.51$ citations) to half of the 2SLS estimate ($10.58$). For a policymaker evaluating uniform examiner scrutiny relaxation, the relevant metric is the PRTE, which RT targets at $11.75$ citations per marginal approval (the $L^2$-closest feasible approximation; identification gap $< 0.03$ citations under non-negative MTE). These gaps are not sampling variability; they reflect which subpopulations the estimand represents.
The primary limitation of the current approach is its restriction to binary treatments and instruments. To address this, combining the response-type framework of AngristSantosTecchio2025 for multinomial treatments with RT could simultaneously accommodate multiple treatments and instruments. On the instrument side, extending to continuous instruments, following the work of MogstadSantosTorgovitsky2018, represents a natural next step. Additionally, our non-negative weight results currently rely on the assumption of PRD. While this holds by construction in the instrument structures that dominate quasi-experimental practice, it may fail in settings with capacity constraints or strategic interactions between instrument sources. Finally, it remains an open question whether the covariate extension (Proposition (ref) in Appendix (ref)) can be freed from the first-stage saturation requirement identified by BlandholBonneyMogstadTorgovitsky2022.