EconBase
← Back to paper

Inference on the TSLS Estimand with Weak Instruments and Treatment Effect Heterogeneity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

58,674 characters · 5 sections · 67 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference on the TSLS Estimand with Weak Instruments and Treatment Effect Heterogeneity

abstractTraditional inference on the coefficient in an instrumental variables regression does not retain size when the instrument set is weak. With constant treatment effects or one instrument, the andersonEstimationParametersSingle1949 AR test, the kleibergenPivotalStatisticsTesting2002--moreiraConditionalLikelihoodRatio2003 LM test, and the Moreira CLR test provide robust alternatives which retain validity. Under treatment effect heterogeneity, no valid inference procedure exists in the overidentified setting. This paper develops the TSLS likelihood ratio (TLR) statistic, for performing inference on the TSLS estimand. When combined with a two-step procedure in the spirit of bergerValuesMaximizedConfidence1994, it retains uniform validity across both the weak- and strong-instrument regimes. The procedure retains power with small choices of first-step level, hence the test can be constructed to numerically coincide with the Wald test in the strong-instrument limit.

Introduction

Consider the problem of performing inference on a parameter of interest,

align[align omitted — 117 chars of source]

when this parameter is the coefficient in an overidentified instrumental variables regression with one endogenous regressor and a potentially weak set of instruments, known as the two-stage least squares (TSLS) estimand,

align[align omitted — 124 chars of source]

The TSLS estimand is a weighted sum of the elements of the reduced-form coefficient vector, \(\delta\), weighted by the elements of the first-stage coefficient vector, \(\gamma\), and the population instrument Gram matrix, \(\Sigma_{ZZ}\). It is a common parameter of interest in economics, political science, sociology, and other fields. Under certain assumptions, it obtains a causal interpretation as a positively weighted average of local average treatment effects, see e.g. angristTwoStageLeastSquares1995 and mogstadCausalInterpretationTwoStage2021.

In the presence of weak instruments, the commonly applied Wald- or t-test is known to suffer from size distortion. This problem is well-studied, and in the setting of a linear model with constant treatment effects (CTE), the AR test (andersonEstimationParametersSingle1949), Kleibergen--Moreira Lagrange multiplier (LM) test (kleibergenPivotalStatisticsTesting2002, moreiraConditionalLikelihoodRatio2003) and conditional likelihood ratio (CLR) test (moreiraConditionalLikelihoodRatio2003) are known to deliver valid inference.\footnote{See finlayImplementingWeakInstrumentRobust2009 and andrewsWeakInstrumentsInstrumental2019 for parsimonious expositions.} In a model with treatment effect heterogeneity, however, these are not valid tests for hypotheses on the TSLS estimand. As pointed out by yapInferenceManyWeak2025, there exists no test proven to be valid (i) in the presence of weak instruments and (ii) treatment effect heterogeneity which (iii) targets a parameter interpretable as a weighted average of local average treatment effects under commonly maintained assumptions when (iv) considering limit experiments that fix the number of instruments.

To understand why standard robust tests are not valid for hypotheses on the TSLS estimand under treatment effect heterogeneity, observe that the AR test performs inference on the hypothesis \(\beta=\beta_0\) by instead performing inference on a set of linear restrictions,

align[align omitted — 124 chars of source]

where the number of restrictions is equal to the number of instruments. The set of restrictions can be rewritten as a composite hypothesis on the following form,

align[align omitted — 155 chars of source]

As is shown, e.g., by moreiraConditionalLikelihoodRatio2003, the AR test statistic can be written as the sum of the LM test statistic and the sarganEstimationEconomicRelationships1958--hansenLargeSampleProperties1982 $J$ test statistic. Maintaining CTE, the $J$ test is a test of instrument exogeneity. Maintaining exogeneity, but allowing for treatment effect heterogeneity, the $J$ test is essentially a test of the CTE assumption. The CLR and LM tests are derived from the AR restrictions, testing \(\beta=\beta_0\) while maintaining the assumption of CTE. As a consequence, they are invalid in the presence of treatment effect heterogeneity when considered as tests for the TSLS hypothesis.

This paper proposes a test that is valid with weak instruments and treatment effect heterogeneity. The test is derived by first observing that a hypothesis about the TSLS estimand, \(\beta=\beta_0\), can be reformulated into a quadratic constraint by multiplying by the denominator on both sides and rearranging,

align[align omitted — 209 chars of source]

While the restriction is similar in form to the AR restriction in that it does not involve division by a potentially small number, it is a single scalar restriction, and does not impose or test any ancillary assumptions, such as CTE.

The estimators of the reduced-form and first-stage coefficients are jointly asymptotically normal, with a variance-covariance matrix that is consistently estimable. This admits a likelihood ratio statistic, formulated as a quadratic optimization program with a quadratic constraint. We call this the TSLS likelihood ratio (TLR) statistic. While the program does not have a closed-form solution, it is a special case of the generalized trust region subproblem, which has a known and unique solution, see e.g. sternIndefiniteTrustRegion1995. When instruments are strong, the TLR statistic is distributed asymptotically chi-squared with one degree of freedom. Combined with the fact that likelihood ratio tests are asymptotically equivalent to their associated Wald test under the null and against local alternatives (ApproximationTheoremsMathematical1980), this makes the implied test numerically indistinguishable from the Wald test in the strong-instrument regime.

When the instrument set is weak, the limiting distribution of the TLR statistic depends on two nuisance parameters. The first nuisance parameter is a structural endogeneity parameter, capturing both selection on gains and levels. Similarly to the one-instrument case (see e.g. vandesijpePowerConditionalLikelihood2023 and leeWhatWhenYou2023) it is consistently estimable under the null. The second nuisance parameter captures the overall identifying power of the instrument set. It is closely related to, but distinct from, the {concentration parameter} often used to gauge instrument strength in the literature. While it cannot be consistently estimated, it is associated with a sample statistic similar to, but distinct from, the first-stage F-statistic. This allows for a two-step procedure similar to the one discussed by bergerValuesMaximizedConfidence1994, and suggested by staigerInstrumentalVariablesRegression1997 for valid Wald inference in the linear constant effects model. Unlike their approach, we leverage more of the information contained in the sample to produce a test that is less conservative. In particular, we can set the first-step level so low as to numerically asymptote to the Wald test with strong instruments without substantial loss of power.

\paragraph*{Literature.}

The paper contributes to the literature on inference with weak instruments. Following the seminal contributions of nelsonDistributionInstrumentalVariables1990,nelsonFurtherResultsExact1990, boundProblemsInstrumentalVariables1995 and staigerInstrumentalVariablesRegression1997, a substantial number of papers have been devoted to understanding and resolving the finite-sample shortcomings of common procedures for inference and estimation with weak instruments. For a comprehensive review, see andrewsWeakInstrumentsInstrumental2019. In the context of constant treatment effects, andersonEstimationParametersSingle1949 provided the first test with uniform validity, shown by moreiraTestsCorrectSize2009 to be uniformly most powerful unbiased with one instrument. Later, kleibergenPivotalStatisticsTesting2002 and moreiraConditionalLikelihoodRatio2003 developed the Kleibergen--Moreira LM and Moreira CLR tests, with better power properties. andrewsOptimalTwoSidedInvariant2006 showed that the latter of these tests is uniformly most powerful unbiased under homoskedasticity. Recently, progress has been made on deriving more intuitive tests (leeValidTRatioInference2022) and tests with more intentional direction of power (leeWhatWhenYou2023) in the one-instrument case, both building on the two-step correction procedure suggested by staigerInstrumentalVariablesRegression1997, which uses the first-stage F-statistic and a Bonferroni correction to produce valid Wald inference.

The paper also contributes to the study of instrumental variables in policy evaluation and models with heterogeneous treatment effects. For a review of this literature, see mogstadInstrumentalVariablesUnobserved2024. While the TSLS estimand is one of many potential ways of aggregating reduced-form estimates when overidentified, following imbensIdentificationEstimationLocal1994b, angristTwoStageLeastSquares1995 and angristInterpretationInstrumentalVariables2000, it has been a common parameter of interest for empirical practitioners. Under certain assumptions, it obtains a weighted average of local average treatment effects. As is shown by kolesarEstimationInstrumentalVariables2013, this is not always the case for estimands associated with other estimators in the literature, such as the LIML estimator introduced by andersonEstimationParametersSingle1949, and the Fuller--$k$ estimator, introduced by fullerPropertiesModificationLimited1977. The methods considered in this paper are robust to violations of the so-called monotonicity (uniformity) assumption, discussed e.g. by heckmanUnderstandingInstrumentalVariables2006, mogstadCausalInterpretationTwoStage2021 and sigstadMonotonicityJudgesEvidence2026, as they do not rely on the causal interpretation of the estimand. While the one-instrument setting is more widespread in the literature, a common setting with more than one instrument is the examiner, judge or leniency design, see chynExaminerJudgeDesigns2025 and goldsmith-pinkhamLeniencyDesignsOperators2025 for reviews. In terms of valid inference in the presence of heterogeneous treatment effects, evdokimovInferenceInstrumentalVariables2018 study the problem in a setting with many instruments, and yapInferenceManyWeak2025 with many weak instruments. To the author's knowledge, this is the first paper to consider the problem of valid inference on the TSLS estimand with a fixed number of weak instruments and potentially heterogeneous treatment effects.

The paper is also related to the literature using two-step procedures to obtain valid inference in the presence of nuisance parameters, see e.g. lohNewMethodTesting1985, bergerValuesMaximizedConfidence1994, silvapulleTestPresenceNuisance1996, hansenTestSuperiorPredictive2005, chernozhukovIntersectionBoundsEstimation2013, romanoPracticalTwoStepMethod2014 and mccloskeyBonferronibasedSizecorrectionNonstandard2017. The two-step test proposed by staigerInstrumentalVariablesRegression1997 is an example of such a testing procedure in the weak-instrument setting.

\paragraph*{Roadmap.} The paper proceeds as follows. In (ref) we set up the linear instrumental variables model with heterogeneous treatment effects and derive a statistic that captures the total identifying power of the instrument set. In (ref) we show that existing weak-instrument robust tests fail as tests for the TSLS hypothesis with treatment effect heterogeneity. In (ref) we introduce the TLR statistic, and derive its limiting distribution. We show that it retains validity when combined with a two-step pretest and study its power properties in the weak- and strong-instrument regime. (ref) concludes.

Set-up and Notation

\paragraph*{Fundamentals.} Let \(\{(Y_i,D_i,Z_i)\}_{i=1}^n\) be an i.i.d. sample, where \(Y_i\) is an outcome, \(D_i\) a treatment, both scalar, and \(Z_i\) a \(d_Z\)-dimensional vector of instruments, with \(d_Z\geq 2\). Assume all variables are demeaned and have finite fourth moments. In the presence of covariates, let \(Y_i\) and \(D_i\) denote variables where their impact is appropriately projected off. We are interested in the linear heterogeneous treatment effects model,

align[align omitted — 69 chars of source]

where \(\beta_i\) is the causal effect of \(D_i\) on \(Y_i\), and \(U_i\) is the model residual. Denote the linear regressions of \(Y_i\) on \(Z_i\) (reduced form) and \(D_i\) on \(Z_i\) (first stage) as,

align[align omitted — 134 chars of source]

where \(\delta_n\) is the vector of reduced-form coefficients, \(\gamma_n\) the vector of first-stage coefficients, and \(W_i\) and \(V_i\) are regression residuals. Subscripts on \(\delta_n\) and \(\gamma_n\) indicate that we do not restrict the data generating process to be invariant with respect to the sample size \(n\). To analyze weak-instrument properties, we follow staigerInstrumentalVariablesRegression1997 in letting \((\delta_n, \gamma_n)\) drift towards zero at rate \(1/\sqrt{n}\), such that \((\delta_n,\gamma_n) \equiv { (\delta,\gamma)}/{\sqrt{n}}\) for some vectors \((\delta, \gamma)\). To study strong-instruments properties, we hold \((\delta_n, \gamma_n)\) fixed as \(n\to\infty\). Let the TSLS estimand, \(\beta\), be defined as in (ref). It is uniquely pinned down by \((\delta, \gamma)\), hence the parameter of interest does not vary with the sample size. Let \(D_i(z)\) denote the potential treatment of an individual with instrument realization \(Z_i=z\). We maintain the following assumptions.

assumptionp{IVR}[Relevance] \(\E[Z_iD_i] \neq \mathbf{0}_{d_Z}\)
assumptionp{IVX}[Strong Exogeneity] \((\beta_i,U_i,\{D_i(z)\}_{z\in\R^{d_Z}} ) \indep Z_i\)

Assumptions (ref) and (ref) are a subset of the assumptions usually maintained in models with heterogeneous treatment effects, see e.g. angristTwoStageLeastSquares1995 and angristInterpretationInstrumentalVariables2000. No assumption is imposed on the monotonicity (uniformity) of \(D_i(z)\) in \(z\). However, treatment effect heterogeneity requires the following refinement to the staigerInstrumentalVariablesRegression1997 weak-instrument limiting sequence.

assumptionp{IVW}[Uniform Weakness] For \(\sqrt{n}(\delta_n,\gamma_n)=(\delta,\gamma)\), as \(n\to\infty\) and \begin{enumerate}[label={(\alph*)}, ref={(\alph*)}] • \(D_i\) binary, \(\lim_{n\to\infty}\sup_{z,z'}\P\bigl[D_i(z)\neq D_i(z')\bigr]=0\)\(D_i\) continuous, \(\lim_{n\to\infty}\sup_{z,z'}\E\bigl[(D_i(z)- D_i(z'))^2\bigr]=0\) \end{enumerate}

Assumption (ref) rules out the case in which an instrument is weak because the share of individuals who are instrumented into more treatment (compliers) and less treatment (defiers) cancel. It is implied by the assumption of monotonicity (uniformity) of \(D_i(z)\) in \(z\), or equivalently by the vytlacilIndependenceMonotonicityLatent2002a single-index model.

\paragraph*{Observed moments.} Let \(\mathfrak{N}(\,\cdot\,,\,\cdot\,)\) denote the Normal distribution, \(\tau_n \equiv \left[

smallmatrix\delta_n \\ \gamma_n

\right]\) the stacked vector of first-stage and reduced-form coefficients and \(\hat \tau\equiv \left[

smallmatrix\hat \delta \\ \hat \gamma

\right]\) its OLS estimator. While the estimator is also a function of the sample size, for brevity of notation we let this be implicit in the following. The estimator has the following limiting distribution.

align[align omitted — 291 chars of source]

Let \(\Sigma_{ZZ}\equiv \E[Z_iZ_i']\) denote the population Gram matrix of \(Z_i\). We can rewrite the constraint defining the TSLS hypothesis from (ref) in terms of \(\tau_n\) as follows.

align[align omitted — 216 chars of source]

With constant treatment effects, the first-stage F-statistic summarizes the strength of the instrument set. With treatment effect heterogeneity, we need to consider a statistic that also captures other facets of the identifying power of the instruments. We record this statistic and its limiting distribution as a lemma.

lemmap{ST}[Summary Statistic] Let \(\hat \Sigma\) denote a consistent estimator for \(\Sigma\), and let \(\hat S \equiv \lVert (\hat\Sigma/n)^{-\frac{1}{2}}\hat \tau\rVert^2/d_Z\) denote the Wald test statistic for the hypothesis \(\tau_n=0\). Let \(\mathfrak{C}^2_{2d_Z}(\xi)\) denote the non-central chi-squared distribution with \(2d_Z\) degrees of freedom and non-centrality parameter \(\xi\). Then, {for \(\sqrt{n}\,\tau_n = \tau\), as \(n\to\infty\),} \[d_Z\hat S \convd \mathfrak{C}^2_{2d_Z}\left(\xi\right),\] where \(\xi \equiv \lVert (\Sigma/n)^{-\frac{1}{2}}\tau_n\rVert^2\).
proofSee Appendix (ref).

The parameter \(\xi\) is related to the concentration parameter often used to enumerate instrument strength in the literature. We record the relation as a lemma.

lemmap{PD}[Parameter Decomposition] Let \(\mu^2\equiv \lVert (\Sigma_{\hat \gamma}/n)^{-\frac{1}{2}}\gamma_n\rVert^2\) denote the concentration parameter, \(\omega\equiv \beta_{\mathrm{ols}}-\beta\), \(\beta_{\mathrm{ols}}\equiv \cov[Y_i,D_i]/\var[D_i]\), the difference between the OLS and TSLS estimands, and \(\nu^2 \equiv \lVert \Sigma_{ZZ}^{{1}/{2}}(\delta_n - \beta \gamma_n)\rVert^2 / \lVert \Sigma_{ZZ}^{{1}/{2}} \gamma_n\rVert^2\) the TSLS-weighted dispersion of local average treatment effects. Under Assumptions (ref) and (ref), as \({\sqrt{n}\,\tau_n=\tau}\) and \(n\to\infty\), we have, \begin{align*} \xi = \mu^2\left(1 + h^{-2}(\nu^2 + \omega^2)\right), \end{align*} where \(h^2\equiv{(\var[W_i]-\beta_{\mathrm{ols}}^{2}\var[V_i])/\var[V_i]}\) is a ratio of residual variances. All parameters are constant along the weak-instrument asymptotic sequence.
proofSee Appendix (ref).

Observe that with one instrument or constant treatment effects, we have \(\nu^2=0\).

\paragraph*{Eigenvalue Representation.}

It will be useful to express the data in terms of a particular eigendecomposition. Let \(\Lambda(\beta_0) \equiv \Sigma^{\frac{1}{2}} \Gamma(\beta_0) \Sigma^{\frac{1}{2}}\) denote the rotation of \(\Gamma(\beta_0)\) in terms of \(\Sigma\). Let \(X\) denote the matrix of eigenvectors of \(\Lambda(\beta_0)\), and \(\hat X\) its analog estimator. Define the rotated first-stage and reduced-form estimators and estimands as

align[align omitted — 146 chars of source]

While \(\Lambda(\beta_0)\) is indefinite, the block-structure of \(\Gamma(\beta_0)\) and (ref) restrict the sign of, and variation in, its eigenvalues.

lemmap{EH}[Eigenvalue Homogeneity] \(\Lambda(\beta_0)\) has \(d_Z\) positive and \(d_Z\) negative eigenvalues. Let \(\{\kappa^{(k)}_-,\kappa^{(k)}_+\}_{k=1}^{d_Z}\) denote these eigenvalues, and let \((q_+,q_-)\) denote the subvectors of \(q\) associated with the positive and negative eigenspace. Under Assumptions (ref) and (ref), along any sequence with \(\sqrt{n}\,\tau_n=\tau\), as \(n\to\infty\), for some common \((\kappa_-,\kappa_+)\), we have \begin{enumerate} • \(\kappa^{(k)}_-\to\kappa_-\) and \(\kappa^{(k)}_+\to\kappa_+\) for all \(k=1,\ldots,d_Z\). • \(\tau_n'\,\Gamma(\beta_0) \,\tau_n = 0 \iff (1+\varrho)\lVert q_+ \rVert^2 = (1-\varrho)\lVert q_- \rVert^2\) \end{enumerate} where \(\varrho \equiv (\kappa_+-\lvert \kappa_-\rvert) / (\kappa_++\lvert\kappa_-\rvert)\) denotes the signed eigenvalue share. If we let \(\omega_0 \equiv \beta_{\mathrm{ols}} - \beta_0 \) denote the value \(\omega\) takes under the null, we have \(\varrho = \omega_0 /\sqrt{h^2+\omega_0^2}\).
proofSee Appendix (ref).

Note that \((q, X, \kappa_\pm, \varrho)\) and their estimators depend on the hypothesized \(\beta_0\). For brevity of notation, we suppress this in the following.

On the Validity of Weak-Instrument Robust Tests

The AR, LM and CLR tests are commonly understood to provide inference that is robust to weak instruments. In the following, we show that this is not the case if the tests are interpreted as tests of the TSLS estimand under treatment effect heterogeneity. We start by recording the constant treatment effect limiting distribution of the three statistics, as derived by andersonEstimationParametersSingle1949, kleibergenPivotalStatisticsTesting2002 and moreiraConditionalLikelihoodRatio2003.

lemmap{ROB-CTE}[Robustness, Constant Treatment Effects] Assume there exists a scalar \(\beta\) such that \(\delta_n=\beta\gamma_n\). Let \(\hat\Omega(\beta_0)\) denote a consistent estimator of the variance of the restriction \(\hat\delta - \beta_0\hat\gamma\), let \(\tilde\gamma(\beta_0)\) denote the first-stage estimator orthogonalized relative to \(\hat\delta - \beta_0\hat\gamma\) and let \(\hat\Psi(\beta_0)\) denote an estimator of its variance. Define \begin{align*} \mathrm{AR}(\beta_0) &\equiv d_Z^{-1}\lVert (\hat\Omega(\beta_0)/n)^{-1/2}(\hat\delta - \beta_0\hat\gamma)\rVert^2, \\[2pt] \mathrm{LM}(\beta_0) &\equiv d_Z^{-1}\bigl[\tilde\gamma(\beta_0)'(\hat\Omega(\beta_0)/n)^{-1}(\hat\delta - \beta_0\hat\gamma)\bigr]^2\,\big/\,\lVert(\hat\Omega(\beta_0)/n)^{-1/2}\tilde\gamma(\beta_0)\rVert^2,\\[2pt] \mathrm{CLR}(\beta_0) &\equiv \tfrac{1}{2}\!\left[d_Z\mathrm{AR}(\beta_0)- \lVert\hat r(\beta_0)\rVert^2 + \sqrt{(d_Z\mathrm{AR}(\beta_0)-\lVert\hat r(\beta_0)\rVert^2)^2 + 4 \,d_Z\mathrm{LM}(\beta_0)\, \lVert\hat r(\beta_0)\rVert^2} \right] \end{align*} where \(\hat r(\beta_0) \equiv (\hat\Psi(\beta_0)/n)^{-1/2}\,\tilde\gamma(\beta_0)\). Let \(\mathfrak{C}^2_{d_Z}\) denote the chi-squared distribution with \(d_Z\) degrees of freedom. We then have, \(d_Z\mathrm{AR}(\beta_0)\convd \mathfrak{C}^2_{d_Z}\), \(d_Z\mathrm{LM}(\beta_0)\convd \mathfrak{C}^2_1\) and \(\mathrm{CLR}(\beta_0)\convd m(\mathfrak{C}^2_{1},\mathfrak{C}^2_{d_Z-1},\lVert \hat r(\beta_0)\rVert^2)\), where \begin{align*} m(x_1,x_{d_Z-1},r) \;\equiv\; \tfrac{1}{2}\!\left[x_{d_Z-1}+x_1- r + \sqrt{(x_{d_Z-1}+x_1-r)^2 + 4\, x_1\, r} \right]. \end{align*}
proofSee Appendix (ref).

With treatment effect heterogeneity, these limits no longer obtain. We record the result as a lemma.

lemmap{ROB-HTE}[Robustness, Heterogeneous Treatment Effects] Let \(\beta=\beta_0\) denote the TSLS hypothesis. Let \(\theta_{\mathrm{AR}} \equiv \mu^2\nu^2/(h^2+\omega^2)\), and let \(\zeta\sim \mathfrak{N}(0,1)\), \(\zeta^2_{1} \sim \mathfrak{C}^2_{1}(\omega^2\theta_{\mathrm{AR}}/h^2)\), and \(\zeta^2_{d_Z-1} \sim \mathfrak{C}^2_{d_Z-1}(\mu^2(h^2+\omega^2)/h^2)\) denote three random variables, where \(\zeta \indep \zeta^2_1 \indep \zeta^2_{d_Z-1}\). Then, as \(\sqrt{n}\tau_n=\tau\), \(n\to\infty\), \begin{enumerate} • \(d_Z\mathrm{AR}(\beta_0)\convd \mathfrak{C}^2_{d_Z}(\theta_{\mathrm{AR}})\)\(d_Z\mathrm{LM}(\beta_0)\convd (\Upsilon+\zeta)^2\), with \( \Upsilon^2 \;\sim\; \theta_{\mathrm{AR}}\cdot{\zeta^2_1}/({\zeta^2_1 + \zeta^2_{d_Z-1}})\)\(\mathrm{CLR}(\beta_0) \convd m\bigl((\Upsilon+\zeta)^2,\;\mathfrak{C}^2_{d_Z}(\theta_{\mathrm{AR}}) - (\Upsilon+\zeta)^2,\;\lVert \hat r(\beta_0)\rVert^2 \bigr)\) \end{enumerate} Let \(c^T_{1-\alpha}\) denote the \((1-\alpha)^{th}\) quantile of the constant-treatment-effects limiting distribution of test statistic \(T\) and assume \(\omega^2 >0\). Then, for \(\nu^2>0\), \[ \limsup_{n\to\infty}\,\P\left[\mathrm{T}(\beta_0)> c^{\mathrm{T}}_{1-\alpha}\right] > \alpha \] where \(T\in\{\mathrm{AR},\mathrm{LM},\mathrm{CLR}\}\). Further, \(\lim_{\mu^2\to\infty}\,\limsup_{n\to\infty}\,\P[\mathrm{T}(\beta_0)> c^{\mathrm{T}}_{1-\alpha}] =1\).
proofSee Appendix (ref).

In (ref), we plot the size of the AR, LM and CLR tests,

figure[figure omitted — 1,625 chars of source]

if used as tests for the TSLS hypothesis, \(\beta=\beta_0\), in a Monte Carlo simulation, with and without treatment effect heterogeneity. We simulate by drawing realizations of \(\hat \tau\) from its limiting distribution, with known variance, consistent with constant treatment effects (panels (a), (c) and (e)), and heterogeneous treatment effects (panels (b), (d) and (f)), and compute the rejection rate under the null. The x-axis is enumerated in terms of the \(d_Z\)-scaled concentration parameter, \(\mu^2/d_Z\). The associated first-stage F-statistic cutoffs that reject this concentration parameter at level \(\alpha=0.05\) are indicated in parentheses.

In panels (a), (c) and (e) of (ref), we see that all three tests retain size under constant treatment effects, irrespective of the level of endogeneity. In panels (b), (d) and (f), we see that under heterogeneous treatment effects, the tests reject more often than their nominal level indicates, and are thus invalid as tests for the TSLS hypothesis, with size distortion increasing in instrument strength. With sufficiently strong instruments and non-zero endogeneity, the tests always reject.

Valid Inference on the TSLS Estimand with Weak Instruments

We now turn to deriving a test for the TSLS hypothesis that is valid with weak instruments and treatment effect heterogeneity. We start by deriving the likelihood ratio statistic for the TSLS hypothesis. We then proceed to derive its strong- and weak-instrument limiting distributions. We conclude by showing that we can perform uniformly valid inference by combining the test with a pretest for the nuisance parameter in a two-step procedure without meaningful loss of power.

\paragraph*{The TLR Statistic.}

Consider the asymptotic log-likelihood of \(\tau_n\) given \(\hat \tau\), given as,

align[align omitted — 128 chars of source]

where \(L\) is a constant. Unconstrained, the log-likelihood is maximized by \(\hat \tau\). Under the TSLS constraint, the maximized log-likelihood is a quadratic program with a quadratic constraint. Considering their difference, we get the likelihood ratio statistic for the TSLS hypothesis. The implied quadratic program has a known solution, due to sternIndefiniteTrustRegion1995. We record the result as a proposition.

propositionp{TLR}[TSLS Likelihood Ratio Statistic] The asymptotic likelihood ratio statistic for the hypothesis \(\beta=\beta_0\), when scaled by \(d^{-1}_Z\), is given as \begin{align*} \mathrm{TLR}(\beta_0) \equiv \min_{\tau_n \in \R^{2d_Z}}\,d_Z^{-1}n(\hat \tau -\tau_n)'\hat\Sigma^{-1}(\hat \tau - \tau_n) \qquad s.t. \qquad \tau_n'\,\hat\Gamma(\beta_0) \,\tau_n = 0 \end{align*} where \(\hat\Gamma(\beta_0)\) is a consistent estimator for \(\Gamma(\beta_0)\). Let \(\hat \Lambda(\beta_0)\) denote a consistent estimator for \(\Lambda(\beta_0)\). The program has a unique solution, \begin{align*} \mathrm{TLR}(\beta_0) = d_Z^{-1}\sum_{j=1}^{2d_Z} \hat Q_j^2\!\left(\frac{\lambda^*\hat\kappa_j}{1+\lambda^*\hat\kappa_j}\right)^{\!2} \quads.t.\quad \sum_{j=1}^{2d_Z}\frac{\hat\kappa_j \hat Q_j^2}{(1+\lambda^*\hat\kappa_j)^2} = 0 \end{align*} where \(\{\hat \kappa_j\}_{j=1}^{2d_Z}\) denote the eigenvalues of \(\hat \Lambda(\beta_0)\), \(\hat Q_j\) the \(j^{\mathrm{th}}\) element of \(\hat Q\) and \(\lambda^*\in(-(\max_{j:\hat\kappa_j>0}\hat\kappa_j)^{-1},-(\min_{j:\hat\kappa_j<0}\hat\kappa_j)^{-1})\) is implicitly defined and unique on this interval.
proofSee Appendix (ref).

We scale by \(d_Z^{-1}\) to simplify comparison between settings with different numbers of instruments. Observe that the above result only relies on the joint asymptotic normality of the reduced form and first stage, and does not require any further assumptions.

Under maintained assumptions, the TLR statistic has a limiting distribution that is well-behaved with strong instruments and depends on two nuisance parameters when the instruments are weak. We record the result.

propositionp{LD}[TLR Limiting Distribution] Let \(S\) be defined such that \(d_Z\hat S \convd S\). Under Assumptions (ref), (ref) and (ref), the TLR statistic obtains the following limiting distribution. \begin{enumerate} • When \(\tau_n\) is fixed and \(n\to\infty\), \(d_Z\,\mathrm{TLR}(\beta_0)\convd\mathfrak{C}^2_1\). • When \(\sqrt n\,\tau_n=\tau\) is fixed and \(n\to\infty\), \[ d_Z\,\mathrm{TLR}(\beta_0)\;\convd\;\tfrac{1}{2}\bigl(\sqrt{(1+\varrho)\,S_+}-\sqrt{(1-\varrho)\,S_-}\bigr)^{\!2}, \] where \(S_+\sim\mathfrak{C}^2_{d_Z}((1-\varrho)\,\xi/2)\), \(S_-\sim\mathfrak{C}^2_{d_Z}((1+\varrho)\,\xi/2)\), \(S_+ \indep S_-\), and \(S_+ + S_-=S\). \end{enumerate}
proofSee Appendix (ref).

The weak-instrument limiting distribution obtains the strong-instrument limit as \(\xi\to \infty\). We record this, and two other useful boundary results, as a proposition.

propositionp{BD}[Boundary Case Distributions] Consider the weak-instrument regime, such that \(\tau_n\sqrt{n}=\tau\) constant as \(n\to\infty\). Then, \begin{enumerate} • As \(\xi\to \infty\), we have \(d_Z\mathrm{TLR}(\beta_0) \convd \mathfrak{C}^2_1\). • As \(\varrho\to \pm 1\), we have \(d_Z\mathrm{TLR}(\beta_0) \convd \mathfrak{C}^2_{d_Z}\). • At \(\xi= 0\), \(\varrho=0\), we have \(d_Z\mathrm{TLR}(\beta_0) \convd \left(\sqrt{\mathfrak{C}^2_{d_Z}} -\sqrt{\mathfrak{C}^2_{d_Z}}\right)^2 \,\big/\,2\). \end{enumerate}
proofSee Appendix (ref).

The TLR statistic is biased, in the sense that its value is minimized away from the truth, i.e. for \(\beta\neq\beta_0\). In the following, we show that we can construct a near-unbiased test by recentering the TLR statistic.

\paragraph*{A recentered TLR statistic.} In order to recenter the TLR statistic, it is necessary to first consider the signed root of the TLR statistic. The signed-root TLR statistic can be used to test the one-sided hypothesis, \(\beta \leq \beta_0\). We record its definition and limiting distribution as a lemma.

lemmap{SLR}[Signed-root TLR Statistic] Denote the signed root of the TLR statistic, \begin{align*} {\mathrm{SLR}}(\beta_0) \;\equiv\; \operatorname{sign}\!\left(\sum_{j=1}^{2d_Z}\hat\kappa_j\,\hat Q_j^2\right)\sqrt{\mathrm{TLR}(\beta_0)}. \end{align*} For \(\sqrt{n}\tau_n=\tau\), as \(n\to\infty\), the limiting distribution is given as, \begin{align*} {\mathrm{SLR}}(\beta_0) \convd \tfrac{1}{\sqrt{2\,d_Z}}\left(\sqrt{(1+\varrho)\,S_+}\,-\,\sqrt{(1-\varrho)\,S_-}\right) . \end{align*} Denote this limiting distribution \( T(\varrho,\xi)\). The distribution has non-zero mean for \(\varrho\neq 0\).
proofSee Appendix (ref).

The bias of the TLR statistic is associated with the non-zero mean of its signed root. We can reduce this bias by subtracting an estimate of this mean. In order to take into account finite-sample deviations from exact eigenvalue homogeneity, we use a parametric bootstrap estimator. We record the definition of the recentered TLR statistic.

definitionp{RTLR}[Recentered TSLS Likelihood-Ratio Statistic] Let \(\hat q^*\) denote the constrained minimizer of the TLR statistic, \(\tilde Q^b \sim\mathfrak{N}(\hat q^*,\,\boldsymbol{\iota}_{2d_Z})\) a parametric bootstrap draw from its limiting distribution, where \(\boldsymbol{\iota}_{2d_Z}\) denotes the identity matrix, and let the bootstrapped TLR statistic be given as, \begin{align*} \mathrm{TLR}^b(\beta_0) = d_Z^{-1}\sum_{j=1}^{2d_Z} (\tilde Q_j^b)^2\!\left(\frac{\tilde\lambda^*\hat\kappa_j}{1+\tilde\lambda^*\hat\kappa_j}\right)^{\!2} \quads.t.\quad \sum_{j=1}^{2d_Z}\frac{\hat\kappa_j (\tilde Q_j^b)^2}{(1+\tilde\lambda^*\hat\kappa_j)^2} = 0. \end{align*} Let \(\hat b(\hat q^*, \hat \kappa)\) denote parametric bootstrap estimators of the asymptotic mean of the signed-root TLR statistic given the vector of eigenvalues, \(\hat\kappa\). Define the recentered TSLS likelihood-ratio statistic as \begin{align*} \mathrm{RTLR}(\beta_0) \equiv {({\mathrm{SLR}}(\beta_0) - \hat b(\hat q^*, \hat \kappa))^2}. \end{align*}

Recentering does not obtain a pivotal statistic. However, simulations suggest that it is approximately centered across large parts of the parameter space. As we show in the following, it is possible to leverage both the raw and recentered TLR statistic to construct uniformly valid tests for the TSLS hypothesis.

\paragraph*{Robust TLR Inference.}

A test that is valid in both the strong- and weak-instrument regimes must be valid for all realizations of the nuisance parameters, \((\varrho,\xi)\). In the following, we introduce a two-step procedure that uses the fact that we can consistently estimate \(\varrho\) under the null, and that \(\hat S\) carries information about \(\xi\), to produce uniformly valid inference, in the spirit of bergerValuesMaximizedConfidence1994.

The statistic \(\hat S\) is such that \(d_Z\hat S\) is distributed asymptotically non-central chi-squared with non-centrality parameter \(\xi\). It follows that we can use \(\hat S\) to produce a level \((1-\alpha_1)\) confidence interval for \(\xi\). We can then use the worst-case \((1-\alpha_2)^{\mathrm{th}}\) quantile in this interval as critical value for the test. If \(\alpha_1+\alpha_2 = \alpha\), this two-step procedure is a valid test for \(\beta=\beta_0\) at level \(\alpha\). We record the result.

propositionp{TS}[Two-step Inference] Let \(\hat \Xi_{1-\alpha_1}\) denote a \(1-\alpha_1\) confidence interval for \(\xi\) obtained by inverting \(\hat S\), and let \(\alpha_2 \equiv \alpha -\alpha_1 \). Let \(\hat T(\beta_0)\) denote the raw, signed-root or recentered TLR statistic, and let \(c_{1-\alpha_2}(\varrho,\xi) \) denote the \((1-\alpha_2)^{th}\) {quantile} of its limiting distribution. We then have: \begin{align*} \limsup_{n\to\infty}\,\sup_{\varrho,\xi}\,\P\!\left(\hat T(\beta_0) > \max_{\xi \in \hat\Xi_{1-\alpha_1}} c_{1-\alpha_2}(\hat\varrho,\xi)\right) \;\leq\; \alpha_1 + \alpha_2 \;=\; \alpha. \end{align*} It follows that \(\hat T(\beta_0) > \max_{\xi \in \hat\Xi_{1-\alpha_1}} c_{1-\alpha_2}(\hat\varrho,\xi)\) is a {uniformly} valid test for \(\beta=\beta_0\).
proofSee Appendix (ref).

When the first-stage F-statistic is sufficiently small, tests based on the TLR statistic fail to reject any \(\beta_0\) far from the origin. It follows that associated confidence intervals have infinite length. We record the result as a lemma.

lemmap{UC}[Unbounded Confidence Interval] Let \(\hat F \equiv \lVert (\hat\Sigma_{\hat\gamma}/n)^{-1/2}\hat\gamma\rVert^2/d_Z\) denote the Wald test statistic for the hypothesis that \(\gamma_n=0\). The two-step TLR test fails to reject \(\beta_0\) for all sufficiently large \(\lvert\beta_0\rvert\) if and only if \[ d_Z\hat F \;\leq\; {\chi^2_{d_Z,1-\alpha_2}}. \] Equivalently, confidence intervals formed by test inversion have infinite length.
proofSee Appendix (ref).

It is instructive to understand how much the limiting distribution of the TLR and recentered TLR statistics differ for different values of \((\varrho,\xi)\) and \(d_Z\). We study this in Monte Carlo simulations. As the limiting distribution of the TLR statistic is symmetric in \(\varrho\), it suffices to consider positive \(\varrho\).

In (ref), we plot the \(95^{\mathrm{th}}\) quantile of the distribution of the TLR statistic and the recentered TLR statistic for different values of \((\varrho, \xi)\) and \(d_Z\).

figure[figure omitted — 760 chars of source]

We make two observations. First, the \(95^{\mathrm{th}}\) quantile of the TLR statistic meaningfully overshoots the \(95^{\mathrm{th}}\) quantile of the chi-squared distribution with one degree of freedom for a wide range of values of \((\varrho,\xi)\). Second, the \(95^{\mathrm{th}}\) quantile of the recentered TLR statistic is consistently located either at or below the \(95^{\mathrm{th}}\) quantile of the chi-squared distribution with one degree of freedom.

From this, we draw two conclusions. First, the relative dominance of the \(95^{\mathrm{th}}\) quantile of the chi-squared distribution over that of the recentered TLR statistic suggest that the combination of the former used as critical value in a test using the recentered TLR statistic will produce an approximately uniformly valid test, albeit conservative for some values of \((\varrho,\xi)\), and small \(d_Z\). Second, a two-step test with a small \(\alpha_1\approx 0\) and little if any change in critical values will produce an exactly uniformly valid test. Given the large-\(\xi\) distribution of the TLR statistic, such a test will numerically recover the Wald test in the strong-instrument limit. By computing conditional critical values, we can increase the power of this test in the weak-instrument regime.

The two-step test from (ref) requires the choice of a first-step level \(\alpha_1\). In order to achieve the strong-instrument Wald limit, it is preferable to choose this to be small. To this end, we let \(\alpha_1=10^{-5}\). However, the performance of the test is relatively invariant to this choice in simulations, and one can equivalently choose e.g. \(\alpha_1=10^{-10}\), giving slightly different second-step conditional critical values.

(ref) plots conditional critical values for a level \(\alpha=0.05\) two-step TLR test (dashed) and two-step recentered TLR test (solid),

figure[figure omitted — 960 chars of source]

for different values of \((\hat \varrho, \hat S)\) and \(d_Z\). We see that conditional critical values for the TLR test differ considerably from its strong-instrument limit. The conditional critical values of the recentered TLR test are either at or below the \(95^{\mathrm{th}}\) quantile of the chi-squared distribution with one degree of freedom for most values of \(\hat \varrho\) and \(\hat S\). In practice, such conditional critical values can either be tabulated, as done for the “VtF” test introduced by leeWhatWhenYou2023, or computed on the fly.

Lastly, it is of interest to understand the power properties of the two-step TLR and recentered tests relative to the Wald test constructed using the TSLS estimator. While this comparison is not prima facie reasonable, as the TSLS Wald test is not valid with weak instruments, it is of interest to understand the extent to which the two-step TLR and recentered TLR tests replicate the power properties of the TSLS Wald test with strong and moderately strong instruments.

In (ref) we plot the power of the TLR and recentered TLR two-step tests in simulations,

figure[figure omitted — 842 chars of source]

and compare them to the power of the TSLS Wald test, for different values of the concentration parameter, \(\mu^2\), and \(\omega\), fixing \(\nu^2=1\), \(h=0.1\) and \(d_Z=5\). We let \(\varrho_{\mathrm{true}}=\omega/\sqrt{h^2+\omega^2}\) denote the true value of \(\varrho\) off the null. For values of the concentration parameter, \(\mu^2\), we indicate the value of the first-stage F-statistic that would reject the hypothesis \(\mu^2\leq\bar \mu^2\) at level \(\alpha=0.05\) in the panel header. Recall from (ref) that \(\xi\) is a deterministic function of \(\omega,h^2,\nu^2\) and \(\mu^2\), and hence pinned down by these parameters. The x-axis is scaled relative to the value of the nuisance parameters, in order for curves to take comparable shapes across panels.

Considering (ref), three observations become apparent. First, the power curve of the recentered TLR statistic is close to being symmetric, and the test is close to being unbiased, for most of the parameter space. The exception is for extreme realizations of \(\varrho_{\mathrm{true}}\) and \(\xi\). Second, we see that unlike the TSLS Wald test, both the TLR and recentered TLR tests reject no more than 5% of the time at the null, \(\beta=\beta_0\). The TSLS Wald test, on the other hand, overrejects under the null at small \(\xi\) and large \(\varrho\). Third, we see that as the instrument set grows stronger, the power curves of the three tests converge to the same limit. This suggests that the power of the TLR and recentered TLR tests asymptote to the power of the Wald test.

Conclusion

In this paper we have shown that existing weak-instrument robust tests are invalid if used as tests for the TSLS estimand in a model with treatment effect heterogeneity and more than one instrument. In particular, we have shown that the AR, Kleibergen--Moreira LM and Moreira CLR tests can overreject at any level of instrument strength, and that the degree of overrejection is increasing as the instrument set grows stronger.

We have further shown that there exists a test statistic that can be leveraged to produce valid inference with treatment effect heterogeneity. In particular, we have constructed the TSLS likelihood-ratio (TLR) statistic for hypotheses on the TSLS estimand, and derived its limiting distribution. We have shown that when combined with a pretest in a two-step procedure, the test obtains uniform validity. Choosing a small first-step level, we have shown in simulations that the test inherits the properties of the Wald test when the instrument set is strong.

thebibliography\bibitem[Anderson and Rubin, 1949]{andersonEstimationParametersSingle1949} Anderson, T. W. and Rubin, H. (1949). \newblock Estimation of the parameters of a single equation in a complete system of stochastic equations. \newblock {\em The Annals of Mathematical Statistics}, 20(1):46--63. \bibitem[Andrews et al., 2006]{andrewsOptimalTwoSidedInvariant2006} Andrews, D. W. K., Moreira, M. J., and Stock, J. H. (2006). \newblock Optimal two-sided invariant similar tests for instrumental variables regression. \newblock {\em Econometrica}, 74(3):715--752. \bibitem[Andrews et al., 2019]{andrewsWeakInstrumentsInstrumental2019} Andrews, I., Stock, J. H., and Sun, L. (2019). \newblock Weak instruments in instrumental variables regression: Theory and practice. \newblock {\em Annual Review of Economics}, 11:727--753. \bibitem[Angrist et al., 2000]{angristInterpretationInstrumentalVariables2000} Angrist, J. D., Graddy, K., and Imbens, G. W. (2000). \newblock The interpretation of instrumental variables estimators in simultaneous equations models with an application to the demand for fish. \newblock {\em The Review of Economic Studies}, 67(3):499--527. \bibitem[Angrist and Imbens, 1995]{angristTwoStageLeastSquares1995} Angrist, J. D. and Imbens, G. W. (1995). \newblock Two-stage least squares estimation of average causal effects in models with variable treatment intensity. \newblock {\em Journal of the American Statistical Association}, 90(430):431--442. \bibitem[Berger and Boos, 1994]{bergerValuesMaximizedConfidence1994} Berger, R. L. and Boos, D. D. (1994). \newblock P values maximized over a confidence set for the nuisance parameter. \newblock {\em Journal of the American Statistical Association}, 89(427):1012--1016. \bibitem[Bound et al., 1995]{boundProblemsInstrumentalVariables1995} Bound, J., Jaeger, D. A., and Baker, R. M. (1995). \newblock Problems with instrumental variables estimation when the correlation between the instruments and the endogenous explanatory variable is weak. \newblock {\em Journal of the American Statistical Association}, 90(430):443--450. \bibitem[Chernozhukov et al., 2013]{chernozhukovIntersectionBoundsEstimation2013} Chernozhukov, V., Lee, S., and Rosen, A. M. (2013). \newblock Intersection bounds: Estimation and inference. \newblock {\em Econometrica}, 81(2):667--737. \bibitem[Chyn et al., 2025]{chynExaminerJudgeDesigns2025} Chyn, E., Frandsen, B., and Leslie, E. (2025). \newblock Examiner and judge designs in economics: A practitioner's guide. \newblock {\em Journal of Economic Literature}, 63(2):401--439. \bibitem[Evdokimov and Koles{\'a}r, 2018]{evdokimovInferenceInstrumentalVariables2018} Evdokimov, K. S. and Koles{\'a}r, M. (2018). \newblock Inference in instrumental variables analysis with heterogeneous treatment effects. \newblock Working Paper. \bibitem[Finlay and Magnusson, 2009]{finlayImplementingWeakInstrumentRobust2009} Finlay, K. and Magnusson, L. M. (2009). \newblock Implementing weak-instrument robust tests for a general class of instrumental-variables models. \newblock {\em The Stata Journal}, 9(3):398--421. \bibitem[Fuller, 1977]{fullerPropertiesModificationLimited1977} Fuller, W. A. (1977). \newblock Some properties of a modification of the limited information estimator. \newblock {\em Econometrica}, 45(4):939--953. \bibitem[{Goldsmith-Pinkham} et al., 2025]{goldsmith-pinkhamLeniencyDesignsOperators2025} {Goldsmith-Pinkham}, P., Hull, P., and Koles{\'a}r, M. (2025). \newblock Leniency designs: An operator's manual. \newblock Working Paper arXiv:2511.03572, arXiv. \bibitem[Hansen, 1982]{hansenLargeSampleProperties1982} Hansen, L. P. (1982). \newblock Large sample properties of generalized method of moments estimators. \newblock {\em Econometrica}, 50(4):1029--1054. \bibitem[Hansen, 2005]{hansenTestSuperiorPredictive2005} Hansen, P. R. (2005). \newblock A test for superior predictive ability. \newblock {\em Journal of Business & Economic Statistics}, 23(4):365--380. \bibitem[Heckman et al., 2006]{heckmanUnderstandingInstrumentalVariables2006} Heckman, J. J., Urzua, S., and Vytlacil, E. (2006). \newblock Understanding instrumental variables in models with essential heterogeneity. \newblock {\em The Review of Economics and Statistics}, 88(3):389--432. \bibitem[Imbens and Angrist, 1994]{imbensIdentificationEstimationLocal1994b} Imbens, G. W. and Angrist, J. D. (1994). \newblock Identification and estimation of local average treatment effects. \newblock {\em Econometrica}, 62(2):467--475. \bibitem[Kleibergen, 2002]{kleibergenPivotalStatisticsTesting2002} Kleibergen, F. (2002). \newblock Pivotal statistics for testing structural parameters in instrumental variables regression. \newblock {\em Econometrica}, 70(5):1781--1803. \bibitem[Koles{\'a}r, 2013]{kolesarEstimationInstrumentalVariables2013} Koles{\'a}r, M. (2013). \newblock Estimation in an instrumental variables model with treatment effect heterogeneity. \newblock Working Paper. \bibitem[Lee et al., 2022]{leeValidTRatioInference2022} Lee, D. S., McCrary, J., Moreira, M. J., and Porter, J. R. (2022). \newblock Valid t-ratio inference for IV. \newblock {\em American Economic Review}, 112(10):3260--3290. \bibitem[Lee et al., 2023]{leeWhatWhenYou2023} Lee, D. S., McCrary, J., Moreira, M. J., Porter, J. R., and Yap, L. (2023). \newblock What to do when you can't use '1.96' confidence intervals for IV. \newblock NBER Working Paper 31893. \bibitem[Loh, 1985]{lohNewMethodTesting1985} Loh, W.-Y. (1985). \newblock A new method for testing separate families of hypotheses. \newblock {\em Journal of the American Statistical Association}, 80(390):362--368. \bibitem[McCloskey, 2017]{mccloskeyBonferronibasedSizecorrectionNonstandard2017} McCloskey, A. (2017). \newblock Bonferroni-based size-correction for nonstandard testing problems. \newblock {\em Journal of Econometrics}, 200(1):17--35. \bibitem[Mogstad and Torgovitsky, 2024]{mogstadInstrumentalVariablesUnobserved2024} Mogstad, M. and Torgovitsky, A. (2024). \newblock Instrumental variables with unobserved heterogeneity in treatment effects. \newblock NBER Working Paper 32927. \bibitem[Mogstad et al., 2021]{mogstadCausalInterpretationTwoStage2021} Mogstad, M., Torgovitsky, A., and Walters, C. R. (2021). \newblock The causal interpretation of two-stage least squares with multiple instrumental variables. \newblock {\em American Economic Review}, 111(11):3663--3698. \bibitem[Moreira, 2003]{moreiraConditionalLikelihoodRatio2003} Moreira, M. J. (2003). \newblock A conditional likelihood ratio test for structural models. \newblock {\em Econometrica}, 71(4):1027--1048. \bibitem[Moreira, 2009]{moreiraTestsCorrectSize2009} Moreira, M. J. (2009). \newblock Tests with correct size when instruments can be arbitrarily weak. \newblock {\em Journal of Econometrics}, 152(2):131--140. \bibitem[Nelson and Startz, 1990a]{nelsonDistributionInstrumentalVariables1990} Nelson, C. R. and Startz, R. (1990a). \newblock The distribution of the instrumental variables estimator and its t-ratio when the instrument is a poor one. \newblock {\em The Journal of Business}, 63(1):S125--S140. \bibitem[Nelson and Startz, 1990b]{nelsonFurtherResultsExact1990} Nelson, C. R. and Startz, R. (1990b). \newblock Some further results on the exact small sample properties of the instrumental variable estimator. \newblock {\em Econometrica}, 58(4):967--976. \bibitem[Romano et al., 2014]{romanoPracticalTwoStepMethod2014} Romano, J. P., Shaikh, A. M., and Wolf, M. (2014). \newblock A practical two-step method for testing moment inequalities. \newblock {\em Econometrica}, 82(5):1979--2002. \bibitem[Sargan, 1958]{sarganEstimationEconomicRelationships1958} Sargan, J. D. (1958). \newblock The estimation of economic relationships using instrumental variables. \newblock {\em Econometrica}, 26(3):393--415. \bibitem[Serfling, 1980]{ApproximationTheoremsMathematical1980} Serfling, R. J. (1980). \newblock {\em Approximation theorems of mathematical statistics}. \newblock Wiley, New York, 1st edition. \bibitem[Sigstad, 2026]{sigstadMonotonicityJudgesEvidence2026} Sigstad, H. (2026). \newblock Monotonicity among judges: Evidence from judicial panels and consequences for judge IV designs. \newblock {\em American Economic Review}, 116(1):189--208. \bibitem[Silvapulle, 1996]{silvapulleTestPresenceNuisance1996} Silvapulle, M. J. (1996). \newblock A test in the presence of nuisance parameters. \newblock {\em Journal of the American Statistical Association}, 91(436):1690--1693. \bibitem[Staiger and Stock, 1997]{staigerInstrumentalVariablesRegression1997} Staiger, D. and Stock, J. H. (1997). \newblock Instrumental variables regression with weak instruments. \newblock {\em Econometrica}, 65(3):557--586. \bibitem[Stern and Wolkowicz, 1995]{sternIndefiniteTrustRegion1995} Stern, R. J. and Wolkowicz, H. (1995). \newblock Indefinite trust region subproblems and nonsymmetric eigenvalue perturbations. \newblock {\em SIAM Journal on Optimization}, 5(2):286--313. \bibitem[{Van de Sijpe} and Windmeijer, 2023]{vandesijpePowerConditionalLikelihood2023} {Van de Sijpe}, N. and Windmeijer, F. (2023). \newblock On the power of the conditional likelihood ratio and related tests for weak-instrument robust inference. \newblock {\em Journal of Econometrics}, 235(1):82--104. \bibitem[Vytlacil, 2002]{vytlacilIndependenceMonotonicityLatent2002a} Vytlacil, E. (2002). \newblock Independence, monotonicity, and latent index models: An equivalence result. \newblock {\em Econometrica}, 70(1):331--341. \bibitem[Yap, 2025]{yapInferenceManyWeak2025} Yap, L. (2025). \newblock Inference with many weak instruments and heterogeneity. \newblock Working Paper.