Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
58,674 characters · 5 sections · 67 citation commands
Inference on the TSLS Estimand with Weak Instruments and Treatment Effect Heterogeneity
Consider the problem of performing inference on a parameter of interest,
when this parameter is the coefficient in an overidentified instrumental variables regression with one endogenous regressor and a potentially weak set of instruments, known as the two-stage least squares (TSLS) estimand,
The TSLS estimand is a weighted sum of the elements of the reduced-form coefficient vector, \(\delta\), weighted by the elements of the first-stage coefficient vector, \(\gamma\), and the population instrument Gram matrix, \(\Sigma_{ZZ}\). It is a common parameter of interest in economics, political science, sociology, and other fields. Under certain assumptions, it obtains a causal interpretation as a positively weighted average of local average treatment effects, see e.g. angristTwoStageLeastSquares1995 and mogstadCausalInterpretationTwoStage2021.
In the presence of weak instruments, the commonly applied Wald- or t-test is known to suffer from size distortion. This problem is well-studied, and in the setting of a linear model with constant treatment effects (CTE), the AR test (andersonEstimationParametersSingle1949), Kleibergen--Moreira Lagrange multiplier (LM) test (kleibergenPivotalStatisticsTesting2002, moreiraConditionalLikelihoodRatio2003) and conditional likelihood ratio (CLR) test (moreiraConditionalLikelihoodRatio2003) are known to deliver valid inference.\footnote{See finlayImplementingWeakInstrumentRobust2009 and andrewsWeakInstrumentsInstrumental2019 for parsimonious expositions.} In a model with treatment effect heterogeneity, however, these are not valid tests for hypotheses on the TSLS estimand. As pointed out by yapInferenceManyWeak2025, there exists no test proven to be valid (i) in the presence of weak instruments and (ii) treatment effect heterogeneity which (iii) targets a parameter interpretable as a weighted average of local average treatment effects under commonly maintained assumptions when (iv) considering limit experiments that fix the number of instruments.
To understand why standard robust tests are not valid for hypotheses on the TSLS estimand under treatment effect heterogeneity, observe that the AR test performs inference on the hypothesis \(\beta=\beta_0\) by instead performing inference on a set of linear restrictions,
where the number of restrictions is equal to the number of instruments. The set of restrictions can be rewritten as a composite hypothesis on the following form,
As is shown, e.g., by moreiraConditionalLikelihoodRatio2003, the AR test statistic can be written as the sum of the LM test statistic and the sarganEstimationEconomicRelationships1958--hansenLargeSampleProperties1982 $J$ test statistic. Maintaining CTE, the $J$ test is a test of instrument exogeneity. Maintaining exogeneity, but allowing for treatment effect heterogeneity, the $J$ test is essentially a test of the CTE assumption. The CLR and LM tests are derived from the AR restrictions, testing \(\beta=\beta_0\) while maintaining the assumption of CTE. As a consequence, they are invalid in the presence of treatment effect heterogeneity when considered as tests for the TSLS hypothesis.
This paper proposes a test that is valid with weak instruments and treatment effect heterogeneity. The test is derived by first observing that a hypothesis about the TSLS estimand, \(\beta=\beta_0\), can be reformulated into a quadratic constraint by multiplying by the denominator on both sides and rearranging,
While the restriction is similar in form to the AR restriction in that it does not involve division by a potentially small number, it is a single scalar restriction, and does not impose or test any ancillary assumptions, such as CTE.
The estimators of the reduced-form and first-stage coefficients are jointly asymptotically normal, with a variance-covariance matrix that is consistently estimable. This admits a likelihood ratio statistic, formulated as a quadratic optimization program with a quadratic constraint. We call this the TSLS likelihood ratio (TLR) statistic. While the program does not have a closed-form solution, it is a special case of the generalized trust region subproblem, which has a known and unique solution, see e.g. sternIndefiniteTrustRegion1995. When instruments are strong, the TLR statistic is distributed asymptotically chi-squared with one degree of freedom. Combined with the fact that likelihood ratio tests are asymptotically equivalent to their associated Wald test under the null and against local alternatives (ApproximationTheoremsMathematical1980), this makes the implied test numerically indistinguishable from the Wald test in the strong-instrument regime.
When the instrument set is weak, the limiting distribution of the TLR statistic depends on two nuisance parameters. The first nuisance parameter is a structural endogeneity parameter, capturing both selection on gains and levels. Similarly to the one-instrument case (see e.g. vandesijpePowerConditionalLikelihood2023 and leeWhatWhenYou2023) it is consistently estimable under the null. The second nuisance parameter captures the overall identifying power of the instrument set. It is closely related to, but distinct from, the {concentration parameter} often used to gauge instrument strength in the literature. While it cannot be consistently estimated, it is associated with a sample statistic similar to, but distinct from, the first-stage F-statistic. This allows for a two-step procedure similar to the one discussed by bergerValuesMaximizedConfidence1994, and suggested by staigerInstrumentalVariablesRegression1997 for valid Wald inference in the linear constant effects model. Unlike their approach, we leverage more of the information contained in the sample to produce a test that is less conservative. In particular, we can set the first-step level so low as to numerically asymptote to the Wald test with strong instruments without substantial loss of power.
\paragraph*{Literature.}
The paper contributes to the literature on inference with weak instruments. Following the seminal contributions of nelsonDistributionInstrumentalVariables1990,nelsonFurtherResultsExact1990, boundProblemsInstrumentalVariables1995 and staigerInstrumentalVariablesRegression1997, a substantial number of papers have been devoted to understanding and resolving the finite-sample shortcomings of common procedures for inference and estimation with weak instruments. For a comprehensive review, see andrewsWeakInstrumentsInstrumental2019. In the context of constant treatment effects, andersonEstimationParametersSingle1949 provided the first test with uniform validity, shown by moreiraTestsCorrectSize2009 to be uniformly most powerful unbiased with one instrument. Later, kleibergenPivotalStatisticsTesting2002 and moreiraConditionalLikelihoodRatio2003 developed the Kleibergen--Moreira LM and Moreira CLR tests, with better power properties. andrewsOptimalTwoSidedInvariant2006 showed that the latter of these tests is uniformly most powerful unbiased under homoskedasticity. Recently, progress has been made on deriving more intuitive tests (leeValidTRatioInference2022) and tests with more intentional direction of power (leeWhatWhenYou2023) in the one-instrument case, both building on the two-step correction procedure suggested by staigerInstrumentalVariablesRegression1997, which uses the first-stage F-statistic and a Bonferroni correction to produce valid Wald inference.
The paper also contributes to the study of instrumental variables in policy evaluation and models with heterogeneous treatment effects. For a review of this literature, see mogstadInstrumentalVariablesUnobserved2024. While the TSLS estimand is one of many potential ways of aggregating reduced-form estimates when overidentified, following imbensIdentificationEstimationLocal1994b, angristTwoStageLeastSquares1995 and angristInterpretationInstrumentalVariables2000, it has been a common parameter of interest for empirical practitioners. Under certain assumptions, it obtains a weighted average of local average treatment effects. As is shown by kolesarEstimationInstrumentalVariables2013, this is not always the case for estimands associated with other estimators in the literature, such as the LIML estimator introduced by andersonEstimationParametersSingle1949, and the Fuller--$k$ estimator, introduced by fullerPropertiesModificationLimited1977. The methods considered in this paper are robust to violations of the so-called monotonicity (uniformity) assumption, discussed e.g. by heckmanUnderstandingInstrumentalVariables2006, mogstadCausalInterpretationTwoStage2021 and sigstadMonotonicityJudgesEvidence2026, as they do not rely on the causal interpretation of the estimand. While the one-instrument setting is more widespread in the literature, a common setting with more than one instrument is the examiner, judge or leniency design, see chynExaminerJudgeDesigns2025 and goldsmith-pinkhamLeniencyDesignsOperators2025 for reviews. In terms of valid inference in the presence of heterogeneous treatment effects, evdokimovInferenceInstrumentalVariables2018 study the problem in a setting with many instruments, and yapInferenceManyWeak2025 with many weak instruments. To the author's knowledge, this is the first paper to consider the problem of valid inference on the TSLS estimand with a fixed number of weak instruments and potentially heterogeneous treatment effects.
The paper is also related to the literature using two-step procedures to obtain valid inference in the presence of nuisance parameters, see e.g. lohNewMethodTesting1985, bergerValuesMaximizedConfidence1994, silvapulleTestPresenceNuisance1996, hansenTestSuperiorPredictive2005, chernozhukovIntersectionBoundsEstimation2013, romanoPracticalTwoStepMethod2014 and mccloskeyBonferronibasedSizecorrectionNonstandard2017. The two-step test proposed by staigerInstrumentalVariablesRegression1997 is an example of such a testing procedure in the weak-instrument setting.
\paragraph*{Roadmap.} The paper proceeds as follows. In (ref) we set up the linear instrumental variables model with heterogeneous treatment effects and derive a statistic that captures the total identifying power of the instrument set. In (ref) we show that existing weak-instrument robust tests fail as tests for the TSLS hypothesis with treatment effect heterogeneity. In (ref) we introduce the TLR statistic, and derive its limiting distribution. We show that it retains validity when combined with a two-step pretest and study its power properties in the weak- and strong-instrument regime. (ref) concludes.
\paragraph*{Fundamentals.} Let \(\{(Y_i,D_i,Z_i)\}_{i=1}^n\) be an i.i.d. sample, where \(Y_i\) is an outcome, \(D_i\) a treatment, both scalar, and \(Z_i\) a \(d_Z\)-dimensional vector of instruments, with \(d_Z\geq 2\). Assume all variables are demeaned and have finite fourth moments. In the presence of covariates, let \(Y_i\) and \(D_i\) denote variables where their impact is appropriately projected off. We are interested in the linear heterogeneous treatment effects model,
where \(\beta_i\) is the causal effect of \(D_i\) on \(Y_i\), and \(U_i\) is the model residual. Denote the linear regressions of \(Y_i\) on \(Z_i\) (reduced form) and \(D_i\) on \(Z_i\) (first stage) as,
where \(\delta_n\) is the vector of reduced-form coefficients, \(\gamma_n\) the vector of first-stage coefficients, and \(W_i\) and \(V_i\) are regression residuals. Subscripts on \(\delta_n\) and \(\gamma_n\) indicate that we do not restrict the data generating process to be invariant with respect to the sample size \(n\). To analyze weak-instrument properties, we follow staigerInstrumentalVariablesRegression1997 in letting \((\delta_n, \gamma_n)\) drift towards zero at rate \(1/\sqrt{n}\), such that \((\delta_n,\gamma_n) \equiv { (\delta,\gamma)}/{\sqrt{n}}\) for some vectors \((\delta, \gamma)\). To study strong-instruments properties, we hold \((\delta_n, \gamma_n)\) fixed as \(n\to\infty\). Let the TSLS estimand, \(\beta\), be defined as in (ref). It is uniquely pinned down by \((\delta, \gamma)\), hence the parameter of interest does not vary with the sample size. Let \(D_i(z)\) denote the potential treatment of an individual with instrument realization \(Z_i=z\). We maintain the following assumptions.
Assumptions (ref) and (ref) are a subset of the assumptions usually maintained in models with heterogeneous treatment effects, see e.g. angristTwoStageLeastSquares1995 and angristInterpretationInstrumentalVariables2000. No assumption is imposed on the monotonicity (uniformity) of \(D_i(z)\) in \(z\). However, treatment effect heterogeneity requires the following refinement to the staigerInstrumentalVariablesRegression1997 weak-instrument limiting sequence.
Assumption (ref) rules out the case in which an instrument is weak because the share of individuals who are instrumented into more treatment (compliers) and less treatment (defiers) cancel. It is implied by the assumption of monotonicity (uniformity) of \(D_i(z)\) in \(z\), or equivalently by the vytlacilIndependenceMonotonicityLatent2002a single-index model.
\paragraph*{Observed moments.} Let \(\mathfrak{N}(\,\cdot\,,\,\cdot\,)\) denote the Normal distribution, \(\tau_n \equiv \left[
\right]\) the stacked vector of first-stage and reduced-form coefficients and \(\hat \tau\equiv \left[
\right]\) its OLS estimator. While the estimator is also a function of the sample size, for brevity of notation we let this be implicit in the following. The estimator has the following limiting distribution.
Let \(\Sigma_{ZZ}\equiv \E[Z_iZ_i']\) denote the population Gram matrix of \(Z_i\). We can rewrite the constraint defining the TSLS hypothesis from (ref) in terms of \(\tau_n\) as follows.
With constant treatment effects, the first-stage F-statistic summarizes the strength of the instrument set. With treatment effect heterogeneity, we need to consider a statistic that also captures other facets of the identifying power of the instruments. We record this statistic and its limiting distribution as a lemma.
The parameter \(\xi\) is related to the concentration parameter often used to enumerate instrument strength in the literature. We record the relation as a lemma.
Observe that with one instrument or constant treatment effects, we have \(\nu^2=0\).
\paragraph*{Eigenvalue Representation.}
It will be useful to express the data in terms of a particular eigendecomposition. Let \(\Lambda(\beta_0) \equiv \Sigma^{\frac{1}{2}} \Gamma(\beta_0) \Sigma^{\frac{1}{2}}\) denote the rotation of \(\Gamma(\beta_0)\) in terms of \(\Sigma\). Let \(X\) denote the matrix of eigenvectors of \(\Lambda(\beta_0)\), and \(\hat X\) its analog estimator. Define the rotated first-stage and reduced-form estimators and estimands as
While \(\Lambda(\beta_0)\) is indefinite, the block-structure of \(\Gamma(\beta_0)\) and (ref) restrict the sign of, and variation in, its eigenvalues.
Note that \((q, X, \kappa_\pm, \varrho)\) and their estimators depend on the hypothesized \(\beta_0\). For brevity of notation, we suppress this in the following.
The AR, LM and CLR tests are commonly understood to provide inference that is robust to weak instruments. In the following, we show that this is not the case if the tests are interpreted as tests of the TSLS estimand under treatment effect heterogeneity. We start by recording the constant treatment effect limiting distribution of the three statistics, as derived by andersonEstimationParametersSingle1949, kleibergenPivotalStatisticsTesting2002 and moreiraConditionalLikelihoodRatio2003.
With treatment effect heterogeneity, these limits no longer obtain. We record the result as a lemma.
In (ref), we plot the size of the AR, LM and CLR tests,
if used as tests for the TSLS hypothesis, \(\beta=\beta_0\), in a Monte Carlo simulation, with and without treatment effect heterogeneity. We simulate by drawing realizations of \(\hat \tau\) from its limiting distribution, with known variance, consistent with constant treatment effects (panels (a), (c) and (e)), and heterogeneous treatment effects (panels (b), (d) and (f)), and compute the rejection rate under the null. The x-axis is enumerated in terms of the \(d_Z\)-scaled concentration parameter, \(\mu^2/d_Z\). The associated first-stage F-statistic cutoffs that reject this concentration parameter at level \(\alpha=0.05\) are indicated in parentheses.
In panels (a), (c) and (e) of (ref), we see that all three tests retain size under constant treatment effects, irrespective of the level of endogeneity. In panels (b), (d) and (f), we see that under heterogeneous treatment effects, the tests reject more often than their nominal level indicates, and are thus invalid as tests for the TSLS hypothesis, with size distortion increasing in instrument strength. With sufficiently strong instruments and non-zero endogeneity, the tests always reject.
We now turn to deriving a test for the TSLS hypothesis that is valid with weak instruments and treatment effect heterogeneity. We start by deriving the likelihood ratio statistic for the TSLS hypothesis. We then proceed to derive its strong- and weak-instrument limiting distributions. We conclude by showing that we can perform uniformly valid inference by combining the test with a pretest for the nuisance parameter in a two-step procedure without meaningful loss of power.
\paragraph*{The TLR Statistic.}
Consider the asymptotic log-likelihood of \(\tau_n\) given \(\hat \tau\), given as,
where \(L\) is a constant. Unconstrained, the log-likelihood is maximized by \(\hat \tau\). Under the TSLS constraint, the maximized log-likelihood is a quadratic program with a quadratic constraint. Considering their difference, we get the likelihood ratio statistic for the TSLS hypothesis. The implied quadratic program has a known solution, due to sternIndefiniteTrustRegion1995. We record the result as a proposition.
We scale by \(d_Z^{-1}\) to simplify comparison between settings with different numbers of instruments. Observe that the above result only relies on the joint asymptotic normality of the reduced form and first stage, and does not require any further assumptions.
Under maintained assumptions, the TLR statistic has a limiting distribution that is well-behaved with strong instruments and depends on two nuisance parameters when the instruments are weak. We record the result.
The weak-instrument limiting distribution obtains the strong-instrument limit as \(\xi\to \infty\). We record this, and two other useful boundary results, as a proposition.
The TLR statistic is biased, in the sense that its value is minimized away from the truth, i.e. for \(\beta\neq\beta_0\). In the following, we show that we can construct a near-unbiased test by recentering the TLR statistic.
\paragraph*{A recentered TLR statistic.} In order to recenter the TLR statistic, it is necessary to first consider the signed root of the TLR statistic. The signed-root TLR statistic can be used to test the one-sided hypothesis, \(\beta \leq \beta_0\). We record its definition and limiting distribution as a lemma.
The bias of the TLR statistic is associated with the non-zero mean of its signed root. We can reduce this bias by subtracting an estimate of this mean. In order to take into account finite-sample deviations from exact eigenvalue homogeneity, we use a parametric bootstrap estimator. We record the definition of the recentered TLR statistic.
Recentering does not obtain a pivotal statistic. However, simulations suggest that it is approximately centered across large parts of the parameter space. As we show in the following, it is possible to leverage both the raw and recentered TLR statistic to construct uniformly valid tests for the TSLS hypothesis.
\paragraph*{Robust TLR Inference.}
A test that is valid in both the strong- and weak-instrument regimes must be valid for all realizations of the nuisance parameters, \((\varrho,\xi)\). In the following, we introduce a two-step procedure that uses the fact that we can consistently estimate \(\varrho\) under the null, and that \(\hat S\) carries information about \(\xi\), to produce uniformly valid inference, in the spirit of bergerValuesMaximizedConfidence1994.
The statistic \(\hat S\) is such that \(d_Z\hat S\) is distributed asymptotically non-central chi-squared with non-centrality parameter \(\xi\). It follows that we can use \(\hat S\) to produce a level \((1-\alpha_1)\) confidence interval for \(\xi\). We can then use the worst-case \((1-\alpha_2)^{\mathrm{th}}\) quantile in this interval as critical value for the test. If \(\alpha_1+\alpha_2 = \alpha\), this two-step procedure is a valid test for \(\beta=\beta_0\) at level \(\alpha\). We record the result.
When the first-stage F-statistic is sufficiently small, tests based on the TLR statistic fail to reject any \(\beta_0\) far from the origin. It follows that associated confidence intervals have infinite length. We record the result as a lemma.
It is instructive to understand how much the limiting distribution of the TLR and recentered TLR statistics differ for different values of \((\varrho,\xi)\) and \(d_Z\). We study this in Monte Carlo simulations. As the limiting distribution of the TLR statistic is symmetric in \(\varrho\), it suffices to consider positive \(\varrho\).
In (ref), we plot the \(95^{\mathrm{th}}\) quantile of the distribution of the TLR statistic and the recentered TLR statistic for different values of \((\varrho, \xi)\) and \(d_Z\).
We make two observations. First, the \(95^{\mathrm{th}}\) quantile of the TLR statistic meaningfully overshoots the \(95^{\mathrm{th}}\) quantile of the chi-squared distribution with one degree of freedom for a wide range of values of \((\varrho,\xi)\). Second, the \(95^{\mathrm{th}}\) quantile of the recentered TLR statistic is consistently located either at or below the \(95^{\mathrm{th}}\) quantile of the chi-squared distribution with one degree of freedom.
From this, we draw two conclusions. First, the relative dominance of the \(95^{\mathrm{th}}\) quantile of the chi-squared distribution over that of the recentered TLR statistic suggest that the combination of the former used as critical value in a test using the recentered TLR statistic will produce an approximately uniformly valid test, albeit conservative for some values of \((\varrho,\xi)\), and small \(d_Z\). Second, a two-step test with a small \(\alpha_1\approx 0\) and little if any change in critical values will produce an exactly uniformly valid test. Given the large-\(\xi\) distribution of the TLR statistic, such a test will numerically recover the Wald test in the strong-instrument limit. By computing conditional critical values, we can increase the power of this test in the weak-instrument regime.
The two-step test from (ref) requires the choice of a first-step level \(\alpha_1\). In order to achieve the strong-instrument Wald limit, it is preferable to choose this to be small. To this end, we let \(\alpha_1=10^{-5}\). However, the performance of the test is relatively invariant to this choice in simulations, and one can equivalently choose e.g. \(\alpha_1=10^{-10}\), giving slightly different second-step conditional critical values.
(ref) plots conditional critical values for a level \(\alpha=0.05\) two-step TLR test (dashed) and two-step recentered TLR test (solid),
for different values of \((\hat \varrho, \hat S)\) and \(d_Z\). We see that conditional critical values for the TLR test differ considerably from its strong-instrument limit. The conditional critical values of the recentered TLR test are either at or below the \(95^{\mathrm{th}}\) quantile of the chi-squared distribution with one degree of freedom for most values of \(\hat \varrho\) and \(\hat S\). In practice, such conditional critical values can either be tabulated, as done for the “VtF” test introduced by leeWhatWhenYou2023, or computed on the fly.
Lastly, it is of interest to understand the power properties of the two-step TLR and recentered tests relative to the Wald test constructed using the TSLS estimator. While this comparison is not prima facie reasonable, as the TSLS Wald test is not valid with weak instruments, it is of interest to understand the extent to which the two-step TLR and recentered TLR tests replicate the power properties of the TSLS Wald test with strong and moderately strong instruments.
In (ref) we plot the power of the TLR and recentered TLR two-step tests in simulations,
and compare them to the power of the TSLS Wald test, for different values of the concentration parameter, \(\mu^2\), and \(\omega\), fixing \(\nu^2=1\), \(h=0.1\) and \(d_Z=5\). We let \(\varrho_{\mathrm{true}}=\omega/\sqrt{h^2+\omega^2}\) denote the true value of \(\varrho\) off the null. For values of the concentration parameter, \(\mu^2\), we indicate the value of the first-stage F-statistic that would reject the hypothesis \(\mu^2\leq\bar \mu^2\) at level \(\alpha=0.05\) in the panel header. Recall from (ref) that \(\xi\) is a deterministic function of \(\omega,h^2,\nu^2\) and \(\mu^2\), and hence pinned down by these parameters. The x-axis is scaled relative to the value of the nuisance parameters, in order for curves to take comparable shapes across panels.
Considering (ref), three observations become apparent. First, the power curve of the recentered TLR statistic is close to being symmetric, and the test is close to being unbiased, for most of the parameter space. The exception is for extreme realizations of \(\varrho_{\mathrm{true}}\) and \(\xi\). Second, we see that unlike the TSLS Wald test, both the TLR and recentered TLR tests reject no more than 5% of the time at the null, \(\beta=\beta_0\). The TSLS Wald test, on the other hand, overrejects under the null at small \(\xi\) and large \(\varrho\). Third, we see that as the instrument set grows stronger, the power curves of the three tests converge to the same limit. This suggests that the power of the TLR and recentered TLR tests asymptote to the power of the Wald test.
In this paper we have shown that existing weak-instrument robust tests are invalid if used as tests for the TSLS estimand in a model with treatment effect heterogeneity and more than one instrument. In particular, we have shown that the AR, Kleibergen--Moreira LM and Moreira CLR tests can overreject at any level of instrument strength, and that the degree of overrejection is increasing as the instrument set grows stronger.
We have further shown that there exists a test statistic that can be leveraged to produce valid inference with treatment effect heterogeneity. In particular, we have constructed the TSLS likelihood-ratio (TLR) statistic for hypotheses on the TSLS estimand, and derived its limiting distribution. We have shown that when combined with a pretest in a two-step procedure, the test obtains uniform validity. Choosing a small first-step level, we have shown in simulations that the test inherits the properties of the Wald test when the instrument set is strong.