EconBase
← Back to paper

Inference on the TSLS Estimand with Weak Instruments and Treatment Effect Heterogeneity

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

58,674 characters

Inference on the TSLS Estimand with Weak Instruments and Treatment Effect Heterogeneity



\maketitle


\begin{abstract}
\noindent
Traditional inference on the coefficient in an instrumental variables regression does not retain size when the instrument set is weak. With constant treatment effects or one instrument, the \cite{andersonEstimationParametersSingle1949} AR test, the \cite{kleibergenPivotalStatisticsTesting2002}--\cite{moreiraConditionalLikelihoodRatio2003} LM test, and the Moreira CLR test provide robust alternatives which retain validity.
Under treatment effect heterogeneity, no valid inference procedure exists in the overidentified setting. This paper develops the TSLS likelihood ratio (TLR) statistic,
for performing inference on the TSLS estimand.
When combined with a two-step procedure in the spirit of \cite{bergerValuesMaximizedConfidence1994}, it retains uniform validity across both the weak- and strong-instrument regimes.
The procedure retains power with small choices of first-step level, hence
the test can be constructed to
numerically coincide with the Wald test in the strong-instrument limit.
\end{abstract}


\clearpage
\section{Introduction}
\label{sec:introduction}

Consider the problem  of performing inference on a parameter of interest,
\begin{align}
    \mathcal{H}_0:\beta = \beta_0 \qquad\qquad \text{vs.} \qquad\qquad \mathcal{H}_1:\beta \neq \beta_0,
\end{align}
when this parameter is the coefficient in an overidentified instrumental variables regression with one endogenous regressor and a potentially weak set of instruments,  known as the two-stage least squares (TSLS) estimand,
\begin{align}
    \quad  \beta \equiv \frac{\gamma'\Sigma_{ZZ}\delta}{\gamma'\Sigma_{ZZ}\gamma}.
    \label{eq:tsls-estimand}
\end{align}
The TSLS estimand is a weighted sum of the elements of the reduced-form coefficient vector, \(\delta\), weighted by the elements of the first-stage coefficient vector, \(\gamma\),
and the population instrument Gram matrix,  \(\Sigma_{ZZ}\).
It is a common parameter of interest in economics, political science, sociology, and other fields. Under certain assumptions, it obtains a causal interpretation as a positively weighted average of local average treatment effects, see e.g. \cite{angristTwoStageLeastSquares1995} and \cite{mogstadCausalInterpretationTwoStage2021}.

In the presence of weak instruments, the commonly applied Wald- or t-test is known to suffer from size distortion.
This problem  is well-studied, and
in the setting of a linear model with constant treatment effects (CTE),
the AR test (\citealp{andersonEstimationParametersSingle1949}), Kleibergen--Moreira Lagrange multiplier (LM) test (\citealp{kleibergenPivotalStatisticsTesting2002}, \citealp{moreiraConditionalLikelihoodRatio2003}) and conditional likelihood ratio (CLR) test (\citealp{moreiraConditionalLikelihoodRatio2003}) are known to deliver valid inference.\footnote{See  \cite{finlayImplementingWeakInstrumentRobust2009} and \cite{andrewsWeakInstrumentsInstrumental2019} for parsimonious expositions.}
In a model with
treatment effect heterogeneity, however, these are not valid tests for hypotheses on the TSLS estimand.
As pointed out by \citet[p. 5]{yapInferenceManyWeak2025}, there exists no  test proven to be valid  (i) in the presence of weak instruments and (ii) treatment effect heterogeneity which (iii) targets a parameter interpretable as a weighted average of local average treatment effects under commonly maintained assumptions when (iv) considering limit experiments that  fix the number of instruments.

To understand why standard robust tests are not valid for hypotheses on the TSLS estimand under treatment effect heterogeneity, observe that
the AR test performs inference on the hypothesis \(\beta=\beta_0\) by instead performing inference on a set of linear restrictions,
\begin{align}
    \mathcal{H}_0: \delta-\beta_0\gamma= 0 \qquad \text{vs.} \qquad \mathcal{H}_1: \delta-\beta_0\gamma \neq 0.
\end{align}
where the number of restrictions is equal to the number of instruments.
The set of restrictions can be rewritten as a composite hypothesis on the following form,
\begin{align}
    \mathcal{H}_{0,(a)}: \beta=\beta_0 \qquad \text{and} \qquad \mathcal{H}_{0,(b)}:\delta-\beta\gamma= 0 \quad \text{for some } \ \beta\in\R.
\end{align}
As is shown, e.g., by \cite{moreiraConditionalLikelihoodRatio2003}, the AR test statistic can be written as the sum of the LM test statistic and the \cite{sarganEstimationEconomicRelationships1958}--\cite{hansenLargeSampleProperties1982} $J$ test statistic. Maintaining CTE, the $J$ test is a test of instrument exogeneity. Maintaining exogeneity, but allowing for treatment effect heterogeneity, the $J$ test is essentially a test of the CTE assumption.
 The CLR and LM tests are derived from the AR restrictions, testing  \(\beta=\beta_0\) while maintaining the assumption of CTE. As a consequence, they are
 invalid in the presence of treatment effect heterogeneity when considered as tests for the TSLS hypothesis.

This paper proposes a test that is valid with weak instruments and treatment effect heterogeneity.
The test is derived by first observing that a hypothesis about the TSLS estimand, \(\beta=\beta_0\), can be reformulated into a quadratic constraint by multiplying by the denominator on both sides and rearranging,
\begin{align}
    \mathcal{H}_0:\gamma'\Sigma_{ZZ}(\delta -\beta_0\gamma) = 0
    \qquad \text{vs.} \qquad
    \mathcal{H}_1:\gamma'\Sigma_{ZZ}(\delta -\beta_0\gamma) \neq 0.
    \label{eq:quad-tsls-constraint}
\end{align}
While the restriction is similar in form to the AR restriction in that it does not involve division by a potentially small number,  it  is a single scalar restriction, and does not impose or test any ancillary assumptions, such as CTE.

The estimators of the reduced-form and first-stage coefficients are jointly asymptotically normal, with a variance-covariance matrix that is  consistently estimable.
This admits a likelihood ratio statistic, formulated as a quadratic optimization program with a quadratic constraint. We call this the TSLS likelihood ratio (TLR) statistic.
While the program does not have a closed-form solution, it is a special case of the generalized trust region subproblem, which has a known and unique solution, see e.g.
\cite{sternIndefiniteTrustRegion1995}.
   When instruments are strong, the TLR statistic is distributed asymptotically chi-squared with one degree of freedom. Combined with the fact that likelihood ratio tests are asymptotically equivalent to their associated Wald test under the null and against local alternatives (\citealp[p. 156]{ApproximationTheoremsMathematical1980}), this makes the implied test numerically indistinguishable from the Wald test in the strong-instrument regime.

When the instrument set is weak, the limiting distribution of the TLR statistic depends on two nuisance parameters.
The first nuisance parameter is a structural endogeneity parameter, capturing both selection on gains and levels. Similarly to the one-instrument case (see e.g. \citealp{vandesijpePowerConditionalLikelihood2023} and \citealp{leeWhatWhenYou2023}) it is consistently estimable under the null.
The second nuisance parameter captures the  overall identifying power of the instrument set. It is closely related to, but distinct from, the {concentration parameter} often used to gauge instrument strength in the literature. While it cannot be consistently estimated, it is associated with a sample statistic similar to, but distinct from, the first-stage F-statistic.
This allows for a two-step procedure  similar to the one discussed by \cite{bergerValuesMaximizedConfidence1994}, and suggested by \cite{staigerInstrumentalVariablesRegression1997}  for valid Wald inference in the linear constant effects model. Unlike their approach, we leverage more of the information contained in the sample to produce a test that is less conservative. In particular, we can set the first-step level so low as to numerically asymptote to the Wald test with strong instruments without substantial loss of power.


\paragraph*{Literature.}

The paper contributes to the literature on inference with weak instruments. Following the seminal contributions of \cite{nelsonDistributionInstrumentalVariables1990,nelsonFurtherResultsExact1990}, \cite{boundProblemsInstrumentalVariables1995} and \cite{staigerInstrumentalVariablesRegression1997}, a substantial number of papers have been devoted to understanding and resolving the finite-sample shortcomings of common procedures for inference and estimation with weak instruments. For a  comprehensive review, see \cite{andrewsWeakInstrumentsInstrumental2019}.
In the context of constant treatment effects,
\cite{andersonEstimationParametersSingle1949}
provided the first test with uniform validity, shown by \cite{moreiraTestsCorrectSize2009}
to be
uniformly most powerful unbiased with one instrument.
Later, \cite{kleibergenPivotalStatisticsTesting2002}
 and \cite{moreiraConditionalLikelihoodRatio2003} developed the Kleibergen--Moreira LM and Moreira CLR tests, with better power properties.
\cite{andrewsOptimalTwoSidedInvariant2006} showed that the latter of these tests is uniformly most powerful unbiased under homoskedasticity.
Recently, progress has been made on deriving
more intuitive tests (\citealp{leeValidTRatioInference2022}) and  tests with more intentional direction of power (\citealp{leeWhatWhenYou2023}) in the one-instrument case, both  building on the
two-step correction procedure suggested by \cite{staigerInstrumentalVariablesRegression1997},
which uses the first-stage F-statistic and a Bonferroni correction to produce valid Wald inference.

The paper also contributes to the study  of instrumental variables in policy evaluation and models with heterogeneous treatment effects. For a review of this literature, see \cite{mogstadInstrumentalVariablesUnobserved2024}. While the TSLS estimand is one of many potential ways of aggregating reduced-form estimates when overidentified, following \cite{imbensIdentificationEstimationLocal1994b}, \cite{angristTwoStageLeastSquares1995} and \cite{angristInterpretationInstrumentalVariables2000}, it has been a common parameter of interest for empirical practitioners. Under certain assumptions, it obtains a weighted average of local average treatment effects. As is shown by \cite{kolesarEstimationInstrumentalVariables2013}, this is not always the case for estimands associated with other estimators in the literature, such as the LIML estimator introduced by \cite{andersonEstimationParametersSingle1949}, and the Fuller--$k$ estimator, introduced by \cite{fullerPropertiesModificationLimited1977}.
The methods considered in this paper are robust to violations of the so-called monotonicity (uniformity) assumption, discussed e.g. by \cite{heckmanUnderstandingInstrumentalVariables2006}, \cite{mogstadCausalInterpretationTwoStage2021} and \cite{sigstadMonotonicityJudgesEvidence2026}, as they do not rely on the causal interpretation of the estimand.
While the one-instrument setting is more widespread in the literature,
a common setting with more than one instrument is
the examiner, judge or leniency design,
see \cite{chynExaminerJudgeDesigns2025} and \cite{goldsmith-pinkhamLeniencyDesignsOperators2025} for reviews. In terms of valid inference in the presence of heterogeneous treatment effects, \cite{evdokimovInferenceInstrumentalVariables2018} study the problem in a setting with many instruments, and \cite{yapInferenceManyWeak2025} with many weak instruments. To the author's knowledge, this is the first paper to consider the problem of valid inference on the TSLS estimand with a fixed number of weak instruments and potentially heterogeneous treatment effects.

The paper is also related to the literature using two-step procedures to obtain valid inference in the presence of nuisance parameters, see e.g. \cite{lohNewMethodTesting1985},
\cite{bergerValuesMaximizedConfidence1994}, \cite{silvapulleTestPresenceNuisance1996}, \cite{hansenTestSuperiorPredictive2005}, \cite{chernozhukovIntersectionBoundsEstimation2013}, \cite{romanoPracticalTwoStepMethod2014} and \cite{mccloskeyBonferronibasedSizecorrectionNonstandard2017}. The two-step test proposed by \cite{staigerInstrumentalVariablesRegression1997} is an example of such a testing procedure in the weak-instrument setting.



\paragraph*{Roadmap.}
The paper proceeds as follows.
In \Cref{sec:set-up-not} we set up the linear instrumental variables model with heterogeneous treatment effects and derive a  statistic that captures the total identifying power of the instrument set. In \Cref{sec:ar-clr} we show that existing weak-instrument robust tests fail as tests for the TSLS hypothesis with  treatment effect heterogeneity. In
\Cref{sec:results} we introduce the TLR statistic, and derive its limiting distribution. We show that it retains validity when combined with a two-step pretest and study its power properties in the weak- and strong-instrument regime. \Cref{sec:conclusion} concludes.



\section{Set-up and Notation}

\label{sec:set-up-not}

\paragraph*{Fundamentals.}
Let \(\{(Y_i,D_i,Z_i)\}_{i=1}^n\) be an i.i.d. sample, where \(Y_i\) is an outcome, \(D_i\) a treatment, both scalar, and \(Z_i\) a \(d_Z\)-dimensional vector of instruments, with \(d_Z\geq 2\).
Assume  all variables are demeaned and have finite fourth moments.
In the presence of covariates, let \(Y_i\) and \(D_i\) denote variables where their impact is appropriately projected off.
We are interested in the linear  heterogeneous treatment effects model,
\begin{align}
       Y_i &= \beta_i D_i + U_i\label{eq:model-equation}
\end{align}
where \(\beta_i\) is the causal effect of \(D_i\) on \(Y_i\), and \(U_i\) is the model residual.
Denote the linear regressions of  \(Y_i\) on \(Z_i\)  (reduced form) and \(D_i\) on \(Z_i\)  (first stage) as,
\begin{align}
Y_i &=  \delta_n' Z_i + W_i \label{eq:reduced-form-equation}\\
D_i &= \gamma_n' Z_i + V_i \label{eq:first-stage-equation}
\end{align}
where \(\delta_n\) is the vector of reduced-form coefficients, \(\gamma_n\) the vector of first-stage coefficients, and \(W_i\) and \(V_i\)
are regression residuals.
Subscripts on \(\delta_n\) and \(\gamma_n\) indicate that we do not restrict
the data generating process to be invariant with respect to the sample size \(n\).
To analyze weak-instrument properties, we follow \cite{staigerInstrumentalVariablesRegression1997} in letting \((\delta_n, \gamma_n)\) drift towards zero at rate \(1/\sqrt{n}\), such that
\((\delta_n,\gamma_n) \equiv  { (\delta,\gamma)}/{\sqrt{n}}\) for some vectors  \((\delta, \gamma)\).
To study strong-instruments properties, we hold \((\delta_n, \gamma_n)\) fixed as \(n\to\infty\). Let the TSLS estimand, \(\beta\), be defined as in \cref{eq:tsls-estimand}. It is uniquely pinned down by  \((\delta, \gamma)\), hence the parameter of interest does not vary with the sample size.
Let \(D_i(z)\) denote the potential treatment of an individual with instrument realization \(Z_i=z\).
We maintain the following assumptions.
\begin{assumptionp}{IVR}[Relevance]\label{ass:ivr}
\(\E[Z_iD_i] \neq \mathbf{0}_{d_Z}\)
\end{assumptionp}

\begin{assumptionp}{IVX}[Strong Exogeneity]\label{ass:ivx}
\((\beta_i,U_i,\{D_i(z)\}_{z\in\R^{d_Z}} ) \indep Z_i\)
\end{assumptionp}

\noindent
Assumptions \ref{ass:ivr} and \ref{ass:ivx}
are a subset of the assumptions usually maintained
in models with heterogeneous treatment effects, see e.g. \cite{angristTwoStageLeastSquares1995} and \cite{angristInterpretationInstrumentalVariables2000}.
No assumption is imposed on  the monotonicity (uniformity) of \(D_i(z)\) in \(z\). However, treatment effect heterogeneity requires the following refinement to the \cite{staigerInstrumentalVariablesRegression1997} weak-instrument limiting sequence.
\begin{assumptionp}{IVW}[Uniform Weakness]\label{ass:ivw}
For \(\sqrt{n}(\delta_n,\gamma_n)=(\delta,\gamma)\), as  \(n\to\infty\) and
\begin{enumerate}[label={\emph{(\alph*)}}, ref={(\alph*)}]
    \item  \label{ass:ivw-a}
    \(D_i\) binary, \(\lim_{n\to\infty}\sup_{z,z'}\P\bigl[D_i(z)\neq D_i(z')\bigr]=0\)
    \item  \label{ass:ivw-b}
    \(D_i\) continuous, \(\lim_{n\to\infty}\sup_{z,z'}\E\bigl[(D_i(z)- D_i(z'))^2\bigr]=0\)
\end{enumerate}
\end{assumptionp}

\noindent
Assumption \ref{ass:ivw} rules out the case in which an instrument is weak because the share of individuals who are  instrumented into more treatment (compliers) and less treatment (defiers) cancel. It is implied by the assumption of monotonicity (uniformity) of \(D_i(z)\) in \(z\), or equivalently by the \cite{vytlacilIndependenceMonotonicityLatent2002a} single-index model.









\paragraph*{Observed moments.}
Let \(\mathfrak{N}(\,\cdot\,,\,\cdot\,)\) denote the Normal distribution, \(\tau_n \equiv \left[\begin{smallmatrix}
    \delta_n \\
    \gamma_n
\end{smallmatrix}\right]\)
  the stacked vector of first-stage and reduced-form coefficients and \(\hat \tau\equiv \left[\begin{smallmatrix}
    \hat \delta \\
    \hat \gamma
\end{smallmatrix}\right]\)
 its OLS estimator. While the estimator is also a function of the sample size, for brevity of notation we let this be implicit in the following.
 The estimator has the following limiting distribution.
\begin{align}
    \sqrt{n}(\hat\tau-\tau_n) \convd \mathfrak{N}(\mathbf{0}_{2d_Z},\Sigma),
    \qquad\quad
    \Sigma   \equiv
   \begin{bmatrix}
       \Sigma_{\hat \delta} & \Sigma_{\hat \delta,\hat \gamma} \\
       \Sigma_{\hat \delta,\hat \gamma}' & \Sigma_{\hat \gamma}
   \end{bmatrix}
\end{align}
Let \(\Sigma_{ZZ}\equiv \E[Z_iZ_i']\) denote the population Gram matrix of \(Z_i\).
We can rewrite
the constraint defining the TSLS hypothesis from \cref{eq:quad-tsls-constraint} in terms of \(\tau_n\) as follows.
\begin{align}
    \tau_n'\,\Gamma(\beta_0) \,\tau_n = 0 \qquad \text{where} \qquad
    \Gamma(\beta_0) \equiv
    \begin{bmatrix}
        0 & \Sigma_{ZZ} \\
        \Sigma_{ZZ} & -2\beta_0\Sigma_{ZZ}
    \end{bmatrix}
\end{align}
\noindent
With constant treatment effects, the first-stage F-statistic summarizes the strength of the instrument set. With treatment effect heterogeneity, we need to consider a statistic that also captures other facets of the identifying power of the instruments. We record this statistic and its limiting distribution as a lemma.

\begin{lemmap}{ST}[Summary Statistic]\label{lem:st}
Let \(\hat \Sigma\) denote a consistent estimator for \(\Sigma\), and let \(\hat S \equiv \lVert (\hat\Sigma/n)^{-\frac{1}{2}}\hat \tau\rVert^2/d_Z\) denote the Wald test statistic for the  hypothesis \(\tau_n=0\). Let \(\mathfrak{C}^2_{2d_Z}(\xi)\) denote the non-central chi-squared distribution with \(2d_Z\) degrees of freedom and non-centrality parameter \(\xi\). Then, {for \(\sqrt{n}\,\tau_n = \tau\), as \(n\to\infty\),}
    \[d_Z\hat S  \convd \mathfrak{C}^2_{2d_Z}\left(\xi\right),\]
where \(\xi \equiv \lVert (\Sigma/n)^{-\frac{1}{2}}\tau_n\rVert^2\).
\end{lemmap}



\begin{proof}
    See Appendix~\ref{sec:proof-st}.
\end{proof}
\noindent
The parameter \(\xi\) is related to the concentration parameter often used to enumerate instrument strength in the literature. We record the relation as a lemma.
\begin{lemmap}{PD}[Parameter Decomposition]\label{lem:pd}
Let \(\mu^2\equiv \lVert  (\Sigma_{\hat \gamma}/n)^{-\frac{1}{2}}\gamma_n\rVert^2\) denote the concentration parameter,
\(\omega\equiv \beta_{\mathrm{ols}}-\beta\), \(\beta_{\mathrm{ols}}\equiv \cov[Y_i,D_i]/\var[D_i]\), the difference between the OLS and TSLS estimands, and
\(\nu^2 \equiv \lVert \Sigma_{ZZ}^{{1}/{2}}(\delta_n - \beta \gamma_n)\rVert^2 / \lVert \Sigma_{ZZ}^{{1}/{2}} \gamma_n\rVert^2\)
the TSLS-weighted dispersion of
local average treatment effects.
Under Assumptions \ref{ass:ivx} and \ref{ass:ivw}, as \({\sqrt{n}\,\tau_n=\tau}\) and \(n\to\infty\), we have,
\begin{align*}
    \xi = \mu^2\left(1 + h^{-2}(\nu^2 + \omega^2)\right),
\end{align*}
where
\(h^2\equiv{(\var[W_i]-\beta_{\mathrm{ols}}^{2}\var[V_i])/\var[V_i]}\) is a  ratio of residual variances.
All parameters are constant along the weak-instrument asymptotic sequence.
\end{lemmap}

\begin{proof}
    See Appendix~\ref{sec:proof-pd}.
\end{proof}
\noindent
Observe that with one instrument or constant treatment effects, we have \(\nu^2=0\).


\paragraph*{Eigenvalue Representation.}

It will be useful to express the data in terms of
a particular eigendecomposition.
Let \(\Lambda(\beta_0) \equiv
\Sigma^{\frac{1}{2}}
\Gamma(\beta_0)
\Sigma^{\frac{1}{2}}\) denote
 the rotation of  \(\Gamma(\beta_0)\) in terms of \(\Sigma\).
Let \(X\) denote the matrix of eigenvectors of \(\Lambda(\beta_0)\), and \(\hat X\) its analog estimator.
 Define the rotated first-stage and reduced-form estimators and estimands as
\begin{align}
    \hat Q \equiv \hat X'\hat\Sigma^{-1/2}\sqrt{n}\,\hat\tau
    \qquad\text{and}\qquad
    q \equiv X'\Sigma^{-1/2}\sqrt{n}\,\tau_n.
\end{align}
While
\(\Lambda(\beta_0)\) is indefinite, the block-structure of \(\Gamma(\beta_0)\) and \Cref{ass:ivx} restrict the sign of, and variation in, its eigenvalues.
\begin{lemmap}{EH}[Eigenvalue Homogeneity]\label{lem:eh}
\(\Lambda(\beta_0)\) has \(d_Z\) positive and \(d_Z\) negative eigenvalues.
Let \(\{\kappa^{(k)}_-,\kappa^{(k)}_+\}_{k=1}^{d_Z}\) denote these eigenvalues, and let \((q_+,q_-)\) denote the subvectors of \(q\) associated with the positive and negative eigenspace.
Under Assumptions \ref{ass:ivx} and \ref{ass:ivw}, along any sequence with \(\sqrt{n}\,\tau_n=\tau\), as \(n\to\infty\), for some common \((\kappa_-,\kappa_+)\), we have
    \begin{enumerate}
    \item
        \(\kappa^{(k)}_-\to\kappa_-\) and  \(\kappa^{(k)}_+\to\kappa_+\)  for all \(k=1,\ldots,d_Z\).
    \item
    \(\tau_n'\,\Gamma(\beta_0) \,\tau_n = 0 \iff (1+\varrho)\lVert q_+ \rVert^2 = (1-\varrho)\lVert q_- \rVert^2\)
    \end{enumerate}
    where \(\varrho \equiv  (\kappa_+-\lvert \kappa_-\rvert) / (\kappa_++\lvert\kappa_-\rvert)\) denotes the signed eigenvalue share. If we let \(\omega_0 \equiv \beta_{\mathrm{ols}} - \beta_0 \) denote the value \(\omega\) takes under the null, we have \(\varrho = \omega_0 /\sqrt{h^2+\omega_0^2}\).
\end{lemmap}

\begin{proof}
    See Appendix~\ref{sec:proof-eh}.
\end{proof}

\noindent
Note that \((q, X, \kappa_\pm, \varrho)\) and their estimators depend on the hypothesized \(\beta_0\). For brevity of notation, we suppress this in the following.






\section{On the Validity of Weak-Instrument Robust Tests}

\label{sec:ar-clr}

The AR, LM and CLR tests are commonly understood to provide inference that is robust to weak instruments. In the following, we show that this  is not the case if the tests are interpreted as tests of the TSLS estimand under treatment effect heterogeneity.
We start by recording the constant treatment effect limiting distribution of the three statistics, as derived by \cite{andersonEstimationParametersSingle1949}, \cite{kleibergenPivotalStatisticsTesting2002} and \cite{moreiraConditionalLikelihoodRatio2003}.



\begin{lemmap}{ROB-CTE}[Robustness, Constant Treatment Effects]\label{lem:rob-cte}
Assume there exists a scalar \(\beta\) such that \(\delta_n=\beta\gamma_n\).
Let \(\hat\Omega(\beta_0)\) denote a consistent estimator of the variance of the restriction \(\hat\delta - \beta_0\hat\gamma\), let \(\tilde\gamma(\beta_0)\)
denote the first-stage estimator orthogonalized relative to \(\hat\delta - \beta_0\hat\gamma\) and let \(\hat\Psi(\beta_0)\) denote an estimator of its variance.
Define
\begin{align*}
\mathrm{AR}(\beta_0)  &\equiv d_Z^{-1}\lVert (\hat\Omega(\beta_0)/n)^{-1/2}(\hat\delta - \beta_0\hat\gamma)\rVert^2,   \\[2pt]
\mathrm{LM}(\beta_0)  &\equiv d_Z^{-1}\bigl[\tilde\gamma(\beta_0)'(\hat\Omega(\beta_0)/n)^{-1}(\hat\delta - \beta_0\hat\gamma)\bigr]^2\,\big/\,\lVert(\hat\Omega(\beta_0)/n)^{-1/2}\tilde\gamma(\beta_0)\rVert^2,\\[2pt]
\mathrm{CLR}(\beta_0) &\equiv \tfrac{1}{2}\!\left[d_Z\mathrm{AR}(\beta_0)- \lVert\hat r(\beta_0)\rVert^2 + \sqrt{(d_Z\mathrm{AR}(\beta_0)-\lVert\hat r(\beta_0)\rVert^2)^2 + 4 \,d_Z\mathrm{LM}(\beta_0)\, \lVert\hat r(\beta_0)\rVert^2} \right]
\end{align*}
where \(\hat r(\beta_0) \equiv (\hat\Psi(\beta_0)/n)^{-1/2}\,\tilde\gamma(\beta_0)\).
Let \(\mathfrak{C}^2_{d_Z}\) denote the chi-squared distribution with \(d_Z\) degrees of freedom. We then have, \(d_Z\mathrm{AR}(\beta_0)\convd \mathfrak{C}^2_{d_Z}\), \(d_Z\mathrm{LM}(\beta_0)\convd \mathfrak{C}^2_1\) and \(\mathrm{CLR}(\beta_0)\convd m(\mathfrak{C}^2_{1},\mathfrak{C}^2_{d_Z-1},\lVert \hat r(\beta_0)\rVert^2)\), where
\begin{align*}
m(x_1,x_{d_Z-1},r) \;\equiv\;
\tfrac{1}{2}\!\left[x_{d_Z-1}+x_1- r + \sqrt{(x_{d_Z-1}+x_1-r)^2 + 4\, x_1\, r} \right].
\end{align*}
\end{lemmap}

\begin{proof}
    See Appendix~\ref{sec:proof-rob-cte}.
\end{proof}
\noindent
With treatment effect heterogeneity, these limits no longer obtain. We record the result as a lemma.
\begin{lemmap}{ROB-HTE}[Robustness, Heterogeneous Treatment Effects]\label{lem:rob-hte}
Let \(\beta=\beta_0\) denote the TSLS hypothesis.
Let \(\theta_{\mathrm{AR}} \equiv \mu^2\nu^2/(h^2+\omega^2)\), and let
 \(\zeta\sim \mathfrak{N}(0,1)\),
\(\zeta^2_{1} \sim \mathfrak{C}^2_{1}(\omega^2\theta_{\mathrm{AR}}/h^2)\), and \(\zeta^2_{d_Z-1} \sim \mathfrak{C}^2_{d_Z-1}(\mu^2(h^2+\omega^2)/h^2)\) denote three random variables, where \(\zeta \indep \zeta^2_1 \indep \zeta^2_{d_Z-1}\).
Then, as \(\sqrt{n}\tau_n=\tau\), \(n\to\infty\),
\begin{enumerate}
    \item
    \(d_Z\mathrm{AR}(\beta_0)\convd \mathfrak{C}^2_{d_Z}(\theta_{\mathrm{AR}})\)
    \item
    \(d_Z\mathrm{LM}(\beta_0)\convd (\Upsilon+\zeta)^2\), with
    \( \Upsilon^2 \;\sim\; \theta_{\mathrm{AR}}\cdot{\zeta^2_1}/({\zeta^2_1 + \zeta^2_{d_Z-1}})\)



    \item
    \(\mathrm{CLR}(\beta_0) \convd m\bigl((\Upsilon+\zeta)^2,\;\mathfrak{C}^2_{d_Z}(\theta_{\mathrm{AR}}) - (\Upsilon+\zeta)^2,\;\lVert \hat r(\beta_0)\rVert^2 \bigr)\)
\end{enumerate}
Let \(c^T_{1-\alpha}\) denote the \((1-\alpha)^{th}\) quantile of the constant-treatment-effects limiting distribution of test statistic \(T\) and assume \(\omega^2 >0\). Then, for \(\nu^2>0\),
\[
\limsup_{n\to\infty}\,\P\left[\mathrm{T}(\beta_0)> c^{\mathrm{T}}_{1-\alpha}\right] > \alpha
\]
where \(T\in\{\mathrm{AR},\mathrm{LM},\mathrm{CLR}\}\). Further, \(\lim_{\mu^2\to\infty}\,\limsup_{n\to\infty}\,\P[\mathrm{T}(\beta_0)> c^{\mathrm{T}}_{1-\alpha}] =1\).
\end{lemmap}

\begin{proof}
    See Appendix~\ref{sec:proof-rob-hte}.
\end{proof}
\noindent
In \Cref{fig:other-tests}, we plot the size of the AR, LM and CLR tests,
\begin{figure}[!htb]
    \centering
    \subbottom[\(\nu^2=0\), High endogeneity]{\includegraphics[width=0.49\textwidth]{figs/size_existing_rho09_hom.pdf}}
    \subbottom[\(\nu^2>0\), High endogeneity]{\includegraphics[width=0.49\textwidth]{figs/size_existing_rho09_het.pdf}}
    \subbottom[\(\nu^2=0\), Intermediate endogeneity]{\includegraphics[width=0.49\textwidth]{figs/size_existing_rho05_hom.pdf}}
    \subbottom[\(\nu^2>0\), Intermediate endogeneity]{\includegraphics[width=0.49\textwidth]{figs/size_existing_rho05_het.pdf}}
    \subbottom[\(\nu^2=0\), No endogeneity]{\includegraphics[width=0.49\textwidth]{figs/size_existing_rho00_hom.pdf}}
    \subbottom[\(\nu^2>0\), No endogeneity]{\includegraphics[width=0.49\textwidth]{figs/size_existing_rho00_het.pdf}}

    \caption{\textbf{Size distortion of the AR, LM and CLR Tests Under Heterogeneous Treatment Effects.} This figure plots the size of the AR, LM and CLR tests if taken as tests of the hypothesis \(\beta=\beta_0\), under constant (\(\nu^2=0\)) and heterogeneous (\(\nu^2>0\)) treatment effects, with \(d_Z=5\) instruments.
    The blue line shows the AR test, the red LM, and the green the CLR test.
    Size is computed by simulating asymptotic realizations of \(\hat \tau\) under high, intermediate and no endogeneity, corresponding to \(\omega/\sqrt{h^2+\omega^2} \in \{0.9, 0.5, 0\}\).
    The y-axis shows rejection rate under the null (size), and the x-axis enumerates instrument strength in terms of the concentration parameter and the \(95^{\mathrm{th}}\) quantile of the associated non-central chi-squared distribution.}
    \label{fig:other-tests}
\end{figure}
if used as tests for the TSLS hypothesis, \(\beta=\beta_0\), in a Monte Carlo simulation, with and without treatment effect heterogeneity. We simulate by drawing realizations of \(\hat \tau\) from its limiting distribution, with known variance,  consistent with constant treatment effects (panels (a), (c) and (e)), and heterogeneous treatment effects (panels (b), (d) and (f)), and compute the rejection rate  under the null. The x-axis is enumerated in terms of the \(d_Z\)-scaled concentration parameter, \(\mu^2/d_Z\). The associated first-stage F-statistic cutoffs that reject this concentration parameter at level \(\alpha=0.05\) are indicated in parentheses.

In panels (a), (c) and (e) of \Cref{fig:other-tests}, we see that all three tests retain size under constant treatment effects, irrespective of the level of endogeneity. In panels (b), (d) and (f), we see that under heterogeneous treatment effects, the tests reject more often than their nominal level indicates, and are thus invalid as tests for the TSLS hypothesis, with size distortion increasing in instrument strength. With sufficiently strong instruments and non-zero endogeneity, the tests always reject.




\section{Valid Inference on the TSLS Estimand with Weak Instruments}
\label{sec:results}

We now turn to deriving a test for the TSLS hypothesis that is valid with weak instruments and treatment effect heterogeneity. We start by deriving the likelihood ratio statistic for the TSLS hypothesis. We then proceed to derive its strong- and weak-instrument limiting distributions. We conclude by showing  that we can perform uniformly valid inference by combining the test with a pretest for the nuisance parameter in a two-step procedure without meaningful loss of power.


\paragraph*{The TLR Statistic.}

Consider the asymptotic log-likelihood of \(\tau_n\) given \(\hat \tau\), given as,
\begin{align}
    \ell(\tau_n\mid \hat \tau)
    =
    L - \tfrac{1}{2}\,n(\hat \tau -\tau_n)'\hat\Sigma^{-1}(\hat \tau - \tau_n)
\end{align}
where \(L\) is a constant.
Unconstrained, the log-likelihood is
maximized by \(\hat \tau\). Under the TSLS constraint, the maximized log-likelihood is a quadratic program with a quadratic constraint. Considering their difference, we get the likelihood ratio statistic for the TSLS hypothesis.
The implied quadratic program has a known solution,  due to \cite{sternIndefiniteTrustRegion1995}.
We record the result as a proposition.

\begin{propositionp}{TLR}[TSLS Likelihood Ratio Statistic]\label{prop:tlr}
The asymptotic likelihood ratio statistic for the hypothesis \(\beta=\beta_0\), when scaled by \(d^{-1}_Z\), is given as
\begin{align*}
\mathrm{TLR}(\beta_0) \equiv \min_{\tau_n \in \R^{2d_Z}}\,d_Z^{-1}n(\hat \tau -\tau_n)'\hat\Sigma^{-1}(\hat \tau - \tau_n) \qquad \text{s.t.} \qquad \tau_n'\,\hat\Gamma(\beta_0) \,\tau_n = 0
\end{align*}
where \(\hat\Gamma(\beta_0)\) is a consistent estimator for \(\Gamma(\beta_0)\).
Let \(\hat \Lambda(\beta_0)\) denote a consistent estimator for \(\Lambda(\beta_0)\).
The program has a unique solution,
    \begin{align*}
    \mathrm{TLR}(\beta_0) = d_Z^{-1}\sum_{j=1}^{2d_Z} \hat Q_j^2\!\left(\frac{\lambda^*\hat\kappa_j}{1+\lambda^*\hat\kappa_j}\right)^{\!2}
    \quad\text{s.t.}\quad
    \sum_{j=1}^{2d_Z}\frac{\hat\kappa_j \hat Q_j^2}{(1+\lambda^*\hat\kappa_j)^2} = 0
\end{align*}
where \(\{\hat \kappa_j\}_{j=1}^{2d_Z}\) denote the eigenvalues of \(\hat \Lambda(\beta_0)\), \(\hat Q_j\) the \(j^{\mathrm{th}}\) element of \(\hat Q\) and \(\lambda^*\in(-(\max_{j:\hat\kappa_j>0}\hat\kappa_j)^{-1},-(\min_{j:\hat\kappa_j<0}\hat\kappa_j)^{-1})\) is implicitly defined and unique on this interval.
\end{propositionp}

\begin{proof}
    See Appendix~\ref{sec:proof-tlr}.
\end{proof}

\noindent
We scale by \(d_Z^{-1}\) to simplify comparison between settings with different numbers of instruments.
Observe that the above result only relies on the joint asymptotic normality of the reduced form and first stage, and does not require any further assumptions.

Under maintained assumptions,
the TLR statistic has a limiting distribution that is well-behaved with strong instruments and depends  on two nuisance parameters when the instruments are weak. We record the result.

\begin{propositionp}{LD}[TLR Limiting Distribution]\label{prop:ld}
Let \(S\) be defined such that \(d_Z\hat S \convd S\).
    Under Assumptions \ref{ass:ivr}, \ref{ass:ivx} and \ref{ass:ivw}, the TLR statistic obtains the following limiting distribution.
    \begin{enumerate}
        \item
        When \(\tau_n\) is fixed and \(n\to\infty\),
        \(d_Z\,\mathrm{TLR}(\beta_0)\convd\mathfrak{C}^2_1\).
        \item
        When \(\sqrt n\,\tau_n=\tau\) is fixed and \(n\to\infty\),
        \[
            d_Z\,\mathrm{TLR}(\beta_0)\;\convd\;\tfrac{1}{2}\bigl(\sqrt{(1+\varrho)\,S_+}-\sqrt{(1-\varrho)\,S_-}\bigr)^{\!2},
        \]
        where \(S_+\sim\mathfrak{C}^2_{d_Z}((1-\varrho)\,\xi/2)\), \(S_-\sim\mathfrak{C}^2_{d_Z}((1+\varrho)\,\xi/2)\), \(S_+ \indep S_-\), and \(S_+ + S_-=S\).
    \end{enumerate}
\end{propositionp}

\begin{proof}
    See Appendix~\ref{sec:proof-ld}.
\end{proof}
\noindent
The weak-instrument limiting distribution  obtains the strong-instrument limit as \(\xi\to \infty\). We record this, and two other useful boundary results, as a proposition.
\begin{propositionp}{BD}[Boundary Case Distributions]\label{prop:bd}
Consider the weak-instrument regime, such that \(\tau_n\sqrt{n}=\tau\) constant as \(n\to\infty\). Then,
\begin{enumerate}
    \item
    As \(\xi\to \infty\), we have
    \(d_Z\mathrm{TLR}(\beta_0) \convd
    \mathfrak{C}^2_1\).
    \item
    As \(\varrho\to \pm 1\), we have
    \(d_Z\mathrm{TLR}(\beta_0) \convd
     \mathfrak{C}^2_{d_Z}\).
    \item
    At  \(\xi= 0\), \(\varrho=0\), we have
    \(d_Z\mathrm{TLR}(\beta_0) \convd
    \left(\sqrt{\mathfrak{C}^2_{d_Z}} -\sqrt{\mathfrak{C}^2_{d_Z}}\right)^2 \,\big/\,2\).
\end{enumerate}
\end{propositionp}

\begin{proof}
    See Appendix~\ref{sec:proof-bd}.
\end{proof}
\noindent
The TLR statistic is biased, in the sense that its value is minimized away from the truth, i.e. for \(\beta\neq\beta_0\). In the following, we show that we can construct a near-unbiased test by recentering the TLR statistic.


\paragraph*{A recentered TLR statistic.}
In order to recenter the TLR statistic, it is necessary to first
consider the signed root of the TLR statistic. The signed-root TLR statistic can be used to test the one-sided hypothesis, \(\beta \leq \beta_0\). We record its definition and limiting distribution as a lemma.


\begin{lemmap}{SLR}[Signed-root TLR Statistic]\label{lem:slr}
Denote the signed root of the TLR statistic,
\begin{align*}
    {\mathrm{SLR}}(\beta_0) \;\equiv\; \operatorname{sign}\!\left(\sum_{j=1}^{2d_Z}\hat\kappa_j\,\hat Q_j^2\right)\sqrt{\mathrm{TLR}(\beta_0)}.
\end{align*}
For \(\sqrt{n}\tau_n=\tau\), as \(n\to\infty\), the limiting distribution is given as,
\begin{align*}
    {\mathrm{SLR}}(\beta_0) \convd \tfrac{1}{\sqrt{2\,d_Z}}\left(\sqrt{(1+\varrho)\,S_+}\,-\,\sqrt{(1-\varrho)\,S_-}\right) .
\end{align*}
Denote this limiting distribution \( T(\varrho,\xi)\).
The  distribution  has non-zero mean for \(\varrho\neq 0\).



\end{lemmap}


\begin{proof}
    See Appendix~\ref{sec:proof-slr}.
\end{proof}
\noindent
The bias of the TLR statistic is associated with the  non-zero mean of its signed root.
We can reduce this bias by subtracting an estimate of this mean. In order to take into account finite-sample deviations from exact eigenvalue homogeneity, we use a parametric bootstrap estimator.
We record the definition of the recentered TLR statistic.









\begin{definitionp}{RTLR}[Recentered TSLS Likelihood-Ratio Statistic]\label{prop:rtlr}
Let \(\hat q^*\) denote the constrained minimizer of the TLR statistic, \(\tilde Q^b \sim\mathfrak{N}(\hat q^*,\,\boldsymbol{\iota}_{2d_Z})\) a parametric bootstrap draw from its limiting distribution, where \(\boldsymbol{\iota}_{2d_Z}\) denotes the identity matrix, and let the bootstrapped TLR statistic be given as,
\begin{align*}
    \mathrm{TLR}^b(\beta_0) = d_Z^{-1}\sum_{j=1}^{2d_Z} (\tilde Q_j^b)^2\!\left(\frac{\tilde\lambda^*\hat\kappa_j}{1+\tilde\lambda^*\hat\kappa_j}\right)^{\!2}
    \quad\text{s.t.}\quad
    \sum_{j=1}^{2d_Z}\frac{\hat\kappa_j (\tilde Q_j^b)^2}{(1+\tilde\lambda^*\hat\kappa_j)^2} = 0.
\end{align*}
Let \(\hat b(\hat q^*, \hat \kappa)\) denote parametric bootstrap estimators of the asymptotic mean of the signed-root TLR statistic given the vector of eigenvalues, \(\hat\kappa\).
Define the recentered TSLS likelihood-ratio statistic as
\begin{align*}
    \mathrm{RTLR}(\beta_0) \equiv {({\mathrm{SLR}}(\beta_0) - \hat b(\hat q^*, \hat \kappa))^2}.
\end{align*}
\end{definitionp}

\noindent
Recentering does not obtain a pivotal statistic. However,
simulations suggest that
 it is approximately centered across large parts of the parameter space.
As we show in the following, it is possible to leverage both the raw and recentered TLR statistic to
construct  uniformly valid tests for the TSLS hypothesis.

\paragraph*{Robust TLR Inference.}

A test that is valid in both the strong- and weak-instrument regimes must be valid for all realizations of the nuisance parameters, \((\varrho,\xi)\).
In the following, we introduce a two-step procedure that uses the fact that we can consistently estimate \(\varrho\) under the null, and that \(\hat S\) carries information about \(\xi\), to produce uniformly valid inference, in the spirit of \cite{bergerValuesMaximizedConfidence1994}.






The statistic \(\hat S\) is such that \(d_Z\hat S\) is distributed asymptotically non-central chi-squared with non-centrality parameter  \(\xi\). It follows that we can
use \(\hat S\) to produce a level \((1-\alpha_1)\) confidence interval for \(\xi\). We can then use the worst-case \((1-\alpha_2)^{\mathrm{th}}\) quantile in this interval as critical value for the test. If \(\alpha_1+\alpha_2 = \alpha\), this two-step procedure is a valid test for \(\beta=\beta_0\) at level \(\alpha\).
We record the result.









\begin{propositionp}{TS}[Two-step Inference]\label{prop:ts}
Let \(\hat \Xi_{1-\alpha_1}\) denote a \(1-\alpha_1\) confidence interval for \(\xi\) obtained by inverting \(\hat S\), and let \(\alpha_2 \equiv \alpha -\alpha_1 \). Let \(\hat T(\beta_0)\) denote the raw, signed-root or recentered TLR statistic, and let
\(c_{1-\alpha_2}(\varrho,\xi) \) denote the \((1-\alpha_2)^{th}\) {quantile} of its limiting distribution.
  We then have:
\begin{align*}
    \limsup_{n\to\infty}\,\sup_{\varrho,\xi}\,\P\!\left(\hat T(\beta_0) > \max_{\xi \in \hat\Xi_{1-\alpha_1}} c_{1-\alpha_2}(\hat\varrho,\xi)\right) \;\leq\; \alpha_1 + \alpha_2 \;=\; \alpha.
\end{align*}
It follows that \(\hat T(\beta_0) > \max_{\xi \in \hat\Xi_{1-\alpha_1}} c_{1-\alpha_2}(\hat\varrho,\xi)\) is a {uniformly} valid test for \(\beta=\beta_0\).
\end{propositionp}
\begin{proof}
    See Appendix~\ref{sec:proof-ts}.
\end{proof}
\noindent
When the first-stage F-statistic is sufficiently small, tests based on the TLR statistic fail to reject any \(\beta_0\) far from the origin. It follows that associated confidence intervals have infinite length.
We record the result as a lemma.
\begin{lemmap}{UC}[Unbounded Confidence Interval]\label{lem:uc}
Let \(\hat F \equiv \lVert (\hat\Sigma_{\hat\gamma}/n)^{-1/2}\hat\gamma\rVert^2/d_Z\) denote the Wald test statistic for the hypothesis that \(\gamma_n=0\).
The two-step TLR test fails to reject \(\beta_0\) for all sufficiently large \(\lvert\beta_0\rvert\) if and only if
\[
    d_Z\hat F \;\leq\; {\chi^2_{d_Z,1-\alpha_2}}.
\]
Equivalently, confidence intervals formed by test inversion have infinite length.
\end{lemmap}
\begin{proof}
    See Appendix~\ref{sec:proof-uc}.
\end{proof}
\noindent
It is instructive to understand how much  the limiting distribution of the TLR and recentered TLR statistics differ for different values of \((\varrho,\xi)\) and \(d_Z\). We study this in Monte Carlo simulations. As the limiting distribution of the TLR statistic is symmetric in \(\varrho\), it suffices to consider positive \(\varrho\).

In \Cref{fig:cutoffs-oracle}, we plot the \(95^{\mathrm{th}}\) quantile of the distribution of the TLR statistic and the recentered TLR statistic for different values of \((\varrho, \xi)\) and \(d_Z\).
\begin{figure}[!t]
    \centering

    \includegraphics[width=\textwidth]{figs/cutoffs_oracle.pdf}

    \caption{\textbf{Quantiles of the TLR and RTLR Statistics.} This figure plots the \(95^{\mathrm{th}}\) quantile of the TLR and RTLR statistics for different \((\lvert\varrho\rvert,\xi)\) and \(d_Z\). The quantiles are plotted relative to the \(95^{\mathrm{th}}\) quantile of the chi-squared distribution with one degree of freedom and divided by \(d_Z\). The blue line plots the quantiles of the TLR statistic and the red line plots the quantiles of the recentered TLR (RTLR) statistic.  The dashed line indicates \(\chi^2_{d_Z,1-\alpha_2}\), i.e. the limit as \(\lvert\varrho\rvert\to 1\). The x axis  is logarithmic.
    }


    \label{fig:cutoffs-oracle}
\end{figure}
We make two observations. First, the \(95^{\mathrm{th}}\) quantile of the TLR statistic meaningfully overshoots the \(95^{\mathrm{th}}\) quantile of the chi-squared distribution with one degree of freedom for a wide range of values of \((\varrho,\xi)\). Second, the \(95^{\mathrm{th}}\) quantile of the recentered TLR statistic is consistently located either at or below the  \(95^{\mathrm{th}}\) quantile of the chi-squared distribution with one degree of freedom.

From this, we draw two conclusions. First, the relative dominance of the  \(95^{\mathrm{th}}\) quantile of the chi-squared distribution  over that of the recentered TLR statistic suggest that
the combination of the former used as critical value in a test using the recentered TLR statistic will produce an approximately  uniformly valid test, albeit conservative for some values of \((\varrho,\xi)\), and small \(d_Z\). Second, a two-step test with a small \(\alpha_1\approx 0\)
and little if any change in critical values will produce an exactly uniformly valid test.
Given the large-\(\xi\) distribution of the TLR statistic, such a test will numerically recover the Wald test in the strong-instrument limit.
By computing conditional critical values, we can increase the power of this test in the weak-instrument regime.

The two-step test from \Cref{prop:ts} requires the choice of a first-step level \(\alpha_1\). In order to achieve the strong-instrument Wald limit, it is preferable to choose this to be small. To this end, we let \(\alpha_1=10^{-5}\). However, the performance of the test is relatively invariant to this choice in simulations, and one can equivalently choose e.g. \(\alpha_1=10^{-10}\), giving slightly different second-step conditional critical values.

\Cref{fig:cutoffs-conditional} plots conditional critical values for a level \(\alpha=0.05\) two-step TLR test (dashed) and two-step recentered TLR  test (solid),
\begin{figure}[!b]
    \centering

    \includegraphics[width=\textwidth]{figs/cutoffs_conditional.pdf}

    \caption{\textbf{Conditional Second-Step Critical Values of the TLR and Recentered TLR Statistics.} This figure plots conditional second-step critical values for an \(\alpha=0.05\) test, with first-step level \(\alpha_1=10^{-5}\), for different values of \((\hat \varrho,\hat S)\) and \(d_Z\).
    The critical values are plotted relative to the \(95^{\mathrm{th}}\) quantile of the chi-squared distribution with one degree of freedom and divided by \(d_Z\). The solid line plots critical values for the recentered TLR statistic and the dashed line plots the quantiles of the TLR statistic.  The thin dashed horizontal line indicates \(\chi^2_{d_Z,1-\alpha_2}\), i.e. the limit as \(\lvert\varrho\rvert\to 1\). The x axis starts at the smallest \(\hat S\) consistent with  \(\hat F > \chi^2_{d_Z,0.95}/d_Z\).
    }


    \label{fig:cutoffs-conditional}
\end{figure}
for different values of \((\hat \varrho, \hat S)\) and \(d_Z\). We see that conditional critical values for the
 TLR test differ considerably from its strong-instrument limit.
The conditional critical values of the recentered TLR test are either at or below the \(95^{\mathrm{th}}\) quantile of the chi-squared distribution with one degree of freedom for most values of  \(\hat \varrho\) and \(\hat S\). In practice, such conditional critical values can either be tabulated, as done for the ``VtF'' test introduced by \cite{leeWhatWhenYou2023}, or computed on the fly.





Lastly, it is of interest to understand the power properties of the two-step TLR and recentered tests relative to the Wald test constructed using the TSLS estimator. While this comparison is not prima facie reasonable, as the TSLS Wald test is not valid with weak instruments, it is of interest to understand the extent to which  the two-step TLR and recentered TLR  tests replicate the power properties of the TSLS Wald test with strong and moderately strong instruments.

In \Cref{fig:lr-regular-asy-all-power} we plot the power of the  TLR and recentered TLR two-step tests in simulations,
\begin{figure}[!b]
    \centering
    \includegraphics[width=\textwidth]{figs/power_tlr.pdf}


    \caption{\textbf{Power of the TLR and Recentered TLR Tests.} This figure shows simulations of the power of the two-step TLR and two-step recentered TLR tests against alternatives with \(d_Z=5\) instruments and different degrees of endogeneity. The red line shows the two-step recentered TLR test, the blue line shows the two-step TLR test and the green line shows the TSLS Wald test.
    The y axis shows the rejection rate, and the x axis enumerates distance from the true parameter, in terms of a natural scaling of nuisance parameters. The simulation is conducted with \(\varrho_{\mathrm{true}}\equiv \omega/\sqrt{h^2+\omega^2} \in\{0.1,0.5,0.8,0.99,0.999\}\), displayed across vertical panels.
    }
    \label{fig:lr-regular-asy-all-power}
\end{figure}
and compare them to the power of the TSLS Wald test, for different values of the concentration parameter, \(\mu^2\), and \(\omega\), fixing \(\nu^2=1\), \(h=0.1\) and \(d_Z=5\). We let  \(\varrho_{\mathrm{true}}=\omega/\sqrt{h^2+\omega^2}\) denote the true value of \(\varrho\) off the null.
For values of the concentration parameter, \(\mu^2\), we indicate the value of the first-stage F-statistic that would reject the hypothesis \(\mu^2\leq\bar \mu^2\) at level \(\alpha=0.05\) in the panel header. Recall from \Cref{lem:pd} that \(\xi\) is a deterministic function of \(\omega,h^2,\nu^2\) and \(\mu^2\), and hence pinned down by these parameters. The x-axis is scaled relative to the value of the nuisance parameters, in order for curves to take comparable shapes across panels.

Considering \Cref{fig:lr-regular-asy-all-power}, three  observations become apparent. First, the power curve of the recentered TLR statistic is close to being symmetric, and the test is close to being unbiased, for most of the parameter space. The exception is for extreme realizations of  \(\varrho_{\mathrm{true}}\) and \(\xi\). Second, we see that unlike the TSLS Wald test, both the TLR and recentered TLR tests reject no more than 5\% of the time at the null, \(\beta=\beta_0\). The TSLS Wald test, on the other hand, overrejects under the null at small \(\xi\) and large \(\varrho\). Third, we see that as the instrument set grows stronger, the power curves of the three tests converge to the same limit. This suggests that the power of the TLR and recentered TLR tests asymptote to the power of the Wald test.





































\section{Conclusion}
\label{sec:conclusion}

In this paper we have shown that existing weak-instrument robust tests are invalid if used as tests for the TSLS estimand in a model with treatment effect heterogeneity and more than one instrument. In particular, we have shown that the AR, Kleibergen--Moreira LM and Moreira CLR tests can overreject at any level of instrument strength, and that the degree of overrejection is increasing as the instrument set grows stronger.

We have further shown that there exists a test statistic that can be leveraged to produce  valid inference with treatment effect heterogeneity. In particular, we have constructed the TSLS likelihood-ratio (TLR) statistic for hypotheses on the TSLS estimand, and derived its limiting distribution. We have shown that when combined  with a pretest in a two-step procedure, the test obtains uniform validity. Choosing a small first-step level, we have shown in simulations that the test inherits the properties of the Wald test when the instrument set is strong.





\clearpage

\begin{thebibliography}{}

\bibitem[Anderson and Rubin, 1949]{andersonEstimationParametersSingle1949}
Anderson, T.~W. and Rubin, H. (1949).
\newblock Estimation of the parameters of a single equation in a complete system of stochastic equations.
\newblock {\em The Annals of Mathematical Statistics}, 20(1):46--63.

\bibitem[Andrews et~al., 2006]{andrewsOptimalTwoSidedInvariant2006}
Andrews, D. W.~K., Moreira, M.~J., and Stock, J.~H. (2006).
\newblock Optimal two-sided invariant similar tests for instrumental variables regression.
\newblock {\em Econometrica}, 74(3):715--752.

\bibitem[Andrews et~al., 2019]{andrewsWeakInstrumentsInstrumental2019}
Andrews, I., Stock, J.~H., and Sun, L. (2019).
\newblock Weak instruments in instrumental variables regression: Theory and practice.
\newblock {\em Annual Review of Economics}, 11:727--753.

\bibitem[Angrist et~al., 2000]{angristInterpretationInstrumentalVariables2000}
Angrist, J.~D., Graddy, K., and Imbens, G.~W. (2000).
\newblock The interpretation of instrumental variables estimators in simultaneous equations models with an application to the demand for fish.
\newblock {\em The Review of Economic Studies}, 67(3):499--527.

\bibitem[Angrist and Imbens, 1995]{angristTwoStageLeastSquares1995}
Angrist, J.~D. and Imbens, G.~W. (1995).
\newblock Two-stage least squares estimation of average causal effects in models with variable treatment intensity.
\newblock {\em Journal of the American Statistical Association},
  90(430):431--442.

\bibitem[Berger and Boos, 1994]{bergerValuesMaximizedConfidence1994}
Berger, R.~L. and Boos, D.~D. (1994).
\newblock P values maximized over a confidence set for the nuisance parameter.
\newblock {\em Journal of the American Statistical Association},
  89(427):1012--1016.

\bibitem[Bound et~al., 1995]{boundProblemsInstrumentalVariables1995}
Bound, J., Jaeger, D.~A., and Baker, R.~M. (1995).
\newblock Problems with instrumental variables estimation when the correlation between the instruments and the endogenous explanatory variable is weak.
\newblock {\em Journal of the American Statistical Association},
  90(430):443--450.

\bibitem[Chernozhukov et~al.,
  2013]{chernozhukovIntersectionBoundsEstimation2013}
Chernozhukov, V., Lee, S., and Rosen, A.~M. (2013).
\newblock Intersection bounds: Estimation and inference.
\newblock {\em Econometrica}, 81(2):667--737.

\bibitem[Chyn et~al., 2025]{chynExaminerJudgeDesigns2025}
Chyn, E., Frandsen, B., and Leslie, E. (2025).
\newblock Examiner and judge designs in economics: A practitioner's guide.
\newblock {\em Journal of Economic Literature}, 63(2):401--439.

\bibitem[Evdokimov and Koles{\'a}r,
  2018]{evdokimovInferenceInstrumentalVariables2018}
Evdokimov, K.~S. and Koles{\'a}r, M. (2018).
\newblock Inference in instrumental variables analysis with heterogeneous treatment effects.
\newblock Working Paper.

\bibitem[Finlay and Magnusson,
  2009]{finlayImplementingWeakInstrumentRobust2009}
Finlay, K. and Magnusson, L.~M. (2009).
\newblock Implementing weak-instrument robust tests for a general class of instrumental-variables models.
\newblock {\em The Stata Journal}, 9(3):398--421.

\bibitem[Fuller, 1977]{fullerPropertiesModificationLimited1977}
Fuller, W.~A. (1977).
\newblock Some properties of a modification of the limited information estimator.
\newblock {\em Econometrica}, 45(4):939--953.

\bibitem[{Goldsmith-Pinkham} et~al.,
  2025]{goldsmith-pinkhamLeniencyDesignsOperators2025}
{Goldsmith-Pinkham}, P., Hull, P., and Koles{\'a}r, M. (2025).
\newblock Leniency designs: An operator's manual.
\newblock Working Paper arXiv:2511.03572, arXiv.

\bibitem[Hansen, 1982]{hansenLargeSampleProperties1982}
Hansen, L.~P. (1982).
\newblock Large sample properties of generalized method of moments estimators.
\newblock {\em Econometrica}, 50(4):1029--1054.

\bibitem[Hansen, 2005]{hansenTestSuperiorPredictive2005}
Hansen, P.~R. (2005).
\newblock A test for superior predictive ability.
\newblock {\em Journal of Business \& Economic Statistics}, 23(4):365--380.

\bibitem[Heckman et~al., 2006]{heckmanUnderstandingInstrumentalVariables2006}
Heckman, J.~J., Urzua, S., and Vytlacil, E. (2006).
\newblock Understanding instrumental variables in models with essential heterogeneity.
\newblock {\em The Review of Economics and Statistics}, 88(3):389--432.

\bibitem[Imbens and Angrist, 1994]{imbensIdentificationEstimationLocal1994b}
Imbens, G.~W. and Angrist, J.~D. (1994).
\newblock Identification and estimation of local average treatment effects.
\newblock {\em Econometrica}, 62(2):467--475.

\bibitem[Kleibergen, 2002]{kleibergenPivotalStatisticsTesting2002}
Kleibergen, F. (2002).
\newblock Pivotal statistics for testing structural parameters in instrumental variables regression.
\newblock {\em Econometrica}, 70(5):1781--1803.

\bibitem[Koles{\'a}r, 2013]{kolesarEstimationInstrumentalVariables2013}
Koles{\'a}r, M. (2013).
\newblock Estimation in an instrumental variables model with treatment effect heterogeneity.
\newblock Working Paper.

\bibitem[Lee et~al., 2022]{leeValidTRatioInference2022}
Lee, D.~S., McCrary, J., Moreira, M.~J., and Porter, J.~R. (2022).
\newblock Valid t-ratio inference for IV.
\newblock {\em American Economic Review}, 112(10):3260--3290.

\bibitem[Lee et~al., 2023]{leeWhatWhenYou2023}
Lee, D.~S., McCrary, J., Moreira, M.~J., Porter, J.~R., and Yap, L. (2023).
\newblock What to do when you can't use '1.96' confidence intervals for IV.
\newblock NBER Working Paper 31893.

\bibitem[Loh, 1985]{lohNewMethodTesting1985}
Loh, W.-Y. (1985).
\newblock A new method for testing separate families of hypotheses.
\newblock {\em Journal of the American Statistical Association},
  80(390):362--368.

\bibitem[McCloskey,
  2017]{mccloskeyBonferronibasedSizecorrectionNonstandard2017}
McCloskey, A. (2017).
\newblock Bonferroni-based size-correction for nonstandard testing problems.
\newblock {\em Journal of Econometrics}, 200(1):17--35.

\bibitem[Mogstad and Torgovitsky,
  2024]{mogstadInstrumentalVariablesUnobserved2024}
Mogstad, M. and Torgovitsky, A. (2024).
\newblock Instrumental variables with unobserved heterogeneity in treatment effects.
\newblock NBER Working Paper 32927.

\bibitem[Mogstad et~al., 2021]{mogstadCausalInterpretationTwoStage2021}
Mogstad, M., Torgovitsky, A., and Walters, C.~R. (2021).
\newblock The causal interpretation of two-stage least squares with multiple instrumental variables.
\newblock {\em American Economic Review}, 111(11):3663--3698.

\bibitem[Moreira, 2003]{moreiraConditionalLikelihoodRatio2003}
Moreira, M.~J. (2003).
\newblock A conditional likelihood ratio test for structural models.
\newblock {\em Econometrica}, 71(4):1027--1048.

\bibitem[Moreira, 2009]{moreiraTestsCorrectSize2009}
Moreira, M.~J. (2009).
\newblock Tests with correct size when instruments can be arbitrarily weak.
\newblock {\em Journal of Econometrics}, 152(2):131--140.

\bibitem[Nelson and Startz, 1990a]{nelsonDistributionInstrumentalVariables1990}
Nelson, C.~R. and Startz, R. (1990a).
\newblock The distribution of the instrumental variables estimator and its t-ratio when the instrument is a poor one.
\newblock {\em The Journal of Business}, 63(1):S125--S140.

\bibitem[Nelson and Startz, 1990b]{nelsonFurtherResultsExact1990}
Nelson, C.~R. and Startz, R. (1990b).
\newblock Some further results on the exact small sample properties of the instrumental variable estimator.
\newblock {\em Econometrica}, 58(4):967--976.

\bibitem[Romano et~al., 2014]{romanoPracticalTwoStepMethod2014}
Romano, J.~P., Shaikh, A.~M., and Wolf, M. (2014).
\newblock A practical two-step method for testing moment inequalities.
\newblock {\em Econometrica}, 82(5):1979--2002.

\bibitem[Sargan, 1958]{sarganEstimationEconomicRelationships1958}
Sargan, J.~D. (1958).
\newblock The estimation of economic relationships using instrumental variables.
\newblock {\em Econometrica}, 26(3):393--415.

\bibitem[Serfling, 1980]{ApproximationTheoremsMathematical1980}
Serfling, R.~J. (1980).
\newblock {\em Approximation theorems of mathematical statistics}.
\newblock Wiley, New York, 1st edition.

\bibitem[Sigstad, 2026]{sigstadMonotonicityJudgesEvidence2026}
Sigstad, H. (2026).
\newblock Monotonicity among judges: Evidence from judicial panels and consequences for judge IV designs.
\newblock {\em American Economic Review}, 116(1):189--208.

\bibitem[Silvapulle, 1996]{silvapulleTestPresenceNuisance1996}
Silvapulle, M.~J. (1996).
\newblock A test in the presence of nuisance parameters.
\newblock {\em Journal of the American Statistical Association},
  91(436):1690--1693.

\bibitem[Staiger and Stock, 1997]{staigerInstrumentalVariablesRegression1997}
Staiger, D. and Stock, J.~H. (1997).
\newblock Instrumental variables regression with weak instruments.
\newblock {\em Econometrica}, 65(3):557--586.

\bibitem[Stern and Wolkowicz, 1995]{sternIndefiniteTrustRegion1995}
Stern, R.~J. and Wolkowicz, H. (1995).
\newblock Indefinite trust region subproblems and nonsymmetric eigenvalue perturbations.
\newblock {\em SIAM Journal on Optimization}, 5(2):286--313.

\bibitem[{Van de Sijpe} and Windmeijer,
  2023]{vandesijpePowerConditionalLikelihood2023}
{Van de Sijpe}, N. and Windmeijer, F. (2023).
\newblock On the power of the conditional likelihood ratio and related tests for weak-instrument robust inference.
\newblock {\em Journal of Econometrics}, 235(1):82--104.

\bibitem[Vytlacil, 2002]{vytlacilIndependenceMonotonicityLatent2002a}
Vytlacil, E. (2002).
\newblock Independence, monotonicity, and latent index models: An equivalence result.
\newblock {\em Econometrica}, 70(1):331--341.

\bibitem[Yap, 2025]{yapInferenceManyWeak2025}
Yap, L. (2025).
\newblock Inference with many weak instruments and heterogeneity.
\newblock Working Paper.

\end{thebibliography}

\clearpage