EconBase
← Back to paper

Extrapolating LATE with Weak IVs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

117,517 characters · 0 sections · 81 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Extrapolating LATE with Weak IVs

bibunit\begin{abstract} To evaluate the effectiveness of a counterfactual policy, it is often necessary to extrapolate treatment effects on compliers to broader populations. This extrapolation relies on exogenous variation in instruments, which is often weak in practice. This limited variation leads to invalid confidence intervals that are typically too short and cannot be accurately detected by classical methods. For instance, the $F$-test may falsely conclude that the instruments are strong. Consequently, I develop inference results that are valid even with limited variation in the instruments. These results lead to asymptotically valid confidence sets for various linear functionals of marginal treatment effects, including LATE, ATE, ATT, and policy-relevant treatment effects, regardless of identification strength. This is the first paper to provide weak instrument robust inference results for this class of parameters. Finally, I illustrate my results using data from agan/doleac/harvey:2023 to analyze counterfactual policies of changing prosecutors' leniency and their effects on reducing recidivism. \begin{description} • Keywords: Identification-robust Inference, Subvector Inference, Marginal Treatment Effects, Examiner Design • JEL classification: C12, C21, C26. \end{description} \end{abstract} \doparttoc \faketableofcontents \part \section{Introduction} This paper provides the first formal treatment of weak identification analysis in the marginal treatment effect (MTE) model. Originally developed in the seminal works of bjorklund/moffitt:1987 and heckman/vytlacil:1999, heckman/vytlacil:2001, heckman/vytlacil:2005, the MTE model has been widely adopted in various studies for extrapolating treatment effects, such as returns to schooling moffitt:2008, carneiro/heckman/vytlacil:2011, analysis of recidivism effects bhuller/dahl/loken/mogstad:2020, agan/doleac/harvey:2023, and evaluation of social insurance programs maestas/mullen/strand:2013, aizawa/mommaerts/corina/rennane:2023 (see Table 6 of mogstad/torgovitsky:2024 for a broad survey of MTE applications). Despite its widespread use in applied economic research, traditional confidence sets for the extrapolated causal effects within this model are often too short when the probability of receiving treatment (i.e., the propensity score) exhibits limited variation across the instrument’s support. Moreover, this variation can be very weak even if the usual $F$-test statistic is very large. To achieve valid coverage of causal effects, this paper establishes the first set of inference results that are robust against weak IV variation in MTE models. For linear MTE models, I develop an asymptotically similar conditional Wald test that delivers uniformly valid confidence sets with exact coverage. For a more general class of polynomial MTE models, I propose a modified linear combination (MLC) test that produces uniformly valid confidence sets while achieving approximate asymptotic efficiency under strong identification. Additionally, I highlight limitations of the additive separability assumption, a functional form often used to address weak variation of propensity scores, by deriving explicit formulas for the bias of estimands when this specification is incorrectly specified. \subsection*{Intuition for weak IVs in MTE models} To demonstrate the consequences of weak IVs in MTE models and to illustrate why the $F$-statistic fails, I plot regression estimates of the outcome variable on propensity scores for treated samples in Figure (ref). These estimators are computed from independent simulations based on a cubic MTE model with a discrete IV that takes four values. As will be shown in section (ref), the MTE can be directly recovered from the conditional regression, allowing us to study the weak IV problem by examining the finite-sample behavior of the estimates in Figure (ref). Figures (ref) through (ref) demonstrate that as propensity score variation diminishes, the regression estimates (shown in gray) become increasingly volatile and diverge from the true population quantity (shown in blue). This behavior implies that the MTE estimator loses consistency when propensity scores exhibit limited variation, which is a common consequence with instruments that weakly influence treatment selection. As a result, the traditional Wald confidence interval, which build on these compromised estimates, may fail to achieve its desired coverage probability. These estimation and inference problems arise not only when propensity scores cluster around one value but also when they approximate binary variation (Figure (ref)) under the cubic MTE design. In such cases, classical confidence intervals become unreliable, but the $F$-statistic can be misleadingly large because it detects the deviation from the null where all propensity scores are equal. \begin{figure}[h!] \caption{Estimators under Different Types of IV Strengths} \begin{subfigure}[t]{0.49\textwidth} \caption{Strong Variation} \end{subfigure} \begin{subfigure}[t]{0.49\textwidth} \caption{Intermediate Variation} \end{subfigure} \vskip\baselineskip \begin{subfigure}[t]{0.49\textwidth} \caption{Limited Variation} \end{subfigure} \begin{subfigure}[t]{0.49\textwidth} \caption{Binary Variation} \end{subfigure} \begin{tablenotes} Note: The figure shows the estimated expected outcomes for treated units conditional on propensity scores under different designs of propensity score variation, based on 20 independent simulations. All designs and simulations use a cubic specification for the MTE curve with a discrete IV that takes four values in its support. The sample size is 2,000. \end{tablenotes} \end{figure} \subsection*{Organization of this paper} In section (ref), I describe the setup in brinch/mogstad/wiswall:2017 and kline/walters:2019 and present the MTE model with a discrete instrument. By imposing a parametric structure on the marginal treatment response (MTR) functions, the MTE is fully characterized by a finite-dimensional parameter, which can be point identified using discrete variation from the instrument. In finite samples, the MTE can be estimated by running separate regressions for both treated and control groups. However, these separate regression estimators, along with the corresponding Wald confidence sets, are highly vulnerable to limited variation in propensity scores. In light of this potential identification failure, this paper constructs uniformly valid confidence sets for causal parameters that are linear functionals of the MTE function, which cover a broad range of causal effects of interest including the MTE function itself, the average treatment effect (ATE), the average treatment effect on the treated (ATT), the local average treatment effect (LATE), and the policy-relevant treatment effect (PRTE). In section (ref), I develop a simple inference method for treatment effects in the linear MTE model, where the MTR functions vary linearly with the selection unobservable brinch/mogstad/wiswall:2017. This special structure offers a novel moment condition for constructing estimators of treatment effects of interest, bypassing the necessity for estimating the full model. Building on this new moment condition, I construct a simple conditional Wald test and demonstrate its uniform validity under weak identification. This approach circumvents challenges posed by weakly identified nuisance parameters in the MTE, thus leading to confidence sets that achieve asymptotic similarity, a key advantage over existing approaches to subvector inference with weak identification as discussed in Online Appendix (ref). For a generic MTE model with a discrete instrument, it is infeasible to directly estimate causal effects of interest without relying on other primitive parameters in the MTE function. In such cases, these weakly identified primitive parameters hamper the ability to perform valid inference on linear functions of them. In section (ref), I build on the improved projection approach developed in I. andrews.i:2018 and propose a MLC test that achieves uniform validity regardless of identification strength (see page 7 for a detailed comparison of my work with the literature). The inverted confidence set from this test has correct coverage under weak identification, is straightforward to compute, and is shown to be approximately as efficient as the Wald confidence set under strong identification. In section (ref), I discuss the issue of incorporating covariates into the MTE model. First, I demonstrate that the commonly used additive separability condition can substantially bias causal effect estimands when it fails. Such bias does not vanish unless (1) unobserved treatment effect heterogeneity does not vary across covariates, or (2) individuals' treatment decisions do not depend on those covariates. If neither of these two assumptions hold, I show that causal effect estimands under additive separability differ from the true effects and potentially have the opposite sign. To deal with the bias induced by imposing additive separability, researchers may use the proposed methods in this paper to conduct inference by conditioning on covariates, while also achieving robustness against limited variation of propensity scores from conditioning. In Online Appendix (ref), I consider a Bonferroni-type size correction for valid inference on aggregated effects in the absence of additive separability. For researchers interested in performing inference under additive separability, I also present an extension of the proposed inference methods to accommodate this framework in Online Appendix (ref). In section (ref), I examine the performance of MLC tests via a sequence of Monte-Carlo simulations in a quadratic MTE model with varying degrees of identification strength. The simulation results show that the MLC test is asymptotically size-correct whenever the model is strongly identified, weakly identified, or partially identified. In contrast, the classical Wald test reports over-rejections of the null value at 14% frequency versus the 5% significance level under partial identification and exhibits trivial power under weak identification. When the instrument strength is sufficiently strong, the proposed MLC test has power close to the asymptotically efficient Wald test. Theoretical justification is provided in Online Appendix (ref). In section (ref), I illustrate the proposed methods by using data from agan/doleac/harvey:2023, who show that not prosecuting defendants with misdemeanor offenses reduces recidivism. While their analysis uses LATE estimates and ATE/ATT estimates under additive separability, I complement their findings by examining recidivism effects under different counterfactual prosecution policies. Specifically, I construct confidence sets for the reduction in recidivism under two scenarios: (1) a homogeneous marginal increase in nonprosecution rates across prosecutors, and (2) implementation of a minimum nonprosecution rate threshold. The results highlight that weak identification is a significant concern in higher-order polynomial models, and empirical conclusions differ substantially between robust and classical inference methods. Section (ref) concludes with a discussion of potential future extensions. \subsection*{Related literature} This paper contributes to three strands of the literature: marginal treatment effects, subvector inference in weakly identified models, and judge/examiner designs. For the rest of the introduction, I review the related literature. Although the MTE model has been used to extrapolate treatment effects in a variety of fields, there are relatively few studies on the theory of estimation and inference in MTE models. In a semiparametric MTE model with continuous propensity scores, heckman/urzua/vytlacil:2006a use bootstrap methods to construct confidence bands for MTE functions. This approach is now common practice\footnote{The Stata implementation of MTE estimation by brave/walstrum:2014 and andresen:2018 uses bootstrap methods to produce confidence sets.}, but its theoretical validity has not yet been established. carneiro/lee:2009 analyze the pointwise limit distributions of MTE estimators using a separate regression approach. Building on this work, sasaki/ura:2023 further derive asymptotic theory for the PRTE using an orthogonalized score for double debiased estimation. Both papers assume a semiparametric MTE model with additive separability and strong IV/covariate variation to achieve $\sqrt{n}$-consistent estimation. mogstad/santos/torgovitsky:2017, mogstad/santos/torgovitsky:2018 propose a bootstrap procedure for conducting inference on the PRTE when the MTE is partially identified. None of these papers study weak identification problems. Thus, compared to this existing literature, the inference approach I propose is the first to achieve robustness against weak instruments in MTE models. To estimate the MTE over a large support of propensity scores, empirical researchers usually assume that the MTR functions---and their difference, the MTE---are additively separable in covariates and unobserved costs of treatment selection brinch/mogstad/wiswall:2017. This assumption allows the MTE to be estimated on the unconditional support of propensity scores, by pooling variation in the propensity scores across different covariate values. Despite its increasing popularity in empirical work, there is little theoretical justification for additive separability. In fact, this assumption can be strong enough to point identify the MTE without exogenous variation from instruments pan/wang/zhang/zhou:2024. In addition, if the MTE is misspecified as linearly additively separable, devereux:2022 provides numerical evidence indicating that omitting higher-order covariate terms can introduce significant bias in the estimated MTE slopes. I contribute to this knowledge by providing the first analytical bias formula for the causal parameters (e.g., ATE, conditional ATE, and MTE slope) when additive separability is misspecified in a widely used class of latent threshold crossing models kline/walters:2019. To avoid this bias, researchers should extrapolate treatment effects conditional on covariates instead of relying on additive separability. Moreover, the robust inference procedures proposed in this paper can help address potential weak IV variation that may arise after conditioning on covariates. The inference problem in this paper is also related to the literature on subvector inference in weakly identified models, where a subvector is a subset of the structural parameters. For inference on subvectors, dufour/mohamed:2005 suggest projecting the robust confidence sets of the full vector onto the subvector of interest. However, this procedure can be very conservative especially when the dimension of the full vector is much larger than that of the target subvector. To this end, there is a sequence of studies trying to reduce the conservativeness of projection inference. chaudhuri/zivot:2011 consider modifying the (Robust) Lagrangian Multiplier statistic kleibergen:2005 such that it is locally equivalent to asymptotic efficient subvector tests under strong identification and propose a Bonferroni method to improve the power of their tests at distant alternatives. Building on this idea, D. andrews.d:2017 improves the Bonferroni method such that the refined tests are asymptotically non-conservative and uniformly valid. However, D. andrews.d:2017 acknowledges six key limitations, including computational challenges and the need for additional tuning parameters to categorize the identification strength, neither of which are required for my proposed MLC test. Moreover, his method does not directly address inference on a linear function of parameters, which is the primary focus of this paper. While his method could conceptually extend to inference on a linear function through model reparametrization, finding a universal reparametrization rule that works for all linear hypotheses of interest while maintaining tractable asymptotic analysis remains challenging. For inference on a function of parameters in a weakly identified model, I. andrews.i:2018 generalizes the results in chaudhuri/zivot:2011 and proposes a two-step confidence set that achieves sequential validity with controlled coverage distortions under a set of high-level conditions imposed on a sequence of data generating processes (DGPs). The MLC test considered in this paper builds on the idea from I. andrews.i:2018 but differs in a few ways: (1) Most importantly, Andrews’ paper only provided sequential validity results, concluding that “conditions for uniform asymptotic validity are an interesting open question’’ in Section V (page 347). I demonstrate that, with a minor modification to the test statistics, the robust test achieves uniform validity for inference on a scalar function. (2) While Andrews derived results for a general class of models using high-level conditions, my focus on the MTE model allows for more specific, primitive conditions for validity. (3) Instead of employing Andrews' two-step approach that alternates between non-robust Wald and robust confidence sets based on identification strength, I focus exclusively on his robust confidence sets. (4) I adapt his robust confidence sets, originally developed for GMM models, to the minimum distance framework arising from separate regressions in the MTE setting. (5) By using the robust confidence set, this approach does not involve pretesting distortion in the two-step approach and maintains the desired asymptotic coverage level $1-\alpha$. (6) I prove that the local power difference from the asymptotically efficient Wald test can be arbitrarily small under strong identification (see Online Appendix (ref)). Applied to MTE models, this approach allows for valid and powerful inference on treatment effects even under limited variation of propensity scores. In Online Appendix (ref), I also discuss other approaches to inference on functions (or subvectors) of parameters in weakly identified models and explain their inapplicability to the MTE model studied here. Finally, the empirical analysis of this paper also connects to the judge/examiner design problems (see the survey by chyn/frandsen/leslie:2024). By using the quasi-random assignment of examiners, researchers can identify the causal effects of judicial decisions for defendants at the margin of being treated, and then extrapolate these effects to the broader population under the assumption of pairwise monotonicity frandsen/lefgren/leslie:2023. While the MTE model is commonly employed for such extrapolation, researchers often report LATE, ATE, ATUT, and ATT to inform potential policy decisions bhuller/dahl/loken/mogstad:2020, agan/doleac/harvey:2023, baron/gross:2023. My paper provides methods for robust inference on all of these parameters, as well as the MTE function itself, and the policy counterfactual parameters studied in heckman/vytlacil:2001, heckman/vytlacil:2005, and carneiro/heckman/vytlacil:2010. \section{MTE Model with Discrete IVs} In this section, I describe the MTE model and the related identification result following brinch/mogstad/wiswall:2017. Based on this result, the weakness of IVs can be characterized by limited variation of propensity scores. Then I introduce the relevant parameter space on which we achieve uniform validity of the proposed inference procedures. \subsection{Setup} Let $Y_1$ be the potential outcome of an individual who receive a binary treatment ($D = 1$) and $Y_0$ denote her potential outcome in the untreated state ($D = 0$). The observed outcome $Y$ is realized through \begin{equation} Y = (1-D)Y_0 + DY_1. \end{equation} We further specify \[ Y_d = \mu_d + V_d, \quad d = 0,1, \] where $\mu_d \equiv \ensuremath{\mathbb{E}}[Y_d]$ is the mean of potential outcome. For simplicity, I leave out additional covariates and discuss them in section (ref). The treatment is determined by a weakly separable selection equation \begin{equation} D = \ensuremath{\mathbbm{1}}[U \leq \nu(Z)], \end{equation} where $\nu(\cdot)$ is an unknown function, $U$ is a continuous random variable representing the unobserved cost of selection into treatment, and $Z \in \operatorname*{supp}(Z) = \{z_0, z_1, \ldots, z_K\}$ is the excluded discrete instrument. Researchers observe the outcome $Y$, the binary treatment $D$, and the excluded instrument $Z$ from data, while the unobservables are the potential outcomes $(Y_0, Y_1)$ and the variable $U$ in the selection equation. The MTE model allows individuals to be selected into treatment based on their information on potential outcomes, which leads to potential dependence between $(V_1, V_0)$ and $U$. The key identifying assumption from brinch/mogstad/wiswall:2017 is as follows: \begin{assumption}[MTE model with discrete IVs] \begin{enumerate} • $Z \mathbin{ \mathpalette{\@indep}{} } U$. • $\ensuremath{\mathbb{E}}[Y_d\mid Z, U] = \ensuremath{\mathbb{E}}[Y_d\mid U]$ and $\ensuremath{\mathbb{E}}|Y_d| < \infty$ for $d \in \{0,1\}$. • $U$ is continuously distributed. • $0 < \ensuremath{\mathbb{P}}(D=1\mid Z=z) <1$ for all $z \in \operatorname*{supp}(Z)$. • Let $\{h_m(\cdot)\}_{m=1}^{M}$ be a set of known continuous functions defined on $(0,1)$. For $d \in \{0,1\}$, the MTR function is given by \[ \ensuremath{\mathbb{E}}[Y_d\mid F_U(U) = u] = \mu_d + \sum_{m=1}^M \rho_{dm} h_m(u)\quad \text{for } u \in (0,1), \] where $F_U(\cdot)$ denotes the distribution function of $U$. • Let $\lambda_{00}(\cdot) = \lambda_{10}(\cdot) \equiv 1$, and define \[ \lambda_{1m}(p) \equiv \frac{1}{p}\int_{0}^p h_m(u) du \quad \text{and} \quad \lambda_{0m}(p) \equiv \frac{1}{1-p}\int_{p}^1 h_m(u) du \quad \text{for $m = 1,\ldots,M$}. \] for $p \in [0,1]$. Then $\{\lambda_{1m}(\cdot)\}_{m=0}^M$ and $\{\lambda_{0m}(\cdot)\}_{m=0}^M$ are unisolvent\footnote{A set of $n$ functions $f_1, f_2, \ldots, f_n$ is unisolvent on domain $\Omega$ if the matrix $F \in \ensuremath{\mathbb{R}}^{n\times n}$ with entries $f_i(x_j)$ has nonzero determinant for any choice of $n$ distinct points $x_1, x_2, \ldots, x_n$ in $\Omega$.} on $(0,1)$. • $\overline{K} \equiv |\{\ensuremath{\mathbb{P}}(D=1\mid Z=z): z=z_0, z_1, \ldots, z_K\}|\geq M+1$. \end{enumerate} \end{assumption} Assumptions (ref).1 and (ref).2 require the excluded instrument $Z$ to be exogenous to both the selection and outcome processes. Assumption (ref).3 allows us to normalize the marginal distribution of $U$ to be uniformly distributed over $[0,1]$. That is, we can transform $U$ to a uniformly distributed variable $\widetilde{U} = F_{U}(U)$. Under the exogeneity of $Z$ in Assumption (ref).1, the function $F_U(\nu(Z))$ can then be interpreted as propensity score: \begin{align*} p(z) &\equiv \ensuremath{\mathbb{P}}(D=1\mid Z = z) \\ &= \ensuremath{\mathbb{P}}(U\leq \nu(Z)\mid Z = z) \\ &= \ensuremath{\mathbb{P}}(U\leq \nu(z)) \\ &= F_U(\nu(z)) \end{align*} where the third line uses the Assumption (ref).1. Define $\widetilde{\nu}(z) \equiv F_U(\nu(Z))$, and then we can work with $(\widetilde{U}, \widetilde{\nu}(z))$ in place of $(U, \nu(z))$ without affecting the empirical content of the selection model. For simplicity, we drop out the tilde and assume $U$ is uniformly distributed throughout our analysis, and therefore $\nu(z) = p(z)$. Assumption (ref).4 validates the overlap condition, so that we observe both treated and untreated individuals for each group defined by the value of the instrument. Assumption (ref).5 imposes a parametric restriction on the MTR function $\ensuremath{\mathbb{E}}[Y_d\mid F_U(U)=u]$ so that the MTE can be extrapolated outside the discrete support of the instrument. Assumption (ref).6 is a weak condition that rules out redundant specifications in $\{h_m(\cdot)\}_{m=0}^M$ that cause multicollinearity in $\{\lambda_{dm}(\cdot)\}_{m=0}^M$. In particular, the usual polynomial specification $h_m(u) = u^m - \frac{1}{m+1}$ satisfies this condition. Assumption (ref).7 requires sufficient variation of the exogenous instrument to point identify the structural parameters. While $\overline{K} \leq K+1$, with strict inequality when multiple instrument values yield identical propensity scores, point identification requires the number of distinct propensity scores, not instrument values, to exceed the order of the MTE model. Define $\theta_d \equiv (\mu_d, \rho_{d1}, \ldots, \rho_{dM})'$ for $d = 0,1$, and let $\theta \equiv (\theta_1', \theta_0')'$. Then we have the following identification result: \begin{theorem}[Identification] Suppose Assumption (ref) holds. Then $\theta$ is point identified. \end{theorem} \begin{proof} Based on the model (ref)--(ref) and Assumption (ref), we have \begin{align} \ensuremath{\mathbb{E}}[Y\mid D=1, Z=z] &= \ensuremath{\mathbb{E}}[Y_1\mid U\leq p(z), Z=z] \notag \\ &= \ensuremath{\mathbb{E}}[Y_1\mid U\leq p(z)] \notag \\ &= \frac{1}{p(z)}\int_{0}^{p(z)} \ensuremath{\mathbb{E}}[Y_1\mid U=u] du \notag \\ &= \mu_1 + \sum_{m=1}^M \frac{\rho_{1m}}{p(z)}\int_{0}^{p(z)} h_m(u) du. \notag \\ &= (\lambda_{10}(p(z)), \lambda_{11}(p(z)), \ldots, \lambda_{1M}(p(z)))\theta_1. \end{align} The first line holds by equations (ref), (ref), and the normalization that $\nu(z) = p(z)$, which is jointly implied by Assumption (ref).1 and (ref).3. The second line holds by the exogeneity condition in Assumption (ref).2. The third line follows by the overlap condition in Assumption (ref).4 and the normalization that $U$ is uniformly distributed over $[0,1]$. The fourth line holds by parametric restriction in Assumption (ref).5. Likewise, we have \begin{equation} \ensuremath{\mathbb{E}}[Y\mid D=0, Z=z] = \mu_0 + \sum_{m=1}^M \frac{\rho_{0m}}{1-p(z)}\int_{p(z)}^{1} h_m(u) du = (\lambda_{00}(p(z)), \lambda_{01}(p(z)), \ldots, \lambda_{0M}(p(z)))\theta_0. \end{equation} Taking $z = z_0,z_1,\ldots,z_K$ in equations (ref) and (ref) then yields two matrix equalities: \begin{equation} \beta_d = A_d\theta_d \quad for d = 0,1, \end{equation} where \[ A_d = \begin{bmatrix} \lambda_{d0}(p(z_0)) & \lambda_{d1}(p(z_0)) & \ldots & \lambda_{dM}(p(z_0)) \\ \lambda_{d0}(p(z_1)) & \lambda_{d1}(p(z_1)) & \ldots & \lambda_{dM}(p(z_1)) \\ \vdots & \vdots & \ddots & \vdots \\ \lambda_{d0}(p(z_K)) & \lambda_{d1}(p(z_K)) & \hdots & \lambda_{dM}(p(z_K)) \end{bmatrix} \] and \[ \beta_d = (\ensuremath{\mathbb{E}}[Y\mid D=d, Z=z_0], \ldots, \ensuremath{\mathbb{E}}[Y\mid D=d, Z=z_K])'. \] Note that Assumption (ref).6 and (ref).7 guarantees that there exists a full-rank submatrix of $A_d$ with $\overline{K}$ rows and $M+1$ columns for each $d = 0,1$. Therefore, both $A_1$ and $A_0$ have full column rank. Then $\theta_1$ and $\theta_0$ are point identified by equation (ref). \end{proof} \begin{remark} This identification result closely aligns with the proof of brinch/mogstad/wiswall:2017, who shows that Assumption (ref).7 is a necessary condition for identifying $\theta$ in a MTE model satisfying Assumption (ref).1--(ref).5. Building on their result, I add Assumption (ref).6 to establish sufficient conditions for point identification. \end{remark} Based on the parametric restriction in Assumption (ref).5, treatment effects can often be written as linear functions of the primitive parameter $\theta$. For example, $\text{ATE} = \mu_1 - \mu_0$ and \begin{align*} \operatorname*{MTE}(u) &= \ensuremath{\mathbb{E}}[Y_1 - Y_0\mid U=u] \\ &= c(u)'\theta_1 - c(u)'\theta_0, \end{align*} where $c(u) = (1, h_1(u),\ldots,h_M(u))$. Moreover, suppose we are interested in the parameters that inform potential policy decisions such as average treatment effects on the (un)treated groups or policy-relevant treatment effects. In that case, the weight $c$ is often unknown and depends on the underlying DGP (see Table 1 and 2 in mogstad/torgovitsky:2018). If we write $c = (c_1', c_0')'$ with $c_{d} = (c_{d,0},\ldots,c_{d,M})'$ for $d = 0,1$ and denote $h_0(u) \equiv 1$, Table (ref) summarizes the weights of various treatment effects under Assumption (ref). In this paper, I consider inference on treatment effects of the form $c'\theta$ where $c_1 = -c_0$ since many causal effects of interest including those in Table (ref) have symmetric weights: $c_{1,m} = -c_{0,m}$ for all $m = 1,\ldots,M$. \begin{table}[ht] \caption{Weights for Treatment Effects} \begin{tabular}{ccl} \toprule Target parameter & Expression & Weights \\ \midrule ATE & $\ensuremath{\mathbb{E}}[Y_1 - Y_0]$ & $c_{1,m} = \ensuremath{\mathbbm{1}}[m=0]$ \\ MTE & $\ensuremath{\mathbb{E}}[Y_1 - Y_0\mid U=u]$ & $c_{1,m} = h_m(u)$\\ ATT & $\ensuremath{\mathbb{E}}[Y_1 - Y_0\mid D=1]$ & $c_{1,m} = \frac{\ensuremath{\mathbb{E}}(\int_0^{p(Z)}h_m(u)du)}{\ensuremath{\mathbb{P}}(D=1)}$ \\ ATU & $\ensuremath{\mathbb{E}}[Y_1 - Y_0\mid D=0]$ & $c_{1,m} = \frac{\ensuremath{\mathbb{E}}(\int_{p(Z)}^1 h_m(u)du)}{\ensuremath{\mathbb{P}}(D=0)}$ \\ LATE & $\ensuremath{\mathbb{E}}[Y_1 - Y_0\mid p(z_0) < U < p(z_k)]$ & $c_{1,m} = \frac{\int_{p(z_0)}^{p(z_k)} h_m(u)du}{p(z_k) - p(z_0)}$ \\ Additive PRTE & PRTE with $p^\epsilon(z) = p(z) + \epsilon$ & \multirow{3}{*}{$c_{1,m} = \frac{\ensuremath{\mathbb{E}}[\int_{p(Z)}^{p^\epsilon(Z)} h_m(u) du]}{\ensuremath{\mathbb{E}}[p^\epsilon(Z) - p(Z)]}$} \\ Proportional PRTE & PRTE with $p^\epsilon(z) = (1+\epsilon) p(z)$ & \\ Quota & PRTE with $p^\epsilon(z) = \min\{p(z),\epsilon\}$ & \\ \bottomrule \end{tabular} \end{table} Define $q(z_\ell) = \ensuremath{\mathbb{P}}(Z=z_\ell)$, i.e., the probability mass function of the discrete IV. For the commonly used polynomial MTE model in empirical studies where $h_m = u^{m} - \frac{1}{m+1}$, the weights of ATT, LATE, additive PRTE, and proportional PRTE become { \begin{align*} &c^{ATT}_{1,m} = \frac{1}{m+1} \left(\frac{\sum_{\ell=0}^K p(z_\ell)^{m+1} q(z_\ell)}{\sum_{\ell=0}^K p(z_\ell)q(z_\ell)} - 1\right)\\ &c^{\text{LATE}}_{1,m} = \frac{1}{m+1} \left(\sum_{j=0}^m p(z_k)^j p(z_0)^{m-j} - 1\right)\\ &c^{\text{A-PRTE}}_{1,m} = \frac{1}{m+1} \left(\sum_{\ell = 0}^K q(z_\ell) \sum_{j = 0}^m \left[p(z_\ell)^{j} (p(z_\ell) + \epsilon)^{m-j}\right] - 1 \right) \\ &c^{\text{P-PRTE}}_{1,m} = \frac{1}{m+1} \left[\frac{(1+\epsilon)^{m+1} - 1}{\epsilon} \cdot \frac{\sum_{\ell = 0}^K p(z_\ell)^{m+1} q(z_\ell)}{\sum_{\ell = 0}^K p(z_\ell) q(z_\ell)} - 1\right] \end{align*} } for $m = 1,\ldots,M$, and the first element $c_{1,0} = 1$ for all the above causal effects. \subsection{Weak identification} The identification strategy relies on the conditions in Assumption (ref), particularly the existence of sufficiently many distinct propensity scores to point identify the structural parameter $\theta$ through Assumption (ref).7. However, variation in propensity scores is often limited. This limited variation may not be enough to guarantee correct asymptotic approximation in the construction of classical confidence sets. In section (ref) and (ref), I develop robust inference procedures for treatment effects of the form $c'\theta$ without requiring the potentially restrictive Assumptions (ref).7. My results have two implications. Along a sequence of DGPs: (1) When Assumption (ref).7 holds but is close to fail, the parameters are point identified, and the robust confidence set achieves correct coverage for the causal parameter of interest; (2) When Assumption (ref).7 fails, the parameters are partially identified, and the proposed robust confidence set mantains correct coverage of the causal parameter $c'\theta$, where $\theta$ lies in the set defined by the linear system (ref). Additionally, this uniform validity result does not rely on Assumption (ref).6, though this assumption is typically satisfied under researchers' common specifications of the MTR functions. \subsection{Parameter space restriction} This section introduces the parameter space for the joint distribution of $(Y,D,Z)$, on which the proposed robust inference procedures are uniformly valid (without covariates). First, consider the usual i.i.d. sampling distribution of the observables. \begin{assumption} The random vectors $(Y_i, D_i, Z_i)$ for $i = 1,\ldots,n$ are i.i.d. with distribution $F$. \end{assumption} Next, I introduce a set of regularity conditions to be imposed on the joint distribution $F$. \begin{definition}[Parameter Space] For some $\delta, \zeta > 0$ and $\epsilon \in (0,1/2)$, define the parameter space $\mathcal{P}$ as the set of pairs $(\theta,F)$ satisfying the following properties: \begin{enumerate} • Equation (ref) is satisfied with $K \geq M$, where $\theta = (\theta_1', \theta_0')' \in \operatorname{int}(\Theta) \subseteq \ensuremath{\mathbb{R}}^{2(M+1)}$ for some compact set $\Theta$ with nonempty interior; • $\sup_{d=0,1}\sup_{z\in \operatorname*{supp}(Z)}\ensuremath{\mathbb{E}}_F[|Y|^{2+\delta}\mid D=d, Z=z] \leq \zeta$; • $\epsilon \leq \inf_{z \in \operatorname*{supp}(Z)}\ensuremath{\mathbb{P}}_F(D=1\mid Z=z) \leq \sup_{z \in \operatorname*{supp}(Z)}\ensuremath{\mathbb{P}}_F(D=1\mid Z=z) \leq 1-\epsilon$; • $\epsilon \leq \inf_{z\in\operatorname*{supp}(Z)}\ensuremath{\mathbb{P}}_F(Z=z) \leq \sup_{z\in\operatorname*{supp}(Z)} \ensuremath{\mathbb{P}}_{F}(Z=z) \leq 1-\epsilon$; • $\epsilon \leq \inf_{d=0,1}\inf_{z\in\operatorname*{supp}(Z)}\operatorname{var}_F(Y\mid D=d, Z=z)$. \end{enumerate} \end{definition} The first condition assumes the correct model specification derived from Assumptions (ref).1-(ref).5, and the order of the MTE model does not exceed the support of the instrument.\footnote{The condition $K\geq M$ is necessary but not sufficient for point identification, particularly when Assumption (ref).7 does not hold.} The second condition is a mild restriction on the existence of moments of $Y$ conditional on $D$ and $Z$, which is essential for establishing the uniform version of the law of large numbers and central limit theorems. The third and fourth conditions, also known as the strong overlap conditions, require observing units for each value of the treatment and the instrument. Together with the final condition regarding sufficient variation in outcomes, these conditions rule out the singularity or near-singularity of the asymptotic variance of the moment equation (ref). Since our parameter of interest has the form of $c'\theta$, let $\mathcal{P}_0 = \{(\lambda,F): \lambda = c'\theta, (\theta,F)\in\mathcal{P}\}$ denote the null parameter space for this linear function, in which the weight $c$ could possibly depend on underlying DGP $F$. It is important to emphasize that our parameter space does not require the relevance Assumptions (ref).6 and (ref).7, the latter of which is likely to fail. For a given $\lambda \in \ensuremath{\mathbb{R}}$, our goal is to develop uniformly valid tests for assessing the hypothesis $H_0: c'\theta = \lambda$ that is robust to the limited variation of propensity scores. \subsection{Notation and preliminaries} The asymptotic behavior of test statistics to be proposed will be unified by the asymptotic distribution of estimators on marginal distribution of instrument $q(z_\ell) \equiv \ensuremath{\mathbb{P}}(Z=z_\ell)$, propensity score $p(z_\ell) = \ensuremath{\mathbb{P}}(D=1\mid Z=z_\ell)$, and the expected outcome conditional on the treatment and instrument $\beta_{d\ell} = \ensuremath{\mathbb{E}}[Y\mid D=d, Z=z_\ell]$ for $d=0,1$ and $\ell = 0,1,\ldots,K$. Under the parameter space restriction outlined in section (ref), we have the following sample-analog estimators for these quantities: \begin{align*} &\hat{q}(z_\ell) = \frac{1}{n}\sum_{i=1}^n \ensuremath{\mathbbm{1}}[Z_i =z_\ell] \\ &\hat{p}(z_\ell) = \frac{1}{n}\sum_{i=1}^n \frac{\ensuremath{\mathbbm{1}}[D_i = 1, Z_i = z_\ell]}{\hat{q}(z_\ell)} \\ &\hat{\beta}_{d\ell} = \frac{1}{n}\sum_{i=1}^n\frac{Y_i\ensuremath{\mathbbm{1}}[D_i = d, Z_i = z_\ell]}{\hat{q}(d,z_\ell)}, \end{align*} where $\hat{q}(1,z_\ell) = \hat{p}(z_\ell)\hat{q}(z_\ell)$ and $\hat{q}(0,z_\ell) = (1-\hat{p}(z_\ell))\hat{q}(z_\ell)$. Collecting these estimators into vectors, we have $\hat{p} = (\hat{p}(z_\ell), \ldots, \hat{p}(z_K))'$, $\hat{q} = (\hat{q}(z_0),\ldots,\hat{q}(z_K))'$, $\hat{\beta}_d = (\hat{\beta}_{d0}, \ldots, \hat{\beta}_{dK})'$ for $d = 0,1$, and $\hat{\beta} = (\hat{\beta}_1', \hat{\beta}_0')'$. Lemma (ref)(a) provides the asymptotic distribution of the estimators $\hat{p}, \hat{q}$, and $\hat{\beta}$ under a drifting sequence in the parameter space $\mathcal{P}$. Moreover, Lemma (ref)(b) also shows that their asymptotic variances can be consistently estimated by \begin{align*} &\hat{\Sigma}_p = \operatorname*{diag}\left\{\frac{\hat{p}(z_\ell)(1-\hat{p}(z_\ell))}{\hat{q}(z_\ell)}: \ell = 0,1,\ldots,K\right\} \\ &\hat{\Sigma}_q = \{\hat{\Sigma}_q[i,j]\}_{i,j = 0,1,\ldots,K} \\ &\hat{\Sigma}_{\beta_d} = \operatorname*{diag}\left\{\frac{\hat{\sigma}_{d\ell}^2}{\hat{q}(d,z_\ell)}: \ell = 0,1,\ldots,K\right\} \\ &\hat{\Sigma}_\beta = \operatorname*{diag}\{\hat{\Sigma}_{\beta_1}, \hat{\Sigma}_{\beta_0}\}, \end{align*} where \begin{align*} & \hat{\sigma}_{d\ell}^2 \equiv \frac{1}{n} \sum_{i=1}^n \frac{(Y_i - \hat{\beta}_{d\ell})^2\ensuremath{\mathbbm{1}}[D_i=d, Z_i=z_\ell]}{\hat{q}(d,z_\ell)} \\ & \hat{\Sigma}_q[i,j] \equiv \begin{cases} \hat{p}(z_i) (1-\hat{p}(z_i)) & \text{ if } i = j\\ -\hat{p}(z_i) \hat{p}(z_j) & \text{ if } i \neq j. \end{cases} \end{align*} Throughout the paper, I use $\partial_x f \in \ensuremath{\mathbb{R}}^{k}$ to denote the gradient of a scalar function $f$ with respect to its argument $x\in \ensuremath{\mathbb{R}}^k$. If $f$ maps to a vector in $\ensuremath{\mathbb{R}}^l$, then $\partial_x f \in \ensuremath{\mathbb{R}}^{l\times k}$ denotes the Jacobian of $f$ with respect to $x\in\ensuremath{\mathbb{R}}^k$. Let $P_A = A(A'A)^{-1}A'$ denote the projection matrix onto the column spaces of matrix $A$ and define $M_A = I - P_A$ as the corresponding annihilator matrix. Let $\operatorname*{supp}(X)$ denote the support of a random element $X$. \section{Robust Inference in Linear MTE Models} In this section, I start with the simplest functional specification and impose linearity on the MTR function. \begin{assumption}[Linear MTE model] Assumption (ref).5 holds with \[ h_m(u) = u - \frac{1}{2} \] and $M=1$. \end{assumption} This assumption was first introduced by olsen:1980 to characterize sample selection bias and later generalized by brinch/mogstad/wiswall:2017 to model MTE functions. This linear MTE specification was recently adopted by kowalski:2023 to extrapolate treatment effects between two experimental studies. The parameters $\theta = (\mu_1, \rho_{11}, \mu_0, \rho_{01})$ can be identified using any pair of propensity scores $(p(z_0), p(z_k))$ for $k = 1,\ldots,K$ that differ from each other. Point identification fails if and only if all propensity scores ${p(z_0),\ldots,p(z_K)}$ are identical, i.e., when there is no variation in propensity scores. The linear MTE specification allows us to derive closed-form estimands for $\theta$, in which case equation (ref) implies \begin{align*} \begin{pmatrix} \mu_1 \\ \rho_{11} \end{pmatrix} &= \frac{1}{p(z_k) - p(z_0)} \begin{pmatrix} p(z_k) - 1 & -(p(z_0) - 1) \\ -2 & 2 \end{pmatrix} \begin{pmatrix} \ensuremath{\mathbb{E}}[Y\mid D=1, Z=z_0] \\ \ensuremath{\mathbb{E}}[Y\mid D=1, Z=z_k] \end{pmatrix} \end{align*} and \begin{align*} \begin{pmatrix} \mu_0 \\ \rho_{01} \end{pmatrix} &= \frac{1}{p(z_k) - p(z_0)} \begin{pmatrix} p(z_k) & -p(z_0) \\ -2 & 2 \end{pmatrix} \begin{pmatrix} \ensuremath{\mathbb{E}}[Y\mid D=0, Z=z_0] \\ \ensuremath{\mathbb{E}}[Y\mid D=0, Z=z_k] \end{pmatrix}. \end{align*} Recall $\beta_{dk} = \ensuremath{\mathbb{E}}[Y\mid D=d, Z=z_k]$. Then the ATE and the slope of MTE can be written as \[ \begin{pmatrix} \mu_1 - \mu_0 \\ \rho_{11} - \rho_{01} \end{pmatrix} = \frac{1}{p(z_k) - p(z_0)} \begin{pmatrix} p(z_k)\left[\beta_{10} - \beta_{00}\right] - p(z_0)\left[\beta_{1k} - \beta_{0k}\right] + \beta_{1k} - \beta_{10} \\ 2\left(\beta_{00} - \beta_{10} + \beta_{1k} - \beta_{0k}\right) \end{pmatrix}. \] Denote the numerator of the right-hand-side equation by \[ \Delta_\mu(z_0, z_k) \equiv p(z_k)\left[\beta_{10} - \beta_{00}\right] - p(z_0)\left[\beta_{1k} - \beta_{0k}\right] + \beta_{1k} - \beta_{10} \] and \[ \Delta_\rho(z_0, z_k) \equiv 2\left(\beta_{00} - \beta_{10} + \beta_{1k} - \beta_{0k}\right). \] Since these quantities can be directly identified from data, we can express the treatment effects parameters $\lambda = c'\theta$ with $c = (c_\mu, c_\rho, -c_\mu, -c_\rho)'$ as the solution of the following moment function: \begin{equation} g_k(\lambda) = \left[p(z_k) - p(z_0)\right]\lambda - c_\mu\Delta_\mu(z_0,z_k) - c_\rho \Delta_\rho(z_0,z_k) = 0. \end{equation} If the instrument $Z$ is binary, then $c'\theta$ is just identified by this linear moment restriction, otherwise, we can construct a vector of moment functions: \begin{equation} g(\lambda) = (g_1(\lambda), \cdots, g_K(\lambda))' = 0_{K\times 1} \end{equation} to improve the efficiency for inference on $\lambda = c'\theta$. It is worth noting that other components of primitive parameter $\theta$ do not enter into (ref). As a result, their nonstandard asymptotic behavior (under weak identification) is irrelevant in this context. To this end, this linear moment function $g(\lambda)$ will be used to construct asymptotically similar tests for treatment effects in linear MTE models. To fix ideas, I first consider inference for the parameter $c'\theta$ with a known weight $c$ and exploit the binary variation $\{z_0, z_k\}$ for $k = 1,\ldots,K$. This known-weight setting allows for inference on parameters such as the ATE and the MTE function. Online Appendix (ref) discusses the extension to cases where $c$ needs to be estimated. Equation (ref) shows that $g_k(\lambda)$ suffices to estimate $\lambda$, the causal effects of interest. Let $\hat{g}_k(\lambda)$ be the sample analog of this moment function after plugging in the estimators $\{\hat{p}(z_\ell), \hat{\beta}_{0\ell}, \hat{\beta}_{1\ell}\}_{\ell = 0}^{K}$, i.e., \[ \hat{g}_k(\lambda) = \left[\hat{p}(z_k) - \hat{p}(z_0)\right]\lambda - c_\mu\hat{\Delta}_\mu(z_0,z_k) - c_\rho \hat{\Delta}_\rho(z_0,z_k) \] where \begin{align*} \hat{\Delta}_\mu(z_0,z_k) &= \hat{p}(z_k)[\hat{\beta}_{10} - \hat{\beta}_{00}] - \hat{p}(z_0)[\hat{\beta}_{1k} - \hat{\beta}_{0k}] + \hat{\beta}_{1k} - \hat{\beta}_{10} \\ \hat{\Delta}_\rho(z_0, z_k) &= 2(\hat{\beta}_{00} - \hat{\beta}_{10} + \hat{\beta}_{1k} - \hat{\beta}_{0k}). \end{align*} Based on this moment condition, one might consider an Anderson-Rubin (AR) test statistic to assess the null hypothesis $H_0: c'\theta = \lambda$ as follows \begin{equation} \text{AR}_{n,k}(\lambda) = \left|\frac{\sqrt{n}\hat{g}_k(\lambda)}{\hat{s}_{k}(\lambda)}\right|^2 \end{equation} where $\hat{s}_{k}^2(\lambda)$ consistently estimates the asymptotic variance of the sample moment $\hat{g}_k(\lambda)$. To construct this variance estimator, note that an asymptotic linear expansion of $\hat{g}_k(\lambda) - g_k(\lambda)$ gives \[ \hat{g}_k(\lambda) - g_k(\lambda) = \begin{pmatrix} \partial_{p'} {g}_k(\lambda)[\hat{p} - p] \\ + ~\partial_{\beta_1'} {g}_k(\lambda)[\hat{\beta}_1 - \beta_1] \\ + ~\partial_{\beta_0'} {g}_k(\lambda)[\hat{\beta}_0 - \beta_0] \end{pmatrix} + o_p(n^{-1/2}). \] with coefficients $\partial_{p} {g}_k(\lambda)$, $\partial_{\beta_1} {g}_k(\lambda)$, and $\partial_{\beta_0}{g}_k(\lambda)$ defined as the gradient of ${g}_k(\lambda)$ with respect to $p$, $\beta_1$, and $\beta_0$, respectively. It is natural to estimate these coefficients by plugging in sample analog estimators below: \begin{equation} \begin{aligned} & \partial_{p} \hat{g}_k(\lambda) = (-\lambda + (\hat{\beta}_{1k} - \hat{\beta}_{0k})c_\mu, 0, \ldots, 0, \lambda - (\hat{\beta}_{10} - \hat{\beta}_{00})c_\mu, 0, \ldots, 0)' \\ & \partial_{\beta_1} \hat{g}_k(\lambda) = ((1-\hat{p}(z_k))c_\mu + 2c_\rho, 0, \ldots, 0, -(1-\hat{p}(z_0))c_\mu - 2c_\rho, 0, \ldots, 0)' \\ & \partial_{\beta_0} \hat{g}_k(\lambda) = (\hat{p}(z_k)c_\mu - 2c_\rho, 0, \ldots, 0, -\hat{p}(z_0)c_\mu + 2c_\rho, 0, \ldots, 0)', \end{aligned} \end{equation} where nonzero elements appear on the first and the $(k+1)$'th elements of vectors. This leads to a consistent variance estimator \[ \hat{s}_{k}^2(\lambda) = \partial_{p'}\hat{g}_k(\lambda) \hat{\Sigma}_p \partial_{p}\hat{g}_k(\lambda) + \partial_{\beta_1'}\hat{g}_k(\lambda) \hat{\Sigma}_{\beta_1} \partial_{\beta_1}\hat{g}_k(\lambda) + \partial_{\beta_0'}\hat{g}_k(\lambda) \hat{\Sigma}_{\beta_0} \partial_{\beta_0}\hat{g}_k(\lambda). \] Let $\alpha \in (0,1)$ be the significant level. The proposition below establishes the uniform validity and asymptotic similarity of the AR test for $H_0:c'\theta=\lambda$ with $\lambda \in \ensuremath{\mathbb{R}}$. \begin{proposition} Let Assumption (ref) and (ref) hold, and suppose that the weight $c = (c_\mu, c_\rho, -c_\mu, -c_\rho)'$ is a nonzero fixed vector, then \[ \liminf_{n\to\infty}\inf_{(\lambda,F)\in \mathcal{P}_0} \ensuremath{\mathbb{P}}_F\left(\text{AR}_{n,k}(\lambda) > q_{\chi_1^2}(1-\alpha)\right) = \limsup_{n\to\infty}\sup_{(\lambda,F)\in \mathcal{P}_0} \ensuremath{\mathbb{P}}_F\left(\text{AR}_{n,k}(\lambda) > q_{\chi_1^2}(1-\alpha)\right) = \alpha, \] where $q_{\chi_1^2}(1-\alpha)$ denotes the $(1-\alpha)$-quantile of the $\chi_1^2$ distribution. \end{proposition} \begin{remark} The above result shows that the proposed AR test is valid and (uniformly) asymptotically similar in the sense of andrews/cheng/guggenberger:2020. Constructing an asymptotically valid subvector test is already challenging in models with weakly identified nuisance parameters\footnote{Since $c'\theta$ is the parameter of interest rather than $\theta$ itself, the vector of primitive parameters $\theta = (\mu_1, \rho_{11}, \mu_0, \rho_{01})$ becomes the weakly identified nuisance parameter under limited variation of propensity scores. Although we can show some functions of $\theta$ are strongly identified using the reparametrization technique proposed by han/mccloskey:2019, two parameters $(\rho_{11}, \rho_{01})$ remain weakly identified in the transformed model (see Online Appendix (ref)).}, because the asymptotic distributions of many test statistics depend on those unknown nuisance parameters that cannot be consistently estimated andrews/mikusheva:2016b. However, I introduce a novel moment function that isolates the causal effects of interest while being independent of $\theta$, thereby making the simple AR test feasible in this context. Unlike other existing approaches to subvector inference with weak identification (discussed in Online Appendix (ref)), this method achieves asymptotic similarity, ensuring the inverted confidence set maintains $1-\alpha$ coverage asymptotically under both strong and weak identification. \end{remark} If $Z$ is a binary instrument taking values in $\{z_0, z_1\}$, there exists a unique AR test statistic $\text{AR}_{n,1}(\lambda)$ for robust inference on the causal parameter $c'\theta$. When the instrument is discrete with multiple values $Z\in\{z_0, \ldots, z_K\}$ where $K > 1$, we obtain multiple AR statistics $\{\text{AR}_{n,k}(\lambda)\}_{k=1}^K$. The informativeness of each statistic depends on the variation in propensity scores $p(z_0) - p(z_k)$. In this case, combining tests based on different propensity score pairs yields greater statistical power. Let \[ \hat{\pi} = (\hat{p}(z_1) - \hat{p}(z_0), \ldots, \hat{p}(z_K) - \hat{p}(z_0))' \] and \[ \hat{\gamma} = (c_\mu \hat{\Delta}_\mu(z_0, z_1) + c_\rho \hat{\Delta}_\rho (z_0, z_1), \ldots, c_\mu \hat{\Delta}_\mu(z_0, z_K) + c_\rho \hat{\Delta}_\rho (z_0, z_K))'. \] Then our goal is to obtain a valid inference procedure based on a vector of linear moment functions: \[ \hat{g}(\lambda) = \hat{\pi}\lambda - \hat{\gamma} = (\hat{g}_1(\lambda), \ldots, \hat{g}_K(\lambda))' \in \ensuremath{\mathbb{R}}^K. \] Under strong identification where $\hat{\pi}$ converges to a nonzero limit, a Wald statistic based on efficient weighting matrix $\hat{S}(\lambda)^{-1}$ for testing $H_0: c'\theta = \lambda$ can be constructed below: \[ W_n(\lambda) = \frac{n\hat{g}(\lambda)'\hat{S}(\lambda)^{-1}\hat{\pi}\hat{\pi}'\hat{S}(\lambda)^{-1}\hat{g}(\lambda)}{\hat{\pi}'\hat{S}(\lambda)^{-1}\hat{\pi}}, \] where \[ \hat{S}(\lambda) = \partial_{p}\hat{g}(\lambda) \hat{\Sigma}_p \partial_{p'}\hat{g}(\lambda) + \partial_{\beta_1}\hat{g}(\lambda) \hat{\Sigma}_{\beta_1}\partial_{\beta_1'}\hat{g}(\lambda) + \partial_{\beta_0}\hat{g}(\lambda) \hat{\Sigma}_{\beta_0}\partial_{\beta_0'}\hat{g}(\lambda) \] is a consistent estimator of the asymptotic covariance matrix of $\hat{\pi}\lambda - \hat{\gamma}$ under null hypothesis. Under weak identification, the propensity score differences $\hat{\pi}$ may converge in probability to a zero vector, resulting in a nonstandard asymptotic distribution for $W_n(\lambda)$. To illusrate this problem, consider a sequence of DGPs where $\sqrt{n}\pi$ converges to a constant vector $\pi_0$. In this case, $\sqrt{n}\hat{\pi}$ converges in distribution to a multivariate normal distribution with mean $\pi_0$, which cannot be consistently estimated from the data. Consequently, the asymptotic distribution of $W_n(\lambda)$ contains the non-estimable nuisance parameter $\pi_0$, making it infeasible to consistently estimate the unconditional quantiles of $W_n(\lambda)$'s limiting distribution. However, similar to the arguments of moreira:2003 for linear IV models and andrews/mikusheva:2016b for quasi-likelihood ratio tests in GMM models, the asymptotic distribution of $W_n(\lambda)$ becomes independent of $\pi_0$ when conditioned on a sufficient statistic, making it feasible to approximate the conditional distribution of $W_n(\lambda)$. Define \[ \hat{h}(\lambda) = \sqrt{n}\hat{\pi} - [\partial_{p} {\pi}]\hat{\Sigma}_p [\partial_{p'} \hat{g}(\lambda)] \hat{S}(\lambda)^{-1} \sqrt{n}\hat{g}(\lambda). \] The statistic $\hat{h}(\lambda)$ is asymptotically independent of $\sqrt{n}\hat{g}(\lambda)$ and contains sufficient information about $\pi_0$ such that the (asymptotic) distribution of $W_n(\lambda)$ conditional on $\hat{h}(\lambda)$ is free of $\pi_0$. To simulate this conditional distribution, I construct a simulated counterpart of $\sqrt{n}\hat{\pi}$, denoted as ${\pi}_s$, which is given by \[ {\pi}_s = \hat{h}(\lambda) + [\partial_{p} {\pi}]\hat{\Sigma}_p [\partial_{p'} \hat{g}(\lambda)] \hat{S}(\lambda)^{-1/2}\eta^* \] where $\eta^* \sim \ensuremath{\mathcal{N}}(0_{K\times 1},I_{K\times K})$ is drawn independently of the data. We then approximate the distribution of $W_n(\lambda)$ by replacing $\sqrt{n}\hat{\pi}$ with ${\pi}_s$ and $\sqrt{n}\hat{g}(\lambda)$ with $\hat{S}(\lambda)^{1/2}\eta^*$, which yields \[ W_n^*(\lambda) = \frac{(\eta^*)'\hat{S}(\lambda)^{-1/2}\pi_s\pi_s'\hat{S}(\lambda)^{-1/2}(\eta^*)}{\pi_s'\hat{S}(\lambda)^{-1}\pi_s}. \] It follows that $W_n^*(\lambda)$ has the same asymptotic distribution as $W_n(\lambda)$ conditional on the realization of $\hat{h}(\lambda)$. Let $\hat{q}_{W^*}(1-\alpha)$ denote the $(1-\alpha)$-quantile of the distribution of $W_n^*(\lambda)$ conditional on the data. This critical value can be used to construct valid conditional Wald test: \begin{theorem} Let Assumption (ref) and (ref) hold, and suppose that the weight $c = (c_\mu, c_\rho, -c_\mu, -c_\rho)'$ is a nonzero fixed vector, then \[ \liminf_{n\to\infty}\inf_{(\lambda,F)\in\mathcal{P}_0} \ensuremath{\mathbb{P}}_{F}\left(W_n(\lambda) > \hat{q}_{W^*}(1-\alpha)\right) = \limsup_{n\to\infty}\sup_{(\lambda,F)\in\mathcal{P}_0} \ensuremath{\mathbb{P}}_{F}\left(W_n(\lambda) > \hat{q}_{W^*}(1-\alpha)\right) = \alpha. \] \end{theorem} \begin{remark} The statistic $W_n(\lambda)$ can be interpreted as a linear combination of AR statistics in equation (ref), with weight $\hat{S}^{-1}\hat{\pi}$. Under strong identification, this weighting vector can be shown to be optimal among all linear combinations of AR statistics, yielding the highest local asymptotic power. \end{remark} \begin{remark} In Lemma (ref), I show that the asymptotic variance of the moment function is uniformly positive definite as long as the weight $c = (c_\mu, c_\rho, -c_\mu, -c_\rho)$ is nonzero. If we incorporate additional moment conditions induced by the difference of propensity scores of the form $p(z_k) - p(z_j)$ for $k, j \neq 0$, then the asymptotic variance may become singular for some nonzero weight $c$. For example, if we set $c_\mu = 0$ and $c_\rho = 1$, it follows that \[ [\hat{p}(z_k) - \hat{p}(z_j)]\lambda - c_\rho\hat{\Delta}_\rho(z_j,z_k) = \hat{g}_k(\lambda) - \hat{g}_j(\lambda), \] implying that the moment condition constructed with the variation between $z_k$ and $z_j$ (on the left-hand side) is the difference of the moments constructed by using $(z_k, z_0)$ and $(z_j, z_0)$ (on the right-hand side). \end{remark} \begin{remark} An alternative statistic one could consider is the quasi-likelihood ratio statistic \[ \text{QLR}_n(\lambda) = n\left[\hat{g}(\lambda)'\hat{S}(\lambda)^{-1}\hat{g}(\lambda) - \inf_{\lambda \in \ensuremath{\mathbb{R}}} \hat{g}(\lambda)'\hat{S}(\lambda)^{-1}\hat{g}(\lambda)\right] \] and its corresponding conditional approach discussed in andrews/mikusheva:2016b. Here I focus on the conditional Wald approach because it is simpler to implement in practice and is first-order asymptotic equivalent to QLR test under strong identification. \end{remark} \section{Robust Inference in General MTE Models} In this section, I relax the linearity Assumption (ref) and extend the analysis to allow for any functional forms specified in Assumption (ref).5. This includes the linear MTE model and the polynomial specification $h_m(\cdot) = u^{m} - \frac{1}{m+1}$ that is commonly used in empirical studies. For this broader class of models, constructing moment conditions that only involve the target causal parameter is less obvious. Therefore, I develop an alternative \textit{improved projection} approach that builds on I. andrews.i:2018 to achieve valid inference. When compared to the conditional Wald test in section (ref), this new approach applies to a wider range of MTE models but is more computationally intensive and is potentially more conservative. On the theoretical front, I argue that the high-level conditions imposed by I. andrews.i:2018 cannot be directly verified in the MTE setting. Therefore, I extend his sequential validity result (that based on high-level conditions) by establishing the uniform validity under primitive conditions on the parameter space outlined in Definition (ref). \subsection{Improved projection inference} First, I describe how to adapt the improved projection test developed by I. andrews.i:2018 for conducting inference on $c'\theta$ using the following linear system of equations: \begin{equation} A\theta = \beta \end{equation} where \begin{align*} A = \begin{pmatrix} A_1 & 0_{(K+1)\times (M+1)} \\ 0_{(K+1)\times (M+1)} & A_0 \end{pmatrix} \qquad \theta = \begin{pmatrix} \theta_1 \\ \theta_0 \end{pmatrix} \qquad \beta = \begin{pmatrix} \beta_1 \\ \beta_0 \end{pmatrix}, \end{align*} as defined below equation (ref). Note that $A$ is a matrix of transformed propensity scores and $\beta$ is a vector of conditional expectations, both of which can be consistently estimated by their sample analogs $\hat{A}$ and $\hat{\beta}$ under appropriate assumptions. Since the matrix $A$ captures the variation in propensity scores, it determines the strength of identification. The conventional Wald test for conducting inference on $c'\theta$ uses the first-order asymptotic efficient estimator obtained by minimizing the following minimum-distance objective function: \[ \hat{\theta}^{\text{eff}} \in \operatorname*{argmin}_{\theta\in \Theta} n(\hat{A}\theta - \hat{\beta})'{\hat{\Omega}(\theta)^{-1}}(\hat{A}\theta - \hat{\beta}) \] where $\hat{\Omega}(\theta)$ is a consistent estimator for the asymptotic variance of $\sqrt{n}(\hat{A}\theta - \hat{\beta})$, formally defined below in equation (ref). Then the classical Wald statistic for assessing $H_0: c'\theta = \lambda$ is \begin{equation} \text{Wald}_n(\lambda) = n(c'\hat{\theta}^{\text{eff}} - \lambda)' (c'(\hat{A}'\hat{\Omega}(\hat{\theta}^{\text{eff}})^{-1}\hat{A})^{-1}c)^{-1}(c'\hat{\theta}^{\text{eff}} - \lambda). \end{equation} When Assumption (ref).7 is nearly violated, $\hat{A}$ may converge in probability to a matrix with deficient rank. For example, consider the following simple example on the linear MTE model with a binary IV. \begin{example}[Linear MTE model with a binary IV] Suppose $K = M = 1$ and set $h_1(u) = u-\frac{1}{2}$. Along a sequence of DGPs $\{F_n\}$, let $p_{F_n}(z_0) = p \in (0,1)$ and $p_{F_n}(z_1) = p + \upsilon_{n} \in (0,1)$ with $\upsilon_n \to 0$. In such case, the probability limit of $\hat{p}(z_0)$ and $\hat{p}(z_1)$ are equal to $p$, and we have \[ \hat{A} = \begin{pmatrix} \begin{matrix} 1 & \frac{1}{2} (\hat{p}(z_0) - 1) \\ 1 & \frac{1}{2} (\hat{p}(z_1) - 1) \end{matrix} & 0_{2\times 2} \\ 0_{2\times 2} & \begin{matrix} 1 & \frac{1}{2}\hat{p}(z_0) \\ 1 & \frac{1}{2}\hat{p}(z_1) \end{matrix} \end{pmatrix} ~~ \xrightarrow{p} ~~ \begin{pmatrix} \begin{matrix} 1 & \frac{1}{2} (p - 1) \\ 1 & \frac{1}{2} (p - 1) \end{matrix} & 0_{2\times 2} \\ 0_{2\times 2} & \begin{matrix} 1 & \frac{1}{2}p \\ 1 & \frac{1}{2}p \end{matrix} \end{pmatrix} \] Note that the probability limit of $\hat{A}$ becomes a singular matrix. \end{example} This singularity may result in multiple minimizers of the limiting objective function when defining $\hat{\theta}^{\text{eff}}$, suggesting that the efficient estimator $\hat{\theta}^{\text{eff}}$ may not exhibit the usual properties of consistency and asymptotic normality. To address this issue, one can derive inference results based on the moment condition: \[ {m}(\theta) \equiv {A}\theta - {\beta} = 0_{2(K+1)\times 1} \] instead of using the estimator $\hat{\theta}^{\text{eff}}$. Let $\hat{m}(\theta) \equiv \hat{A}\theta - \hat{\beta}$ denote the sample moment function. By substituting the true (or hypothesized) value $\theta$ for $\hat{\theta}^{\text{eff}}$ in $\hat{\Omega}(\hat{\theta}^{\text{eff}})$ and plugging the following first-order asymptotic expansion: \[ \sqrt{n}(\hat{\theta}^{\text{eff}} - \theta) = -(\hat{A}'\hat{\Omega}(\theta)^{-1}\hat{A})^{-1}\hat{A}'\hat{\Omega}(\theta)^{-1}\sqrt{n}(\hat{A}\theta - \hat{\beta}) + o_p(1) \] into equation (ref), this gives a locally equivalent Lagrangian Multiplier (LM) test statistic: \[ \text{LM}_n(\theta) \equiv n(\hat{A}\theta - \hat{\beta})'\hat{\Omega}(\theta)^{-1/2} P_{\hat{\Omega}(\theta)^{-1/2}\hat{A}(\hat{A}'\hat{\Omega}(\theta)^{-1}\hat{A})^{-1}c}\hat{\Omega}(\theta)^{-1/2}(\hat{A}\theta - \hat{\beta}) \] which does not require $\hat{\theta}^{\text{eff}}$ to be consistent. However, the singularity of $\hat{A}'\hat{\Omega}(\theta)^{-1}\hat{A}$ persists under the projection operator inside the LM statistic. Following the insights of kleibergen:2005, one can orthogonalize columns of $\hat{A} = (\hat{a}_1,\ldots, \hat{a}_{2(M+1)})$ with respect to the variation in the moment condition $\hat{m}(\theta)$ to obtain a new gradient estimator \[ \hat{D}(\theta) = (\hat{d}_1(\theta), \hat{d}_2(\theta), \ldots, \hat{d}_{2(M+1)}(\theta)) \] where \[ \hat{d}_j(\theta) \equiv \hat{a}_j - \hat{\Gamma}_j(\theta)\hat{\Omega}(\theta)^{-1}(\hat{A}\theta - \hat{\beta}) \quad \text{for } j = 1,2,\ldots, 2(M+1). \] Here $\hat{\Gamma}_j(\theta)$ is a consistent estimator of the asymptotic covariance between $\sqrt{n}\hat{m}(\theta)$ and the $j$-th column in $\sqrt{n}(\hat{A} - A)$, formally defined below in equation (ref). By this orthogonalization, $\sqrt{n}(\hat{D}(\theta) - A)$ is asymptotically independent of the moment condition $\sqrt{n}\hat{m}(\theta)$ in large samples. Replacing $\hat{A}$ with $\hat{D}(\theta)$ in the LM statistic then leads to the “subvector” version of Robust LM (RLM) statistics for inference on $c'\theta$: \begin{equation} \text{RLM}_n(\theta) = n(\hat{A}\theta - \hat{\beta})'\hat{\Omega}(\theta)^{-1/2} P_{\hat{\Omega}(\theta)^{-1/2}\hat{D}(\theta)(\hat{D}(\theta)'\hat{\Omega}(\theta)^{-1}\hat{D}(\theta))^{-1}c}\hat{\Omega}(\theta)^{-1/2}(\hat{A}\theta - \hat{\beta}). \end{equation} If $\sqrt{n}A$ converges to a fixed matrix $\mathcal{A}$ as in kleibergen:2005, so that $A$ has a zero limit, then $\sqrt{n}\hat{D}(\theta)$ converges to a Gaussian matrix with mean $\mathcal{A}$ that is asymptotically independent of the moment vector. Consequently, the projection matrix in the RLM statistic becomes asymptotically independent of the moments on both sides, implying that $\text{RLM}_n(\theta)$ converges to the standard $\chi_1^2$ limiting distribution for a sequence of DGPs that induces a zero limit of $\hat{A}$ at a rate $n^{-1/2}$. Compared to the classical RLM statistic for full vector inference in kleibergen:2005, this subvector RLM statistic (ref) only attains power for deviations in the linear function $c'\theta$ rather than the full vector $\theta$. This feature has two important implications. On the one hand, when identification is sufficiently strong, the subvector RLM statistic is locally equivalent to the efficient subvector Wald statistic (ref) when $\theta$ approaches its true value; On the other hand, this equivalence may fail since the subvector RLM statistic cannot distinguish alternative values of $\theta$ from the true value when they yield the same value of $c'\theta$. To overcome this limitation, I. andrews.i:2018 introduces a linear combination (LC) statistic that combines the RLM and AR statistics: \begin{equation} \text{LC}_n(\theta) = \text{RLM}_n(\theta) + a \cdot \text{AR}_n(\theta) \end{equation} where \[ \text{AR}_n(\theta) = n(\hat{A}\theta - \hat{\beta})'\hat{\Omega}(\theta)^{-1}(\hat{A}\theta - \hat{\beta}) \] and $a > 0$ is a tuning parameter on the weights attached to the AR statistic. By incorporating the AR term, the LC statistic diverges to infinity outside the $n^{-1/2}$ neighborhood of the true parameter $\theta$ under strong identification. This property guarantees power against any deviations from the true value. Moreover, when $a$ is sufficiently small and the model is strongly identified, the projection test based on the LC statistic becomes approximately equivalent to the subvector Wald test. \subsection{Contributions and Modifications} In a GMM model, I. andrews.i:2018 shows that $\text{LC}_n(\theta)$ converges to a mixture of two independent chi-squared distributions $(1+a)\chi^2_{1} + a\chi^2_{2K+1}$ in a sequence of DGPs $\{F_n\}_{n\geq 1}$ satisfying certain high-level conditions. As Andrews notes in his conclusion, the uniform validity of his result under more primitive conditions remains an open question. In this section, I make two key contributions: I show that his high-level conditions are not trivially satisfied in the MTE framework, and I establish the uniform validity of the LC test through a simple modification. First, consider Assumption 4 of I. andrews.i:2018. This assumption requires two convergence conditions by the existence of normalizing sequences: a sequence of full-rank matrices $\{\Lambda_{1,n}\}\subseteq \ensuremath{\mathbb{R}}^{2(M+1)\times 2(M+1)}$ and a sequence of nonzero constants $\{\Lambda_{2,n}\} \subseteq \ensuremath{\mathbb{R}}$. Under these sequences, the normalized matrix $\hat{D}(\theta)\Lambda_{1,n}$ must converge in distribution to a Gaussian matrix of full rank almost surely, and the normalized weight $\Lambda_{1,n}'c\Lambda_{2,n}$ must converge to a nonzero vector. The convergence part of this assumption can be verified straightforwardly in the context of kleibergen:2005 as discussed above, who considers the sequence \begin{equation} \sqrt{n}A_{F_n} \to \mathcal{A} \in \ensuremath{\mathbb{R}}^{2(K+1)\times 2(M+1)} \end{equation} in which case $\Lambda_{1,n} = \sqrt{n}I$ and $\Lambda_{2,n} = \frac{1}{\sqrt{n}}$. More generally, stock/wright:2000 and chaudhuri/zivot:2011 consider a sequence \begin{equation} A_{F_n} = (0_{2(K+1)\times q}, A_{\text{full}}) + A_{\text{sing},F_n} \quad \text{and} \quad \sqrt{n}A_{\text{sing},F_n} \to \mathcal{A} \end{equation} in which case $A_{\text{full}}$ has full column rank, representing those parameters that are strongly identified. We can set $\Lambda_{1,n} = \operatorname*{diag}\{\sqrt{n}I_{q},I_{2(M+1) - q}\}$ and $\Lambda_{2,n} = \frac{1}{\sqrt{n}}$ such that Assumption 4 still holds. However, the sequences (ref) and (ref) are inappropriate for the MTE setup. To see this, consider again the example on a linear MTE model with a binary IV as discussed above: \begin{example}[Linear MTE model with a binary IV] Suppose $K=M=1$ and set $h_{1}(u) = u - \frac{1}{2}$. Along a sequence of DGPs $\{F_n\}$, let $p_{F_n}(z_0) = p_{F_n}(z_1) = p \in (0,1)$ for each $n\geq 1$. Since $\overline{K} = |\{p_{F_n}(z_0), p_{F_n}(z_1)\}| = 1 < M+1 = 2$, Assumption (ref).7 fails in this example. Note that the matrix $A_{F_n}$ becomes \[ A_{F_n} = \begin{pmatrix} \begin{matrix} 1 & \frac{1}{2} (p - 1) \\ 1 & \frac{1}{2} (p - 1) \end{matrix} & 0_{2\times 2} \\ 0_{2\times 2} & \begin{matrix} 1 & \frac{1}{2}p \\ 1 & \frac{1}{2}p \end{matrix} \end{pmatrix} \] In this case, neither (ref) nor (ref) holds here. \end{example} More generally, a singular but nonzero limit of $A_{F_n}$ would not satisfy conditions (ref) or (ref). Similar examples of weakly identified models that fail to meet these conditions are discussed in andrews/guggenberger:2017, where a singular value decomposition (SVD) technique is introduced to establish uniform validity of the classical RLM test for the \textit{full} vector. To achieve the same goal of uniform validity, I generalize their SVD approach to address inference on parameter \textit{functionals} in the MTE framework. My approach proceeds with a SVD of $A_{F_n}$: \[ A_{F_n} = C_{F_n} \underbrace{ \begin{bmatrix} \begin{pmatrix} \tau_{F_n,1} & 0 & \ldots & 0 \\ 0 & \tau_{F_n,2} & \ldots & 0 \\ \vdots & \vdots & \ddots & \vdots \\ 0 & 0 & \ldots & \tau_{F_n, 2(M+1)} \end{pmatrix} \\ 0_{2(K-M)\times 2(M+1)} \end{bmatrix}}_{\Pi_{F_n}} B_{F_n}', \] where $C_{F_n} \in \ensuremath{\mathbb{R}}^{2(K+1)\times 2(K+1)}$ and $B_{F_n} \in \ensuremath{\mathbb{R}}^{2(M+1)\times 2(M+1)}$ are orthogonal matrices, and $\infty \geq \tau_{F_n,1} \geq \tau_{F_n,2} \geq \ldots \geq \tau_{F_n,2(M+1)} \geq 0$ are singular values of $A_{F_n}$ in a descending order. I show that the following normalizing matrices establish convergence of $\hat{D}(\theta)\Lambda_{1,n}$ and $\Lambda_{1,n}'c\Lambda_{2,n}$: \begin{align*} &\Lambda_{1,n} = B_{F_n} \operatorname*{diag}\{(\tau_{F_n,1})^{-1},\ldots, (\tau_{F_n,q})^{-1}, \sqrt{n},\ldots, \sqrt{n}\} \\ &\Lambda_{2,n} = \|\Lambda_{1,n}'c\|^{-1}, \end{align*} where $q$ is the number of scaled singular values $\{\sqrt{n}\tau_{F_n,j}\}_{j=1}^{2(M+1)}$ that diverge to infinity. Despite the convergence of $\hat{D}\Lambda_{1,n}$ under the constructed normalizing matrices, its asymptotic limit may be rank deficient. This issue prevents the use of the efficiently weighted projection matrix in the subvector RLM statistic. A similar issue was also noticed by andrews/guggenberger:2017, who impose additional parameter space restrictions to avoid this problem. To address this problem, I modify the matrix $\hat{D}$ by adding a small noise of size $n^{-1/2}$: \begin{equation} \widetilde{D}(\theta) = \hat{D}(\theta) + \kappa n^{-1/2}\xi, \end{equation} where $\kappa > 0$ is a tuning parameter and $\xi \in \ensuremath{\mathbb{R}}^{2(K+1)\times 2(M+1)}$ is a matrix of i.i.d. standard normal random variables independent of data. This perturbation ensures that $\widetilde{D}(\theta)\Lambda_{1,n}$ achieves full rank almost surely while having asymptotically negligible effects on the test statistic under strong identification. This perturbation technique draws inspiration from the AR/LM test\footnote{In addition to the modification in (ref), D. andrews.d:2017 introduces two tuning parameters, $(K_L^*, K_U^*)$, to categorize identification strength. However, this categorization is not required for our inference procedure, and there is no clear guidance for choosing these parameters.}andrews.d:2017. While AR/LM test focuses specifically on inference for a subset of parameters, my paper addresses a different problem of inference on a linear functional of parameters. Define the new matrix under projection as follows: \[ \hat{Q}(\theta) = \hat{\Omega}(\theta)^{-1/2}\widetilde{D}(\theta)(\widetilde{D}(\theta)'\hat{\Omega}(\theta)^{-1}\widetilde{D}(\theta))^{-1}c, \] and the corresponding modified RLM (MRLM) statistic becomes \[ \text{MRLM}_n(\theta) = n(\hat{A}\theta - \hat{\beta})' \hat{\Omega}(\theta)^{-1/2} P_{\hat{Q}(\theta)} \hat{\Omega}(\theta)^{-1/2}(\hat{A}\theta - \hat{\beta}). \] Then the modified LC (MLC) statistic is given by \begin{align} \text{MLC}_n(\theta) = \text{MRLM}_n(\theta) + a \cdot \text{AR}_n(\theta). \end{align} In section (ref), I establish the uniform validity of using (ref) for conducting inference on the linear function $c'\theta$ even if Assumption (ref).7 might fail or be close to failing. \subsection{Implementation of the MLC test} In this section, I describe the implementation of the projection test based on the MLC statistic as below: \begin{enumerate} • Construct the key quantities for the test statistics \begin{enumerate} • Construct the estimators of $\hat{p}$ and $\hat{\beta}$, as well as their asymptotic variances estimators $\hat{\Sigma}_p$ and $\hat{\Sigma}_\beta$, as in section (ref). Plugging in the estimator $\hat{p}$ into the matrix $A$ then obtains the estimator $\hat{A}$. • Define the asymptotic variance estimator of the moment function \begin{equation} \hat{\Omega}(\theta) = H(\hat{p}, \theta)\hat{\Sigma}_p H(\hat{p}, \theta)' + \hat{\Sigma}_\beta \end{equation} where \[ H(p,\theta) = \begin{bmatrix} \operatorname*{diag}\left\{\sum_{m=0}^M \theta_{1m} \lambda_{1m}'({p}(z_\ell)): \ell = 0,1,\ldots, K\right\} \\ \operatorname*{diag}\left\{\sum_{m=0}^M \theta_{0m} \lambda_{0m}'({p}(z_\ell)): \ell = 0,1,\ldots, K\right\} \end{bmatrix}. \] • For each $m=0,1,\ldots,M$, let \[ L_m(\hat{p}) = \operatorname*{diag}\{\lambda_{1m}'(\hat{p}(z_0)), \ldots, \lambda_{1m}'(\hat{p}(z_K))\} \] and \[ R_m(\hat{p}) = \operatorname*{diag}\{\lambda_{0m}'(\hat{p}(z_0)), \ldots, \lambda_{0m}'(\hat{p}(z_K))\}. \] For each $j=1,\ldots,2(M+1)$, define a column vector $\hat{d}_j(\theta)$ and the covariance estimator $\hat{\Gamma}_j(\theta)$ as below \begin{align} & \hat{d}_j(\theta) \equiv \hat{a}_j - \hat{\Gamma}_j(\theta)\hat{\Omega}(\theta)^{-1}(\hat{A}\theta - \hat{\beta}) \notag \\ & \hat{\Gamma}_{j}(\theta) = M_{j} (\hat{p}) \hat{\Sigma}_p H(\hat{p},\theta)' \end{align} where \begin{equation} M_j(\hat{p}) = \begin{cases} \begin{bmatrix} L_{j-1}(\hat{p}) \\ 0_{(K+1)\times (K+1)} \end{bmatrix} & \text{if } j \leq M+1 \\ \begin{bmatrix} 0_{(K+1)\times (K+1)} \\ R_{j-M-2}(\hat{p}) \end{bmatrix} & \text{if } j > M+1 \\ \end{cases}. \end{equation} Let $\hat{D}(\theta) = (\hat{d}_1(\theta),\ldots,\hat{d}_{2(M+1)}(\theta))$ and $\widetilde{D}(\theta) = \hat{D}(\theta) + \kappa n^{-1/2}\xi$, where $\xi\in \ensuremath{\mathbb{R}}^{2(K+1)\times 2(M+1)}$ is a matrix of standard normal random variables independent of data. $\kappa$ is a positive tuning parameter chosen as $10^{-6}$ following D. andrews.d:2017. \end{enumerate} • Given the estimators $\hat{A}$, $\hat{\beta}$, $\hat{\Omega}(\theta)$, and $\widetilde{D}(\theta)$ constructed from step 1, compute the modified LC statistics as in equation (ref) for $\theta \in \Theta$. For the empirical application and simulation studies, I set the AR weight coefficient as 0.05, which is close to the suggested weight $a(\gamma)$ introduced by I. andrews.i:2018 by setting $\alpha = 5\%$, $\gamma = 10\%$, and $K = 15$, consistent with the application considered in section (ref). • Let $\alpha \in (0,1)$ be the significance level. The modified LC test rejects the null hypothesis $H_0: c'\theta = \lambda$ if the profiled test statistic over the linear manifold $c'\theta = \lambda$ is larger than their respective critical values, i.e., \begin{align*} \hat{\phi}_{\text{MLC}}(\lambda) = \ensuremath{\mathbbm{1}}\left[\inf_{c'\theta = \lambda} \text{MLC}_n(\theta) > q_{(1+a)\chi_1^2 + a\chi_{2K+1}^2}(1-\alpha)\right] \end{align*} where $q_{(1+a)\chi_1^2 + a\chi_{2K+1}^2}(1-\alpha)$ denotes the $(1-\alpha)$-quantile of the mixture chi-square distribution $(1+a)\chi_1^2 + a\chi_{2K+1}^2$ with $\chi_1^2 \mathbin{ \mathpalette{\@indep}{} } \chi_{2K+1}^2$. Additionally, define $\hat{\phi}_{\text{MLC}}(\lambda) = 1$ for $\lambda \not\in \{c'\theta: \theta \in \Theta\}$. \end{enumerate} A robust confidence set can be obtained by inverting the test, that is, \[ \mathcal{C}_{\text{MLC}} = \{\lambda\in\ensuremath{\mathbb{R}}: \hat{\phi}_{\text{MLC}}(\lambda) = 0\}. \] In practice, I report the convex hull of the confidence set so that it becomes an interval. \subsection{Uniform validity} The following theorem establishes the uniform validity of the MLC test based on the profiling the MLC statistic. \begin{theorem} Let Assumption (ref) hold, and suppose that the weight $c$ is a nonzero fixed vector. Then we have \[ \limsup_{n\to\infty}\sup_{(\lambda,F)\in\mathcal{P}_0} \ensuremath{\mathbb{E}}_F[\hat{\phi}_{\text{MLC}}(\lambda)] \leq \alpha. \] \end{theorem} By this theorem, the MLC test has an asymptotic size less than or equal to the nominal level $\alpha \in (0,1)$ for the parameter space $\mathcal{P}_0$. In other words, the MLC test is uniformly valid regardless of the model's identification status. In Online Appendix (ref), I further demonstrate that the MLC test is consistent for distant (or fixed) alternatives and has comparable power to the asymptotic efficient Wald test under strong identification when the AR weight $a$ is sufficiently small. These findings highlight the usefulness of the MLC test for researchers seeking robust inference against weak identification. \section{Incorporating Covariates} In this section, I consider a set of covariates $W \in \ensuremath{\mathbb{R}}^L$ included in the MTE model. I first establish an augmented set of assumptions under which the MTE function can be point identified conditional on $W=w$ for each $w \in \operatorname*{supp}(W)$. Then I argue that a commonly used additive separability condition can introduce large bias in the estimand of target causal effects when unobserved heterogeneity interacts with the covariates. To this end, I recommend conducting the robust inference procedure by conditioning on covariates rather than imposing additive separability. \subsection{Identifying assumptions with covariates} Consider the following MTE model with an additional set of covariates $W$: \begin{align*} & Y_d = \mu_d(W) + V_d \\ & D = \ensuremath{\mathbbm{1}}[U \leq \nu(W,Z)], \end{align*} where $\mu_d(W) \equiv \ensuremath{\mathbb{E}}[Y_d\mid W]$ denotes the conditional expectation of potential outcomes. Similar to Theorem (ref), we can identify the MTE function $\ensuremath{\mathbb{E}}[Y_1-Y_0\mid F_{U|W}(U|W)=u, W=w]$ under the following assumptions. \begin{assumption}[MTE model with covariates] \begin{enumerate} • $Z \mathbin{ \mathpalette{\@indep}{} } U\mid W$. • $\ensuremath{\mathbb{E}}[Y_d\mid Z, W, U] = \ensuremath{\mathbb{E}}[Y_d\mid W, U]$ and $\ensuremath{\mathbb{E}}|Y_d| < \infty$ for $d \in \{0,1\}$. • $U\mid W = w$ is continuously distributed for all $w \in \operatorname*{supp}(W)$. • $0 < \ensuremath{\mathbb{P}}(D=1\mid W=w, Z=z) <1$ for all $(w,z) \in \operatorname*{supp}(W, Z)$. • $\ensuremath{\mathbb{E}}[Y_d\mid F_{U\mid W}(U|W) = u, W=w] = \mu_d(w) + \sum_{m=1}^M \rho_{dm}(w) h_m(u)$ for some known continuous functions $\{h_m(\cdot)\}_{m=1}^{M}$ for each $d = 0,1$ and $u\in(0,1)$, where $F_{U\mid W}(\cdot|w)$ denotes the distribution function of $U$ conditional on $W=w$. • $\{\lambda_{1m}(\cdot)\}_{m=0}^{M}$ and $\{\lambda_{0m}(\cdot)\}_{m=0}^M$ are unisolvent on $(0,1)$, where $\lambda_{00}(\cdot) = \lambda_{10}(\cdot) \equiv 1$, and \[ \lambda_{1m}(p) \equiv \frac{1}{p}\int_{0}^p h_m(u) du \quad \text{and} \quad \lambda_{0m}(p) \equiv \frac{1}{1-p}\int_{p}^1 h_m(u) du \quad \text{for $m = 1,\ldots,M$}. \]$|\{\ensuremath{\mathbb{P}}(D=1\mid Z=z, W=w): z \in \operatorname*{supp}(Z\mid W=w)\}|\geq M+1$ for all $w \in \operatorname*{supp}(W)$. \end{enumerate} \end{assumption} Under the above assumption, we can normalize the distribution of $U$ conditional on the random covariate $W$ to be uniformly distributed over the unit interval (i.e., $F_{U\mid W}(u|W) = u$) and thus $\nu(w,z) = p(w,z)\equiv \ensuremath{\mathbb{P}}(D=1\mid W=w, Z=z)$. The structural parameter \[ \theta(w) = (\mu_1(w), \{\rho_{1m}(w)\}_{m=1}^M, \mu_0(w),\{\rho_{0m}(w)\}_{m=1}^M) \] can be identified for all $w\in\operatorname*{supp}(W)$ by the following separate regressions: \begin{equation} \begin{aligned} \ensuremath{\mathbb{E}}[Y\mid D=1, Z=z, W=w] & = \mu_1(w) + \sum_{m=1}^M \rho_{1m}(w) {\int_0^{p(w,z)} \frac{h_m(u)}{p(w,z)} du} \\ & = \mu_1(w) + \sum_{m=1}^M \rho_{1m}(w) \lambda_{1m}(p(w,z))\\ \ensuremath{\mathbb{E}}[Y\mid D=0, Z=z, W=w] &= \mu_0(w) + \sum_{m=1}^M \rho_{0m}(w) {\int_{p(w,z)}^1 \frac{h_m(u)}{1-p(w,z)} du} \\ & = \mu_0(w) + \sum_{m=1}^M \rho_{0m}(w) \lambda_{0m}(p(w,z)). \end{aligned} \end{equation} For some $w\in \operatorname*{supp}(W)$, the variation of $\{p(w,z): z\in\operatorname*{supp}(Z\mid W=w)\}$ can be weak in practice, especially when conditioning on certain groups of individuals who have very high (or low) probability of being treated. To address this problem, researchers commonly impose the additive separability condition as follows. \begin{assumption}[Additive separability] The MTR function is linear and additively separable in covariates and selection unobservable. That is, for $d = 0,1,$ \[ \ensuremath{\mathbb{E}}[Y_d\mid W, U] = \mu_d + W'\tau_d + \ensuremath{\mathbb{E}}[V_d\mid U]. \] \end{assumption} This assumption eliminates heterogeneous effects of covariates $W$ interacted with selection unobservable $U$ on potential outcomes, i.e., $\rho_{dm}(w)$ does not vary with $w$. Up to a linear term $W'\tau_d$, the covariate $W$ is mean independent of $Y_d$ conditional on $U$. This enables point identification of MTEs on the unconditional support of propensity scores $p(W,Z)$ by using variation from covariates in addition to exogenous variation from discrete instruments carneiro/heckman/vytlacil:2011. Next, I show that the failure of Assumption (ref) can lead to a biased estimator of treatment effects unless $W$ is uncorrelated with the treatment and propensity score. Given this result, I consider Assumption (ref) to be a strong assumption and implement inference in the empirical analysis by conditioning on covariates rather than relying on additive separability. \subsection{Bias from additive separability} I show the population bias of treatment effect estimands for a specific class of DGP as follows: \begin{equation} \ensuremath{\mathbb{E}}[Y_d \mid W = w, U = u] = \mu_d + w'\tau_d + \rho_d(w) h(u), \end{equation} where $\rho_d(w) = \rho_d + w'\eta_d$ for $d = 0,1$, and $h(\cdot)$ is strictly increasing and integrates to zero over $[0,1]$. This model can be viewed as an MTE model satisfying Assumption (ref) with $M=1$ and with $\mu_d(w)$ linear in $w$. A related specification of this form has appeared in kline/walters:2019 under additive separability, where $\rho_d(w)$ is treated as constant in $w$. So the model (ref) can be considered as a direct generalization of their setting by allowing the effect of $W$ to vary with the selection unobservable $U$, therefore potentially violating the additive separability Assumption (ref). Moreover, the bias derived under the specific model class (ref) indicates that the worst-case bias becomes even larger when extending to a broader class of models where additive separability does not hold. To obtain correct estimates on the model parameters, one should implement separate regressions as follows: \begin{equation} \begin{aligned} &\ensuremath{\mathbb{E}}[Y\mid D=1, W=w, P=p] = \mu_1 + w'\tau_1 + \rho_1 \lambda_1(p) + [w\lambda_1(p)]'\eta_1 \\ &\ensuremath{\mathbb{E}}[Y\mid D=0, W=w, P=p] = \mu_0 + w'\tau_0 + \rho_0 \lambda_0(p) + [w\lambda_0(p)]'\eta_0 \end{aligned} \end{equation} where $P = \ensuremath{\mathbb{P}}(D=1\mid Z, W)$ is the propensity score, and \[ \lambda_1(p) = \frac{1}{p} \int_0^p h(u) du \quad \text{and} \quad \lambda_0(p) = \frac{1}{1-p} \int_p^1 h(u) du. \] However, suppose a researcher mistakenly imposes additive separability. Thus $\rho_d(w)$ is considered as a constant in the model with $\eta_0 = \eta_1 = 0_{L\times 1}$. As a result, shorter regressions with omitted interaction terms $W\lambda_d(P)$ will be implemented: \begin{equation} \begin{aligned} &Y = \tilde{\mu}_1 + W'\tilde{\tau}_1 + \tilde{\rho}_1 \lambda_1(P) + Y^{\perp W, \lambda_1(P)\mid D=1} &\text{conditional on } D = 1\\ &Y = \tilde{\mu}_0 + W'\tilde{\tau}_0 + \tilde{\rho}_0 \lambda_0(P) + Y^{\perp W, \lambda_0(P)\mid D=0} &\text{conditional on } D = 0 \end{aligned} \end{equation} where $Y^{\perp X\mid D=d}$ denotes the OLS residual of $Y$ from regressing $Y$ on $(1,X)$ conditional on subsamples $D=d$: \[ Y^{\perp X\mid D=d} \equiv Y - \tilde{X}'\ensuremath{\mathbb{E}}[\tilde{X}\tilde{X}'\mid D=d]^{-1}\ensuremath{\mathbb{E}}[\tilde{X}Y\mid D=d]. \] where $\tilde{X} = (1,X')'$. The next lemma compares the true parameters $(\rho_d, \tau_d)$, obtained as estimands from the correctly specified model in (ref), with those from the misspecified regressions in (ref). \begin{lemma} Under the DGP (ref), the bias on $\rho_d$ equals \begin{equation} \tilde{\rho}_d - \rho_d = \frac{\operatorname{cov}(W'\lambda_d(P), \lambda_d(P)^{\perp W\mid D=d} \mid D=d)}{\operatorname{var}(\lambda_d(P)^{\perp W\mid D=d}\mid D=d)} \eta_d \end{equation} and the bias on $\tau_d$ equals \begin{equation} \tilde{\tau}_d - \tau_d = \ensuremath{\mathbb{E}}[(W^{\perp \lambda_d(P)\mid D=d})(W^{\perp \lambda_1(P)\mid D=d})'\mid D=d]^{-1}\ensuremath{\mathbb{E}}[(W^{\perp \lambda_d(P)\mid D=d})(W'\lambda_d(P))\mid D=d] \eta_d \end{equation} \end{lemma} The bias formulas are direct consequences of Frisch–Waugh–Lovell (FWL) theorem. It shows that the degree of treatment effects heterogeneity of the covariate $W$, measured by $\eta_d$, has a nontrivial impact on the bias of estimands. The bias formulas (ref) and (ref) simplify further under the following additional assumptions: \begin{assumption} Under the DGP (ref), the following conditions hold: \begin{enumerate} • $W$ is uncorrelated with the control functions of propensity score conditional on treatment status: For $d = 0,1,$ \begin{align*} & \operatorname{cov}(W, \lambda_d(P)\mid D=d) = 0_{L\times 1} \\ &\operatorname{cov}(WW',\lambda_d(P)\mid D=d) = 0_{L\times L} \\ &\operatorname{cov}(W, \lambda_d(P)^2\mid D=d) = 0_{L\times 1}. \end{align*} • $W$ is uncorrelated with treatment: $\ensuremath{\mathbb{E}}[W\mid D=1] = \ensuremath{\mathbb{E}}[W\mid D=0].$ \end{enumerate} \end{assumption} Assumptions (ref).1 posits that the covariate $W$ and the control function $\lambda_d(P)$ are uncorrelated up to their second moment, given the treatment or control groups. This assumption is notably strong, suggesting that the covariate $W$ cannot cause significant variation on the propensity scores within both groups. However, the bias of treatment effect estimands is shown to persist even under this stringent assumption. Only if researchers are willing to assume that $W$ is completely “irrelavant” under an additional Assumption (ref).2, the misspecified model (ref) would yield correct estimands of average treatment effects and the slope of MTEs as the ones obtained from the true model (ref). I focus on the bias for ATE and slope of MTE curve because the shape of the unconditional MTE, $\ensuremath{\mathbb{E}}[Y_1-Y_0\mid U=u]$, is uniquely determined by these quantities. I also consider the conditional ATE that captures treatment effect heterogeneity induced by covariates. These parameters are defined below: \begin{align*} & \text{ATE} = \ensuremath{\mathbb{E}}[Y_1 - Y_0] = \mu_1 - \mu_0 + \ensuremath{\mathbb{E}}[W'(\tau_1 - \tau_0)] \\ & \text{CATE} = \ensuremath{\mathbb{E}}[Y_1 - Y_0\mid W=w] = \mu_1 - \mu_0 + w'(\tau_1 - \tau_0) \\ & \text{Slope} = \ensuremath{\mathbb{E}}[\rho_1(W) - \rho_0(W)] = \rho_1 - \rho_0 + \ensuremath{\mathbb{E}}[W'(\eta_1 - \eta_0)]. \end{align*} Under (misspecified) additive separability, the researcher may estimate these effects by \begin{align*} & \widetilde{\text{ATE}} = \tilde{\mu}_1 - \tilde{\mu}_0 + \ensuremath{\mathbb{E}}[W'(\tilde{\tau}_1 - \tilde{\tau}_0)] \\ & \widetilde{\text{CATE}} = \tilde{\mu}_1 - \tilde{\mu}_0 + w'(\tilde{\tau}_1 - \tilde{\tau}_0) \\ & \widetilde{\text{Slope}} = \tilde{\rho}_1 - \tilde{\rho}_0. \end{align*} The next theorem shows the difference between the causal parameters and their corresponding estimands under Assumption (ref). \begin{theorem} Under the model (ref), \begin{enumerate} • Suppose Assumption (ref).1 holds, then \begin{align*} \widetilde{\text{ATE}} - \text{ATE} &= (\ensuremath{\mathbb{E}}[W] - \ensuremath{\mathbb{E}}[W\mid D=1])'\eta_1\times \ensuremath{\mathbb{E}}[\lambda_1(P)\mid D=1] \\ &\quad - (\ensuremath{\mathbb{E}}[W] - \ensuremath{\mathbb{E}}[W\mid D=0])'\eta_0 \times \ensuremath{\mathbb{E}}[\lambda_0(P)\mid D=0] \\ \widetilde{\text{CATE}} - \text{CATE} &= (w - \ensuremath{\mathbb{E}}[W\mid D=1])'\eta_1\times \ensuremath{\mathbb{E}}[\lambda_1(P)\mid D=1] \\ &\quad - (w - \ensuremath{\mathbb{E}}[W\mid D=0])'\eta_0 \times \ensuremath{\mathbb{E}}[\lambda_0(P)\mid D=0] \\ \widetilde{\text{Slope}} - \text{Slope} &= \left(\ensuremath{\mathbb{E}}[W\mid D=1] - \ensuremath{\mathbb{E}}[W\mid D=0]\right)'(\ensuremath{\mathbb{P}}(D=0) \eta_1 + \ensuremath{\mathbb{P}}(D=1) \eta_0). \end{align*} • Suppose Assumption (ref).1 and (ref).2 hold, then $\widetilde{\text{ATE}} = \text{ATE}$ and $\widetilde{\text{Slope}} = \text{Slope}$. \end{enumerate} \end{theorem} \begin{remark} While the estimand of the ATE may remain unbiased under both conditions outlined in Assumption (ref), it is important to note that the bias on CATE does not necessarily disappear under such assumption. Consequently, researchers should be cautious about interpreting CATE estimates in short regressions (ref) even if they have justified the validity of Assumption (ref). \end{remark} From this theorem, it becomes evident that the bias on ATE and the slope of MTE are driven by two main factors: the magnitude of $\eta_d$, which measures the heterogeneous effects of $W$, and the extent to which the covariate $W$ is unbalanced between the treatment and control groups. Therefore, the bias on those estimands can be quite significant if we omit heterogeneous effects of covariates that vary with unobserved heterogeneity $U$. In Online Appendix (ref), I provide a numerical example illustrating that such bias can even alter the sign of the ATE estimand when additive separability is mistakenly imposed. Due to potential bias from misspecified additive separability, Online Appendix (ref) shows how to apply the proposed methodology when conditioning on a set of discrete covariates using the Šidák–Bonferroni correction. For researchers who wish to impose additive separability in MTE models, Online Appendix (ref) provides guidelines on how to extend the robust inference procedure to that setting. \section{Monte Carlo Simulation} In this section, I compare the finite-sample performance of the proposed MLC test with the classical Wald test using simulated data from a simple quadratic MTE model with a three-valued instrument. The power comparison between conditional Wald tests, MLC tests, and classical Wald tests in a linear MTE model can be found in Online Appendix (ref). Consider a DGP specified as follows: For each $d = 0,1,$ \begin{align*} & Y_d = \mu_d + V_d \\ & D = \ensuremath{\mathbbm{1}}[U \leq p(Z)] \\ & V_d = \rho_{d1}\left(U - \frac{1}{2}\right) + \rho_{d2}\left(U^2 - \frac{1}{3}\right) + e_d, \end{align*} where $Z$ is uniformly distributed over three points $\{z_0,z_1,z_2\}$ and is independent of $(U, e_1, e_0)$. The error terms $(e_1, e_0)$ follow a joint normal distribution with zero mean and covariance matrix $\Sigma_e = 0.5 \cdot I_{2\times 2}$. For simplicity, I set $\mu_1 = \mu_0 = 0$ and impose strong endogeneity by specifying \[ \rho_{11} = \rho_{12} = -5 \quad \text{ and } \quad \rho_{01} = \rho_{02} = 5. \] Regarding the specification of the propensity score $p(z)$ for $z \in \{z_0,z_1,z_2\}$, I fix $p(z_0) = 0.5$ and let $(p(z_1), p(z_2))$ vary across $(0.05,0.95)^2$ to generate a variety of degrees/directions of weak identification. By drawing 2,000 i.i.d. simulation samples, I implement the Wald test and the MLC test with AR weight $a = 0.05$ to assess the null hypothesis on testing ATE $H_0: \mu_1 - \mu_0 = 0$. The average null rejection rates based on 5% significance level are displayed in Figure (ref). \begin{figure}[h!] \caption{3D Plot of Empirical Size for MLC and Wald Tests} \begin{tablenotes} Note: The plots show empirical rejection rates for testing the null hypothesis $H_0: \text{ATE} = 0$ at the 5% significance level. The upper panel presents three-dimensional surfaces of rejection rates, while the lower panel shows the corresponding contour plots. Each plot varies $p(z_1)$ and $p(z_2)$ across $(0.05,0.95)$ with $p(z_0)$ fixed at 0.5. The contour lines correspond to rejection rates of 2.5%, 5%, 7.5%, 10%, and 12.5%, with regions between these lines representing intermediate rejection rates. Darker regions indicate lower rejection rates. \end{tablenotes} \end{figure} The null rejection rates reveal distinct patterns in tests performance across propensity score variation. The conventional Wald test exhibits under-rejection when the three propensity scores are similar, indicating trivial power in this weakly identified scenario to be shown below. When the propensity scores cluster at two points (i.e., when either $p(z_1)$ or $p(z_2)$ equals $p(z_0) = 0.5$), the ATE parameter becomes partially identified, and the Wald test's rejection rates reach to 14%, substantially exceeding the nominal 5% significance level. This demonstrates the Wald test's invalidity under identification failure. In contrast, the proposed MLC test maintains proper size control across all propensity score combinations, demonstrating its robustness to weak identification. \begin{figure}[htbp] \caption{Power Curves of the MLC and Wald Test} \begin{subfigure}{0.95\textwidth} \caption{$p(z) = [0.2,0.5,0.8]$} \end{subfigure} \begin{subfigure}{0.95\textwidth} \caption{$p(z) = [0.2,0.5,0.5]$} \end{subfigure} \begin{subfigure}{0.95\textwidth} \caption{$p(z) = [0.4,0.5,0.6]$} \end{subfigure} \begin{tablenotes} Note: {Testing ATE at values on $[-5,5]$ with the true effects fixed at zero. The significance level is set at 5%. The sample size equals 2,000 and the average rejection rates are computed with 2,000 independent Monte-Carlo simulations. The two dashed lines around zero in subfigure (b) characterize the identified set of ATE by imposing a parameter space $[-10,10]$ for each element of $\theta$.} \end{tablenotes} \end{figure} Figure (ref) compares the power of the MLC and Wald tests across different identification scenarios. In Panel (a), where propensity scores are well-separated (strong identification), the Wald test achieves asymptotic efficiency, and the MLC test delivers comparable rejection power. In Panel (b), where the instrument takes only two distinct values (partial identification), the Wald test exhibits not only size distortion at the null but also power loss at alternatives. In contrast, the MLC test maintains correct size and almost surely rejects distant alternatives outside the identified set. In Panel (c), where propensity scores are clustered together (weak identification), the Wald test's rejection rates fall consistently below the nominal level, indicating negligible power, while the MLC test retains nontrivial power at distant alternatives. \section{Empirical Application to Misdemeanor Prosecution} In this section, I revisit the empirical analysis by agan/doleac/harvey:2023, who examined the causal effects of misdemeanor prosecution on defendants' subsequent criminal activity. Based on the quasi-randomized assignment of nonviolent misdemeanor cases to assistant district attorneys (ADAs), their LATE and MTE estimates demonstrate that nonprosecution leads to a large reduction in the likelihood of defendants' future criminal involvement over the next two years. To study the policy effects of increasing nonprosecution, they analyzed the impact of imposing a presumption of misdemeanor nonprosecution by relying on an existing policy change issued by a new district attorney. In this section, I use their data to answer some policy-relevant questions concerning the exogenous change on ADA nonprosecution rates. My approach does not require the additional data or information related to an actual policy that is already implemented, but instead extrapolates causal effects to address this issue. Similar to their analysis, I use the MTE framework to extrapolate treatment effects outside compliers. However, my goal is to analyze causal effects of implementing several counterfactual policies on ADA leniency, which were not explored in their empirical studies. To adapt their data into the framework considered in this paper, I implement the proposed inference procedure conditional on each court and use ADAs' identity as the discrete instrument since ADAs are randomly assigned conditional on courts and time. To avoid bias introduced by many instruments (see the references in mikusheva/sun:2022 for further discussions), I combine ADAs into a total of 15 groups for each court, excluding the smaller courts (BRI, CHE, and CHA) due to insufficient ADA counts. Let the treatment $D$ be the nonprosecution status of defendants (that takes value 1 if the defendant is not prosecuted) and the outcome $Y$ be the indicator of subsequent criminal complaints within two years postarraignment. To validate the strong overlap condition on propensity scores, I focus on ADAs that handle more than 50 cases and have nonprosecution rate at least 0.025 within each court. The summary statistics of ADAs' nonprosecution rates conditional on each court are provided in Table (ref) below. From this table we can see that the range of propensity scores varies from 0.21 to 0.50. When employing a high-order polynomial MTE model such as the cubic specification in Figure IV of agan/doleac/harvey:2023 in this setup, the finite-sample performance of confidence sets may be too poor to guarantee valid coverage. Specifically, in Online Appendix (ref), I show the evidence of weak instruments in cubic and quartic MTE models, while the classical $F$ test (with threshold 10) does not deliver the same conclusion. The reason is that $F$ statistic aims to detect deviations from the null where all propensity scores are equal to each other. However, this is not sufficient to strongly identify MTE models with flexible structure that require three or more propensity score to be well separated from each other (see Figure (ref)). \begin{table}[htbp] \caption{Summary Statistics of Nonprosecution Rates across Courts} \begin{tabular}{lccccc} \toprule \multirow{2}[2]{*}{Court} & \multirow{2}[2]{*}{Number of ADA} & \multicolumn{3}{c}{Nonprosecution Rate } & \multirow{2}[2]{*}{Sample Size} \\ & & Min & Mean & Max & \\ \midrule SBO & 16 & 0.04 & 0.24 & 0.54 & 3,921 \\ EBOS & 22 & 0.08 & 0.32 & 0.53 & 7,566 \\ WROX & 43 & 0.05 & 0.37 & 0.55 & 8,905 \\ BMC & 65 & 0.04 & 0.17 & 0.4 & 9,593 \\ ROX & 66 & 0.06 & 0.16 & 0.33 & 13,333 \\ DOR & 73 & 0.04 & 0.13 & 0.43 & 13,523 \\ CHA & 4 & 0.09 & 0.24 & 0.42 & 362 \\ CHE & 12 & 0.07 & 0.16 & 0.3 & 872 \\ BRI & 5 & 0.14 & 0.29 & 0.48 & 885 \\ & & & & & \\ Total & 262 & 0.04 & 0.22 & 0.55 & 58,960 \\ \bottomrule \end{tabular} \begin{tablenotes} Notes: Courts include South Boston (SBO), East Boston (EBOS), West Roxbury (WROX), Central (BMC), Roxbury (ROX), Dorchester (DOR), Charlestown (CHA), Chelsea (CHE), and Brighton (BRI). \end{tablenotes} \end{table} For all the counterfactual experiments, I consider an exogenous change to the ADA nonprosecution rates. Let $p^\epsilon(z)$ be the counterfactual propensity score of not prosecuting defendants, where $\epsilon$ is a nonnegative scalar (or vector) denoting the degree of deviation from status quo. I compare the induced outcome $Y^\epsilon$, defined as \[ Y^{\epsilon} = Y_1 \ensuremath{\mathbbm{1}}[U \leq p^{\epsilon}(Z)] + Y_0 \ensuremath{\mathbbm{1}}[U > p^{\epsilon}(Z)] \] with the observed outcome $Y$ in the data. Similar to the policy invariance assumption imposed by heckman/vytlacil:2005, suppose that this counterfactual change on propensity score does not shift the distribution of potential outcomes and unobserved cost $U$ of being selected into treatment, therefore omitting any “general equilibrium” effects that may arise after the change of prosecution rates. I consider three different comparisons between the counterfactual outcome $Y^\epsilon$ and the observed outcome $Y$: \begin{enumerate} • Non-normalized policy effects: \begin{equation} \alpha(\epsilon) \equiv \ensuremath{\mathbb{E}}[Y^\epsilon - Y]. \end{equation} This criterion follows from heckman/vytlacil:2001. The term “non-normalized” distinguishes these effects from the policy-relevant treatment effects discussed below. • Policy-relevant treatment effects (PRTE) \begin{equation} \bar{\alpha}(\epsilon) \equiv \frac{ \ensuremath{\mathbb{E}}[Y^\epsilon - Y]}{\ensuremath{\mathbb{E}}[p^{\epsilon}(Z) - p(Z)]}. \end{equation} This criterion follows from heckman/vytlacil:2005. This quantity can be interpreted as the treatment effect for units shifted into treatment via the counterfactual policy. Note that the additive/proportional PRTEs suggested by carneiro/heckman/vytlacil:2010,carneiro/heckman/vytlacil:2011 and mogstad/torgovitsky:2018 are special cases of $\bar{\alpha}(\epsilon)$ for particular choices of $p^{\epsilon}$, as described in equations (ref) and (ref) below. • Marginal policy-relevant treatment effects (MPRTE) \begin{equation} \alpha_{+}(0) \equiv \lim_{\epsilon \searrow 0} \frac{ \ensuremath{\mathbb{E}}[Y^\epsilon - Y]}{\ensuremath{\mathbb{E}}[p^{\epsilon}(Z) - p(Z)]}. \end{equation} This criterion follows from carneiro/heckman/vytlacil:2010. When $p^{\epsilon}(z) \geq p(z)$ for all $z\in\operatorname*{supp}(Z)$, this parameter can be interpreted as the average treatment effects for marginal defendants at the edge of being prosecuted. The equivalence between average MTE and MPRTE $\alpha_+(0)$ has been established by carneiro/heckman/vytlacil:2010. \end{enumerate} Next, I discuss the counterfactual propensity scores of interests and present confidence sets for these policy effects. \subsection{Marginal policy relevant treatment effects} First, suppose policymakers increase the ADAs' nonprosecution rates up to a positive constant $\epsilon$ or to a proportion $\epsilon$, where $\epsilon > 0$. Consider the additive or proportional marginal PRTE $\alpha_+(0)$ under the new policy: \begin{align} \text{Additive change:} \quad & p^{\epsilon}(z) = p(z) + \epsilon, \\ \text{Proportional change:} \quad & p^{\epsilon}(z) = (1 + \epsilon) p(z). \end{align} \begin{table}[h!] \caption{95% (top) and 90% (bottom) Confidence Sets for Additive MPRTE ${\alpha}_+(0)$} \begin{tabular}{lccc} \toprule & MTE polynomial & Wald & MLC \\ \midrule Uncond. & Linear & [-0.19, -0.11] & [-0.17, -0.13] \\ Uncond. & Quadratic & [-0.19, -0.11] & [-0.18, -0.13] \\ Uncond. & Cubic & [-0.17, -0.06] & [-0.15, -0.08] \\ Uncond. & Quartic & [-0.18, -0.05] & [-0.17, -0.06] \\ & & & \\ Average & Linear & [-0.24, -0.03] & [-0.23, -0.03] \\ Average & Quadratic & [-0.29, -0.01] & [-0.33, 0.09] \\ Average & Cubic & [-0.32, 0.03] & [-0.79, 0.59] \\ Average & Quartic & [-0.39, 0.01] & [-0.76, 0.61] \\ \midrule Uncond. & Linear & [-0.18, -0.12] & [-0.16, -0.14] \\ Uncond. & Quadratic & [-0.19, -0.12] & [-0.17, -0.14] \\ Uncond. & Cubic & [-0.17, -0.07] & [-0.14, -0.09] \\ Uncond. & Quartic & [-0.17, -0.06] & [-0.13, -0.08] \\ & & & \\ Average & Linear & [-0.22, -0.04] & [-0.21, -0.06] \\ Average & Quadratic & [-0.27, -0.04] & [-0.28, 0.01] \\ Average & Cubic & [-0.29, -0.00] & [-0.60, 0.34] \\ Average & Quartic & [-0.35, -0.02] & [-0.63, 0.38] \\ \bottomrule \end{tabular} \begin{tablenotes} Note: {Results for the classical Wald and MLC confidence sets under polynomial specifications up to the fourth order. The counterfactual policy considered is a uniformly additive increase in nonprosecution rates by a marginal amount. The top panel reports the 95% confidence sets, and the bottom panel reports the 90% confidence sets. The “Uncond.” confidence sets are obtained without controlling for court identities, whereas the “Average’’ confidence sets are weighted averages of the court-specific confidence sets, with weights proportional to court size.} \end{tablenotes} \end{table} Table (ref) collects the 95% and 90% confidence sets on marginal policy effects $\alpha_+(0)$ for the additive leniency increase in equation (ref) by letting $\epsilon$ approach zero. The “unconditional” confidence sets are obtained without controlling for court identities, whereas the “average” confidence sets are weighted average of court-specific confidence sets weighted by the size of courts. The differences between these two sets of results highlight the importance of adjusting for court identity as a confounding variable. Incorporating court identity generally increases the uncertainty of causal effect predictions. Comparing average Wald confidence sets with average MLC confidence sets, the results show that weak identification is a potential concern for models with polynomial orders higher than cubic. In such cases, robust confidence sets may not be informative about the sign of causal effects due to the inability to estimate a highly flexible model by using limited variation of propensity scores. Although the ranges of both confidence sets are similar under the quadratic specification, Wald confidence sets remain negatively significant, whereas MLC confidence sets lose such significance due to weak identification. In Online Appendix (ref) (Table (ref)), I evaluate the strength of identification for the marginal PRTEs based on the additive change in (ref). The findings further confirm the presence of weak identification for MTE models with cubic or quartic orders. In addition to these aggregated results, Table (ref) in Online Appendix (ref) provides confidence sets conditional on each court, revealing significant heterogeneity in the causal impacts of increasing ADAs' leniency across courts. It is worth noting that increasing ADA leniency might be harmful to court ROX, indicating that defendants may be more likely to be deterred from increasing punishment after prosecution, rather than committing to further crimes. Similar results for proportional marginal PRTE based on (ref) can be found in Table (ref) in Online Appendix (ref). \subsection{Quota} Instead of mandating all ADAs to increase their nonprosecution rates simultaneously, it might be more convenient for policymakers to set up lower and upper bounds for ADA nonprosecution rates. For example, let $\epsilon = (\underline{\epsilon}, \bar{\epsilon}) \in [0,1]^2$ with $\underline{\epsilon} \leq \bar{\epsilon}$ denoting the lower and upper bounds of ADAs' nonprosecution rate, respectively. Then the counterfactual propensity score becomes \[ p^\epsilon(z) = \min(\max\{\underline{\epsilon}, p(z)\}, \bar{\epsilon}) \] Since this counterfactual propensity score is non-smooth in $p(z)$ which may invalidate the inferential results, I employ a smooth approximation as follows: \[ p^\epsilon(z; \phi) = -\frac{1}{\phi}\log\left(\frac{1}{e^{\phi p(z)} + e^{\phi \underline{\epsilon}}} + \frac{1}{e^{\phi \bar{\epsilon}}}\right). \] One can show that $p^\epsilon(z; \phi) \to p^\epsilon(z)$ as $\phi \to \infty$. In practice, I set $\phi = 30$ to approximate the counterfactual propensity score $p^\epsilon(z)$. I analyze the (non-)normalized policy effects when policymakers implement a lower bound $\underline{\epsilon} \in [0.05, 0.3]$ while maintaining $\bar{\epsilon} = 1$ for each court. For notational consistency, I use $\epsilon$ to denote the value of this lower bound instead of $\underline{\epsilon}$. \begin{figure}[h!] \caption{95% (top) and 90% (bottom) Average Confidence Sets for Non-normalized Policy Effects $\alpha(\epsilon)$ when Setting a Lower Bound $\underline{\epsilon}$ for Nonprosecution Rate} \begin{tablenotes} Note: {Results for the classical Wald and MLC confidence sets under polynomial specifications up to the third order. The counterfactual policy under consideration is a universal lower bound on the nonprosecution rates shown on the $x$-axis. The $y$-axis depicts changes in recidivism rates (in percentage points). The top row presents the 95% confidence sets, and the bottom row presents the 90% confidence sets. All confidence sets are computed as weighted averages of the court-specific confidence sets, with weights proportional to court size.} \end{tablenotes} \end{figure} Figure (ref) displays the average confidence sets for the non-normalized policy effects $\alpha(\epsilon)$ of implementing a lower bound on nonprosecution rates. Under linear extrapolation, the results suggest that imposing a lower quota on nonprosecution is beneficial. However, this evidence becomes less conclusive with quadratic and cubic specifications as the lower bound ${\epsilon}$ increases, reflecting limitations of available exogenous variation at higher extrapolation levels. The Wald confidence sets appear “spuriously precise” by failing to account for this limited IV variation. Consequently, policy recommendations derived from the robust MLC approach may differ from those based on the classical Wald approach. For instance, based on the 90% confidence sets from the quadratic MTE model, policymakers concerned with robustness to weak IVs may prefer setting the minimum nonprosecution rate below 20%, since the reductions in recidivism become insignificant above this threshold under robust confidence sets. This sign reversal could potentially be driven by the fact that defendants with greater risk of recidivism are more likely to be released as the lower bound of nonprosecution rate increases. \section{Conclusion} In this paper, I propose two robust inference procedures for a class of causal effects identified by MTEs with discrete instruments. For the linear MTE model, I introduce a conditional Wald test, which is simple to implement and asymptotically similar regardless of identification strength. For a broader class of MTE models that nest polynomial specifications, I propose a modified linear combination test that achieves uniform validity against weak identification and has satisfactory power properties under strong identification. Finally, I use the proposed methods to investigate the counterfactual effects of manipulating ADAs' leniency in the study of misdemeanor prosecution by agan/doleac/harvey:2023. There are several avenues for future research. First, while this paper focuses on using discrete variation from instruments, extending the weak IV analysis to a semiparametric MTE model with continuous propensity scores would be empirically relevant. In addition, the MLC test requires the choice of tuning parameter on the weight of AR statistics, which trades off the power of test under different identification strengths (see more discussion in Online Appendix (ref)). It would be interesting to investigate the optimal choice of this tunning parameter following the regret analysis of andrews:2016. Finally, extending robust inference techniques to accommodate many weak instruments would be valuable, especially considering the typically large number of judges involved in empirical studies jochmans:2023. \putbib