EconBase
← Back to paper

Estimation of Heterogeneous Treatment Effects Using a Conditional Moment Based Approach

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

68,801 characters · 6 sections · 66 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimation of Heterogeneous Treatment Effects Using a Conditional Moment Based Approach

\affil[1]{Department of Economics, Simon Fraser University}

abstractWe propose a new estimator for heterogeneous treatment effects in a partially linear model (PLM) with multiple exogenous covariates and a potentially endogenous treatment variable. Our approach integrates a Robinson transformation to handle the nonparametric component, the Smooth Minimum Distance (SMD) method to leverage conditional mean independence restrictions, and a Neyman-Orthogonalized first-order condition (FOC). By employing regularized model selection techniques like the Lasso method, our estimator accommodates numerous covariates while exhibiting reduced bias, consistency, and asymptotic normality. Simulations demonstrate its robust performance with diverse instrument sets compared to traditional GMM-type estimators. Applying this method to estimate Medicaid's heterogeneous treatment effects from the Oregon Health Insurance Experiment reveals more robust and reliable results than conventional GMM approaches. Keywords: Endogeneity; Conditional mean independence; Neyman-orthogonal moments; Lasso. JEL Classification: C13; C14.

Introduction

Heterogeneous treatment effects have been a central focus in the literature since Imbens1994. These models explore how the effects of policies, programs, or other interventions on a specific outcome may vary across individuals with distinct characteristics. To quantify heterogeneous treatment effects, researchers often augment standard linear models by including interaction terms between the treatment variable and relevant covariates (as discussed in Section 21.4, Chapter 21 of WooldridgeJeffreyM.2010Eaoc and Section 7.6, Chapter 7 of imbens_rubin_2015, Page 125). This extended model becomes linear in the treatment variable, the interaction term, and the covariates, with the parameters of interest typically encompassing the effects of the treatment and its interaction with some covariates. Furthermore, significant attention has been devoted to addressing unobserved heterogeneity (Wooldridge1997b, Wooldridge1999c). In this paper, we adopt a standard and widely used semi-parametric model to demonstrate the advantages of integrating a conditional moment-based approach into the estimation of the observed heterogeneous treatment effects, that is, the parameters in front of the treatment and the interaction terms.

Our work introduces a novel estimator, D-RSMD, extending the R-SMD estimator proposed by AS2021 to accommodate big data settings. Specifically, the D-RSMD estimator combines the debiased method with R-SMD, integrating these approaches for improved estimation procedures. This new estimator offers two key properties essential for applied researchers. While versatile and applicable across various scenarios, we specifically focus on estimating heterogeneous treatment effects to illustrate those properties. We look into the case where the outcome variable adopts a classic partially linear structure encompassing linear and non-parametric components. Within the linear part, the treatment variable is endogenous, and the interaction terms are the treatment multiplied by exogenous covariates.

The primary advantage of the D-RSMD estimator is its reliance on a single valid instrument, bypassing the need to explicitly model the relationship between treatment and the instrument. This eliminates the requirement to generate additional exogenous variables, assume specific functional forms for the first stage (as in 2SLS), or select moments for GMM estimation. The choice of generated variables, model specifications in the first stage, or moments significantly impacts estimation results and the asymptotic properties of treatment effects. This vital advantage is shared by Smooth Minimum Distance (SMD) estimators like the SMD estimator proposed and examined in LavergnePatilea2013 and the R-SMD estimator introduced by AS2021. Such an approach is crucial for avoiding model specification issues and ensuring robust estimation of treatment effects.

For example, consider a scenario where the treatment variable is endogenous and interacts with a covariate. The traditional GMM method typically requires two moments: a valid instrument and an additional generated variable, for instance, the interaction between the instrument and covariate, to estimate both parameters. If we employ 2SLS, we need a model specification assumption. However, if the covariate in the interaction term exhibits limited variation or lacks information, GMM estimates can become inconsistent with high variance, resembling a weak identification problem. In contrast, our D-RSMD method relies solely on a single valid instrument, eliminating the need for additional generated variables, first-stage model specification, or moment selection.

Another approach to addressing this issue brought about by interacting with the covariate is splitting the data set, which is effective when dealing with binary covariates and a known outcome model with a substantial number of observations. However, the D-RSMD method offers distinct advantages by leveraging all information from a single conditional moment restriction across all observations, producing competitive results with smaller standard errors (SEs), sometimes even smaller SEs than those from traditional GMM with the correct model specification. We illustrate this idea using simulations in the appendix.

The second advantage of the D-RSMD estimator is its applicability in big data settings. The R-SMD estimator introduced by AS2021 combines Robinson transformation with SMD estimation to estimate parameters in the linear portion of a partially linear model\footnote{see Robinsontrans and LiRacine2006 for combining Robinson transformation and 2SLS estimators.}. In their work, AS2021 employ the Nadaraya-Watson estimator to estimate conditional means, restricting the number of covariates to fewer than four (refer to Chapter 7 of LiRacine2006). However, when dealing with models involving a large number of covariates, as in their empirical application, they resort to dimension reduction techniques such as Principal Component Analysis (PCA).

In contrast, the D-RSMD method is tailored for big data scenarios, thus relaxing restrictions on the number of covariates. This is achieved by applying a method that works in big data settings, then incorporating the debiased method proposed by ChernozhukovVictor2018Dmlf. In their work, ChernozhukovVictor2018Dmlf employ various techniques such as Lasso, random forests, neural networks, boosted regression trees, ridge regression, and so on. Then they employ the debiased method to reduce the effect of bias introduced by those techniques, which involves Neyman-Orthogonalization. This method yields an orthogonalized first-order condition (FOC) akin to the Robinson transformation proposed in Robinsontrans for partially linear models.

Our work employs the Neyman-Orthogonalized method following Robinson transformation. This adaptation is necessary due to the additional layer introduced by the SMD estimation procedure within the objective function. Orthogonalizing the FOC with this layer helps mitigate bias arising from employing the method in big data settings. Consequently, our D-RSMD estimator extends the debiased method proposed by ChernozhukovVictor2018Dmlf into the settings of SMD and U-statistics. Recent work by escanciano2023debiased provides general results for the Neyman orthogonalization of U-statistic type moments, offering a general construction of orthogonal quadratic moment functions for obtaining debiased estimators and valid inference in settings with U-statistics. While our objective function resembles a special case of the function in escanciano2023debiased, it incorporates a special weighting component based on the Fourier Transform using two observations from SMD, adding another layer of complexity to the construction of the Neyman-orthogonal FOC. With this extra weighting component, our identifying moment cannot be linearized. Hence, the methods and the inference results of escanciano2023debiased assuming the existence of a linearized component in Equation (2.8) of escanciano2023debiased are not applicable in our work.

Our D-RSMD focuses on using the Lasso method. The Lasso method is not the only option we have; as discussed in ChernozhukovVictor2018Dmlf, there are lots of available estimation methods. Here, we use Lasso as an illustration. Some researchers use sieves for the non-parametric part of the model, and they need to specify how to select the number of sieves in their estimation procedure in practice. In comparison, our new method does not need to choose the number of polynomials for the sieve method. When we use the Lasso method, the tuning parameter is inside the penalty term for the estimation of the nuisance parameter. In practice, there are several ways to choose Lasso tuning parameters (see vandeGeerSaraEaTU), and we use cross-validation (ChernozhukovVictor2018Dmlf). When we know there are fewer than four covariates, the Nadaraya-Watson estimator is an alternative method. This estimation procedure also needs one tuning parameter, namely the bandwidth. The bandwidth can be chosen by rule of thumb or cross-validation in practice. We show that our D-RSMD estimator is consistent and $\sqrt{n}$ asymptotically normal under mild regularity conditions. Monte-Carlo simulations show that the D-RSMD estimator outperforms the R-SMD and GMM-type estimators in terms of much smaller bias and lower standard errors.

To demonstrate the benefits of our method, we examine the impact of Medicaid using data from the Oregon Health Insurance Experiment.\footnote{I would like to acknowledge A. Colin Cameron for drawing my attention to the Oregon Health Insurance Experiment and providing access to the dataset. The dataset is sourced from his and Pravin K. Trivedi's book in 2022 and is also available on the website \url{https://www.stata-press.com/data/mus2.html}.} Our estimator, using a single valid instrument in the Oregon Health Insurance Experiment, yields statistically significant results for heterogeneous treatment effects, unlike the GMM estimator. Moreover, our approach produces more reliable results without requiring a generated instrument variable. This section presents an empirical analysis of the new D-RSMD estimator, emphasizing its advantages in estimating treatment effects and associated standard errors. Furthermore, it contributes to the literature on the Oregon Health Insurance Experiment by uncovering heterogeneous treatment effects that traditional methods fail to detect.

The paper is organized as follows: Section (ref) introduces our framework, motivation, and the estimator. Section (ref) states the large sample properties for our estimator. Section (ref) presents the simulation results for finite samples with a large number of covariates. Section (ref) uses the D-RSMD estimator to estimate the effects of enrollment in Medicaid on outcomes of interest, based on the Oregon Medicaid health experiment. Additional results and proofs are in the Appendix.

Framework and Motivation

Our framework is based on a partially linear model featuring a binary treatment variable $W_i$. Our focus lies in understanding the treatment's impact on the outcome variable. However, it is important to note that the treatment is not always exogenous. Take, for instance, the Oregon Medicaid health experiment, where treatment corresponds to enrollment in Medicaid—a decision influenced by the choice of individual lottery winner, rendering enrollment endogenous. Moreover, the treatment effect may not be homogeneous; it can vary depending on certain covariate variables in $X_i$, a vector of $q_X$ covariates ($X_i = [X_{i1}', X_{i2}']'$). Thus, we seek to examine the treatment effect conditional on specific groups determined by the values of the covariates $X_{i1}$. The remaining part of the model consists of an unknown function of exogenous $X_i$. While we maintain a fixed dimension for $X_{i1}$ for illustrative purposes, the dimension of $X_i$ is unrestricted and can be either high or low-dimensional. Additionally, the outcome variable $y_i$ is not constrained to be binary. For studies involving binary outcome variables, a large number of covariates, and similar restricted dimension settings for estimating heterogeneous treatment effects, see NekipelovDenis2018ROML. Generally, we focus solely on scenarios where we have knowledge of the functional form of the treatment and its heterogeneity, representing an initial step toward utilizing conditional moment restriction to estimate unobserved heterogeneity.

equation[equation omitted — 115 chars of source]

The formal definition for $W_i \cdot X_{i1}' \theta_{wx0}$ is $\sum_{k=1}^{q_{X_1}} W_i X_{i1,k}\times \theta_{wx0,k}$ where $X_{i1,k}$ is a $k$-th element inside the vector $X_{i1}$. The treatment part of interest is $W_i + W_i \cdot X_{i1}$. The parameters that measure heterogeneous treatment effects, i.e. $\theta_{w0}$ and $\theta_{wx0}$, are our key parameters.

As in the previous discussion, $W_i$ and the interaction term between $W_i$ and $X_{i1}$ are endogenous. To estimate the key parameters, a vector\footnote{$Z_i$ denotes the general vector of instruments. If we only use one instrument, $Z_i$ will be the random assignment. If we use a vector of instruments, $Z_i$ is the vector including the random assignment and other instruments.} of instruments $Z_i$ is introduced. For instance, in the Oregon Medicaid health experiment, the instrument is the lottery outcome, that is, the random assignment\footnote{In the experiment, the lottery winners are allowed to enroll in the Medicaid program. Not every lottery winner enrolled. The treatment variable is the actual enrollment status.}. We maintain the conditional mean independence assumption for the instrument. With valid instrument and exogenous control variables, the conditional moment restriction is

equation[equation omitted — 47 chars of source]

We consider the single treatment case, where the covariates may be correlated with the instruments. Traditional estimators, such as GMM or 2SLS, typically require $1 + q_{X_1}$ moments or exogenous variables to identify (and estimate) $1 + q_{X_1}$ parameters within $\theta_{w0}$ and $\theta_{wx0}$. These moments or exogenous variables include functions of $Z_i$ and $X_{i}$. For instance, the interaction terms $Z_i \cdot X_{i1}$ are often employed as generated instruments or in constructing moments.

However, as discussed in Section (ref), if there is little variation in $X_{i1}$, the $Z_i \cdot X_{i1}$ interaction terms become less informative. In such cases, traditional GMM methods will fail to provide reliable estimates. For example, if $X_{i1}$ is binary, limited variation means that the majority of observations in $X_{i1}$ are either 0 or 1. With only one valid instrument, GMM encounters the problem of generating new moments when we need to estimate more than one parameter, leading to under-identification. While splitting the sample in this case can solve the problem, it decreases the number of observations, thus resulting in larger standard errors, as discussed in the appendix.

This dilemma is resolved by directly employing the conditional moment restriction. This restriction encapsulates all the information of each instrument, regardless of the number of instruments used. Consequently, this approach allows us to utilize just one instrument, such as the random assignment $Z_i$, to identify and estimate parameters associated with both $W_i$ and $W_i \cdot X_{i1}$. Similar approaches have been employed in the construction of other Bierens-type estimators, such as the Integrated Conditional Moment (ICM) estimators. Notable references include Bierens1982, Antoine2014, AS2021, and Tsyawo2022. We will demonstrate that our estimation strategy, relying solely on a single valid instrument (e.g., the random assignment $Z_i$), yields reliable inference for both parameters.

The function $f_{0,1}(.)$ represents the nuisance parameter, which is not of interest in our analysis. The unknown form of $f_{0,1}(.)$ allows $X_i$ to enter the model flexibly. The initial step in constructing our estimator involves employing a Robinson transformation, as described in Robinsontrans. This transformation involves subtracting the conditional expectation of $y_i$ with respect to the controls $X_i$. As a result, the term $f_{0,1}(X_i)$ vanishes, eliminating the need to assume any specific form for $f_{0,1}(X_i)$. With the Robinson transformation and the exogeneity assumption on $X_i$, Equation (ref) simplifies to: $$y_i - E(y_i|X_i) = (P_i-E(P_i|X_i))'\theta_0+ \epsilon_i$$ where $P_i = [W_i, W_i \cdot X_{i1} ']'$ and $\theta_0 = [\theta_{w0},\theta_{wx0}']'$. Denote

equation[equation omitted — 90 chars of source]

with $\tilde y_j \equiv y_j - E(y_j|X_j)$ and $\tilde P_j \equiv P_j-E(P_j|X_j)$.

This equation also illustrates that $P_j$ may encompass any known function form of the treatment, interactions, covariates, and so forth. The symbol $g_0$ represents all the unknown real-valued functions, i.e., nuisance parameters. At this stage, $g_0$ comprises a vector of $g_{0,y}(X_j)$ (or $E(y_j|X_j)$) and $g_{0,P}(X_j)$ (or $E(P_j|X_j)$). Through this transformation, we exclude $f_{0,1}(X_i)$ from the model and introduce $g_0$. This is also the first step in AS2021.

Even though we only have a conditional moment restriction, we can exploit all the information in it by using an infinite number of unconditional moments. This equivalence between a conditional moment restriction and an infinite number of unconditional moment restrictions is from Bierens1982 and used in constructing SMD-type or ICM-type estimators. Based on Equation ((ref)) and the equivalence, we have the following results:

eqnarray[eqnarray omitted — 168 chars of source]

$t$ is a vector of any real number.\footnote{Using higher order Fourier terms, for instance, $E[\epsilon_j(\theta_0, g_0) e^{it^2 Z_j}]$ will lead to implicit choice of $\mu(t)$ and assumptions about $k(.)$ introduced in the later discussion and in practice. We choose to work with $e^{it'Z_j}$ because of its explicit and easy-to-use practice results. This is also the choice of AS2021.} Next, we define the population objective function in Equation ((ref)). Inside Equation ((ref)), $\mu(t)$ is a strictly positive measure on the vector $t$.

eqnarray[eqnarray omitted — 128 chars of source]

The objective function involves the norm of a complex function. To estimate $\theta_0$, we need to find the derivative of the objective function, which is difficult to compute. Under the independent assumption for the population, the objective function has an alternative expression. In the Appendix, we show the equivalence between these two objective functions.

eqnarray[eqnarray omitted — 229 chars of source]

The objective function defined in Equation ((ref)) depends on the parameters $\theta$ and the nuisance parameters $g_0$, as specified in Equation ((ref)) after the Robinson transformation. Under the regularity assumptions provided subsequently, $\theta_0$ serves as the unique minimizer of the objective function $M_{\infty}(\theta,g_0)$, where $g_0$ represents a vector of $g_{0,y}(X_j)$ and $g_{0,P}(X_j)$ at this stage. AS2021 similarly define an objective function incorporating $\kappa_{j,l}$ as the special weighting component from the SMD method, summarizing all the information from the conditional moment. While this may resemble a special case of the general moment considered in escanciano2023debiased, we find that with $\kappa_{j,l}$, it is not possible to identify a function and constants satisfying Equation (2.8) of escanciano2023debiased. Consequently, the methodology and inference results of escanciano2023debiased do not apply. Therefore, we make assumptions and derive theorems tailored to our specific setting.

The FOC of $M_{\infty}(\theta,g_0)$ with respect to $\theta$ is

eqnarray[eqnarray omitted — 92 chars of source]

When we do not know $g_0$, the FOC of $M_{\infty}(\theta,g)$ is as follows:

eqnarray[eqnarray omitted — 128 chars of source]

The FOC defined in Equation ((ref)) provides an explicit form for $\theta_0$. This FOC extends the framework established by AS2021 to incorporate interaction terms within the parametric component of the model. Thus, the estimator, derived as the sample analog of $\theta_0$, is a direct extension of AS2021. It also represents a specialized instance of the methodology proposed by LavergnePatilea2013 when the bandwidth within $\kappa_{j,l}$ is held constant.

asu(Regularity assumptions) \\ \refstepcounter{subassumption} (\alph{subassumption}) \ignorespaces $E(\epsilon_i|X_i, Z_i) = 0$.\\ \refstepcounter{subassumption} (\alph{subassumption}) \ignorespaces $E(\tilde P_i|Z_i) \neq 0$ a.s. (with probability 1) with $\tilde P_i= P_i- E(P_i|X_i)$. \\ \refstepcounter{subassumption} (\alph{subassumption}) \ignorespaces $E(\tilde P_i \tilde P_i')$ is nonsingular. \\ \refstepcounter{subassumption} (\alph{subassumption}) \ignorespaces Let $f_{Z}(.)$ denote the density function of $Z_j$. We assume that $E(\tilde{P}_{j}|Z_j=.)f_Z(.)$ is $L_q$ for some $1\leq q\leq2$. \\ \refstepcounter{subassumption} (\alph{subassumption}) \ignorespaces $(y_l, W_l, X_l, Z_l)$ is an independent and identical copy of $(y_j, W_j, X_j, Z_j)$ with $l, j \in \{1,..., n\}$. \\ \refstepcounter{subassumption} (\alph{subassumption}) \ignorespaces Let $\mu$ be a given strictly positive measure on $\mathbb{R}^{q_z}$. Let $k(.)$ be the Fourier inversion integral induced by $\mu$, $k(Z_j-Z_l)=\int_{\mathbb{R}^{q_z}} e^{it'(Z_j-Z_l)} d \mu(t)$. We assume that $k(.)$ is a symmetric bounded density function on $\mathbb{R}^{q_z}$ and that its Fourier transform is strictly positive.

Assumption (ref) is the exogeneity assumption for the controls and instruments. In the potential outcomes setting, Assumption (ref) is also the Random Assignment Assumption. The model form implies the Exclusion Restriction Assumption as in imbens_rubin_2015, Page 558, that is, the value of the instrument does not affect the potential outcomes directly. Assumption (ref) is the relevant instrument assumption. Assumption (ref) and the corresponding Assumption in Robinsontrans are not the same. This assumption is made in a high-dimensional setting. For not invertible $E(\tilde P_i \tilde P_i')$, see the solution provided in ChernozhukovVictor2018Dmlf.

Assumption (ref) is for the measure $\mu(.)$. These conditions in Assumption (ref) are not very restrictive. $\mu$ has a $\mathbb{R}^{q_z}$ support and is symmetric. This leads us to have a symmetric $k(Z_j-Z_l)$. Also, the fourier transform of $k(Z_j-Z_l)$ is $\mu$. Thus, the fourier transform of $k(Z_j-Z_l)$ is strictly positive. There are many different available measures. Measures after Fourier inversion have the respective results. For the simplicity of computation, we focus on those $k(.)$ with explicit formulas. Specifically, we use the CDF of the Gaussian distribution in simulations and applications.

Under Assumption (ref), $\theta_0$ is the unique minimizer of $M_{\infty}(\theta,g)$ when $g=g_0$. Specifically, assumptions (ref) and (ref) guarantee the identification of the parameters that we are interested in under the FOC defined in Equation ((ref)). This is established in AS2021 and is easily extended to our setting.

However, given that the nuisance parameters in $g_0$ are unknown, AS2021 consider the sample analog of $\theta_0$ defined in Equation ((ref)) as an infeasible estimator. To obtain a feasible estimator, AS2021 estimate the nuisance parameters as the first step. They achieve this by estimating their nuisance parameters using Nadaraya-Watson estimators. They demonstrate that, under appropriate assumptions on the bandwidth, the bias introduced by the Nadaraya-Watson estimators on the estimation of the key parameters is sufficiently small, ensuring that the feasible and infeasible estimators of the key parameter share the same asymptotic properties. However, the Nadaraya-Watson estimator imposes a constraint on the number of covariates to satisfy these assumptions. This constraint may not be suitable when there are a large number of available covariates. Moreover, even with a small number of covariates, when $g \neq g_0$, the bias resulting from estimating the nuisance parameters affects the estimation of the key parameters, particularly when the sample size is small. This is because the FOC defined in Equation ((ref)) is not orthogonal to the nuisance parameter, leading to a more pronounced effect of the bias on the estimate of the key parameter. This is illustrated in the proof of Appendix.

To address this bias issue, we propose a new FOC that extends the Neyman-orthogonal estimator from ChernozhukovVictor2018Dmlf to the SMD and U-statistic settings. The Neyman-Orthogonal method constructs an FOC that is orthogonal to the nuisance parameters. This new FOC ensures that all partial derivatives with respect to the nuisance parameters are zero, effectively making it orthogonal to the bias introduced by estimators. The concept of partial derivatives with respect to a function is elucidated in ChernozhukovVictor2018Dmlf.

Following Neyman-orthogonalization, the new FOC for $\theta_0$ becomes less sensitive to the bias of the estimator for the nuisance parameters. It is, however, sensitive to the square of the bias. As long as the bias is of the order $o_p(n^{-1/4})$, we still obtain a $\sqrt{n}$-asymptotically normally distributed estimator for the key parameter. It is important to note that under this order, there are numerous estimation methods to choose from, including Lasso, Sieves, and Random Forest, among others.

Our new FOC for $\theta_0$ corresponding to Equation ((ref)) is

equation[equation omitted — 216 chars of source]

with $g_{0,\tilde P_{m}}(X_l) \equiv E[(P_{m} - g_{0,P}(X_m)) \kappa_{m,l} |X_l]$ and $g_{0,\kappa_{m,l}}(X_l) \equiv E[\kappa_{m,l} |X_l]$. $g_{0,\tilde P_{m}}(X_l)$ has the same dimension as $\tilde P_j$, and $g_{0,\kappa_{m,l}}(X_l)$ is a function of $X_l$: $ \mathbb{R}^{q_X} \to \mathbb{R}$.

$g_{0,\tilde P_{m}}(X_l)$ and $g_{0,\kappa_{m,l}}(X_l)$ are two additional parameters inside the nuisance parameter vector. The FOC defined in Equation ((ref)) has a partial derivative with respect to all nuisance parameters equal to zero. This is shown in the Appendix.

asuThe singular values of the matrix $E\left[ \kappa_{j,l} \left( \tilde P_j - g_{0,\tilde P_{m}}(X_l)g^{-1}_{0,\kappa_{m,l}}(X_l) \right) \tilde P_l' \right]$ are between $c_0$ and $c_1$ with $0<c_0<c_1$.

Assumption (ref) is the identification assumption. This is the similar assumption as Assumption 3.1 (e) in ChernozhukovVictor2018Dmlf. As a contrary, Chernozhukov2022 and escanciano2023debiased do not impose the identification assumption, because the structure of the $\Psi(.; \theta, g)$ of their works guarantee the identification. Here, due to the fact that our $\Psi(D; \theta, g)$ is constructed under the SMD method with the special weight $\kappa_{j,l}$, we need to assume additional identification assumption as in ChernozhukovVictor2018Dmlf. Additionally, this assumption can be examined. We show this detailed information in the identification section of the Appendix, Section (ref)\footnote{In the Appendix, we illustrate the identification issue for the group of SMD estimators, such as SMD, RSMD, and D-RSMD. We consider cases involving various types of instruments and models.}. In this section, we show that with the information of the true Data Generating Process (DGP), the determinant of the matrix $E\left[ \kappa_{j,l} \left( \tilde P_j - g_{0,\tilde P_{m}}(X_l)g^{-1}_{0,\kappa_{m,l}}(X_l) \right) \tilde P_l' \right]$ is deducible. Also, in the applied setting, if the matrix is singular, the resulting standard error of the estimator will be huge and abnormal. This provides an intuitive way to determine whether the matrix is invertible.

Equation ((ref)) gives us the identification of the true key parameter and the explicit forms of the infeasible and feasible estimators.

prop(Identification of $\theta_0$ using the orthogonalized FOC) \\ Under Assumptions (ref)-(ref) and FOC defined in Equation ((ref)) $$\theta^*_0 = E\left[\kappa_{j,l}\left( \tilde P_j - \frac{g_{0,\tilde P_{m}}(X_l)}{g_{0,\kappa_{m,l}}(X_l)} \right)\tilde P_l' \right]^{-1} E\left[\kappa_{j,l}\left( \tilde P_j - \frac{g_{0,\tilde P_{m}}(X_l)}{g_{0,\kappa_{m,l}}(X_l)} \right) \tilde y_l \right]$$

If we plug Equation ((ref)) into the formula for $\theta^*_0$ with error term $\epsilon_l$, we will have $\theta^*_0= \theta_0$. Thus, the sample analog of $\theta^*_0$ delivers an estimator for $\theta_0$. Because we have two distinct individuals inside the expectation (e.g., j and l), we need to replace the expectation by the average of a double summation to obtain the infeasible estimator under Equation ((ref)). The closed-form expression for the infeasible estimator, $\tilde \theta_{n,o}$, is:

equation*[equation* omitted — 350 chars of source]

The infeasible estimator $\tilde \theta_{n,o}$ depends on unknown nuisance parameters: $g_{0,\tilde P_{m}}(X_l)$, $g_{0,\kappa_{m,l}}(X_l)$, $E(y_l|X_l)$, and $E(P_l|X_l)$. All of the nuisance parameters are conditional expectation functions on covariates. Hence, in practice, we need to find estimators for these conditional expectation functions to obtain a feasible estimator. Replacing every nuisance parameter $g_0$ with estimators $\hat g$ such that $\hat g$ converges to $g_0$ at a rate of $o_p(n^{-1/4})$ will deliver the feasible estimator on $\theta_0$ with $\sqrt{n}$ asymptotic normality.

The nuisance parameters are functions of $X$ involving $q_X$ variables. If $q_X$ is very large, such as $q_X \geq n$, there will be overfitting. To mitigate this, we need to reduce the number of variables included in the regression. Using Lasso will be one of the solutions. In order to apply Lasso, we assume sparsity for the nuisance parameters; that is, the conditional means can be described with only a few non-zero parameters in front of $X$. The number of non-zero parameters, $s_0$, is allowed to grow at the rate of $o_p(n^{1/2}/log(q_X))$. Under the assumed rate for $s_0$, the Lasso estimation has the following property (as Equation (2.1) of Chapter 2 in vandeGeerSaraEaTU) on the order of mean square error: $$||X(\hat \beta - \beta_0)||_2^2/n = \mathcal{O}_p \left(\frac{s_0 log(q_X)}{n} \right) $$ where $\hat \beta$ is the Lasso estimator (ChernozhukovVictor2018Dmlf and vandeGeerSaraEaTU), and $||.||_2$ is the $\mathcal{L}_2$ norm. When $3<q_X <n$, we can still use the Lasso method with a higher assumed rate for $s_0$ to select variables. If $q_X \leq 3$, we can still use the Nadaraya-Watson estimator to estimate the nuisance parameter with a second-degree kernel. It is worth noting that $q_X \leq 3$ is a constraint that is seldom realistic in practice. Nevertheless, our estimation procedure allows us to accommodate larger values of $q_X$. Upon replacing every nuisance parameter with its estimate, we obtain the feasible estimator $\hat \theta_{n,o}$, referred to as the D-RSMD estimator.

equation[equation omitted — 448 chars of source]

since $E[\kappa_{j,l} \left(\widehat{\tilde{P}}_j - \frac{\widehat{g_{0,\tilde P_{m}}}(X_l)}{\widehat{g_{0,\kappa_{m,l}}}(X_l)} \right) \widehat{\tilde{P_l}}']$ is invertible.

This leads to the algorithm of our D-RSMD estimation procedure.

algorithm[algorithm omitted — 807 chars of source]

All of the nuisance parameters for the D-RSMD estimator in the simulation and empirical application are estimated by the Lasso method with cross validation. The maximum degree of the polynomial for the nuisance parameters is 5, which guarantees that the nuisance parameters can be approximated by 5 degree polynomials in all controls. Lasso with cross validation helps us select the controls and their polynomials.

Large Sample Theory

References such as LavergnePatilea2013 provide a general framework for analyzing the asymptotic properties of SMD estimators, even in cases where explicit forms of these estimators are not available. Building upon this work, AS2021 extend the methodology to accommodate semiparametric models and derive the asymptotic properties specific to the RSMD estimator. In this section, we establish the asymptotic properties of the D-RSMD estimator, demonstrating that the feasible estimator $\hat \theta_{n,o}$ and the infeasible estimator $\tilde \theta_{n,o}$ exhibit equivalent asymptotic behavior. The properties of the infeasible estimator $\tilde \theta_{n,o}$ are outlined as follows:

prop(Consistency and Asymptotic normality of $\tilde \theta_{n,o}$) \\ Under Assumption (ref) and iid assumption for the sample, $\tilde \theta_{n,o}$ is consistent for $\theta_0$, that is $ \tilde \theta_{n,o} \stackrel{p}{\rightarrow} \theta_0 $, and asymptotically normally distributed, $$\sqrt{n}(\tilde \theta_{n,o} - \theta_0) \xrightarrow{d} N \left(0, A^{-1}\Sigma \left(A^{-1} \right)' \right)$$ where $A = E\left[\kappa_{j,l}\left( \tilde P_j - \frac{g_{0,\tilde P_{m}}(X_l)}{g_{0,\kappa_{m,l}}(X_l)} \right)\tilde P_l'\right]$, and $\Sigma$ stands for the variance-covariance matrix involving terms of $(\tilde P_i,\epsilon_i,Z_i,X_i)$, which is defined in the proof section of the Appendix.

The asymptotic properties of $\tilde \theta_{n,o}$ are based on the corresponding properties of U-statistics. This is thoroughly discussed in Hoeffding1948, particularly in Theorem 7.1 on Page 320 of the book.

asu$\hat g$ converges to $g_0$ at a rate of $o_p(n^{-1/4})$. For the Lasso method, the number of non-zero parameters $s_0$ grows at the rate of $o_p(n^{1/2}/log(q_X))$. For the Nadaraya-Watson estimator, $\sqrt n \left( \sum_{s=1}^{q_z} h_s^4+ \left[\frac{1}{n h_1...h_{q_z}} \right] \right) = o(1)$ where $h$ is the bandwidth.
theorem(Consistency and Asymptotic normality of the D-RSMD estimator: $\hat \theta_{n,o}$) Under Assumptions (ref) - (ref), $\hat \theta_{n,o}$ is consistent, and has an asymptotically normal distribution, that is, $$\sqrt{n}(\hat \theta_{n,o} - \theta_0) \xrightarrow{d} N \left(0, A^{-1} \Sigma \left(A^{-1} \right)' \right),$$

Notably, the asymptotic variances of the D-RSMD and R-SMD estimators proposed by AS2021 differ, primarily due to additional terms introduced through the debiasing process in $A$ and $\Sigma$. Analytical comparison of these asymptotic variances is challenging, given the complexity introduced by these extra terms. Furthermore, we provide the explicit form of the estimator for variance under heteroskedasticity in the following:

eqnarray[eqnarray omitted — 386 chars of source]

with $C_n =\frac{1}{n(n-1)} \sum_{j = 1}^n \sum_{l \neq j}^n \kappa_{j,l} \left(\widehat{\tilde{P}}_j - \frac{\widehat{g_{0,\tilde P_{m}}}(X_l)}{\widehat{g_{0,\kappa_{m,l}}}(X_l)} \right) \widehat{\tilde{P_l}}'$.

Based on Theorem (ref), the feasible and infeasible estimators exhibit the same asymptotic distribution and demonstrate consistency. This is from Assumption (ref) on the convergence rate, which ensures that the bias introduced during nuisance parameter estimation has negligible impact on subsequent steps. For example, when the number of covariates ($q_X$) and the sample size are small, the D-RSMD estimator exhibits less bias compared to the R-SMD estimator, especially when both use the same Nadaraya-Watson estimator for nuisance parameter estimation. In ChernozhukovVictor2018Dmlf, simulations illustrate a comparison between Neyman orthogonal estimators and non-orthogonal estimators. Further details on bias analysis are presented in the subsequent section, specifically in the simulation section.

Simulation Study

We consider a partially linear model with one endogenous variable and its interaction in the linear part.

eqnarray[eqnarray omitted — 331 chars of source]

$I(.)$ is the indicator function. It will take the value one if the statement inside is true and zero otherwise. In the simulations, the key parameters $\theta_{w0}$ and $\theta_{wx0}$ are 2 and 3 respectively. The other parameters are in the following. $\theta_{z0} = 3$, $\theta_{z3} = 4$, $\alpha_{3q} =2$, $\beta_{2q} =-3$, $\alpha_{1q}=\beta_{1q} = 1$ for $q \leq S$ with $S$ the number of non-zero parameters, and $q_X$ ($q_X = 30$) the number of covariates. There are nonlinear terms in the model, so $S$ is chosen to be a half of $s_0$, the sparsity level. $\alpha_{1q} = \alpha_{3q} = \beta_{1q} = \beta_{2q} = 0$ when $q > S$. Here, $S = 5$. If $\beta_{2q}$ is 0 for all $q$, $y_i$ is linear in $X_i$. Otherwise, the model is partially linear.

We use a binary instrument to create a treatment variable and explore simulation outcomes. The results when we use a categorical instrument in the DGP are in the appendix. $Z_i$ is the instrument with $E(Z_i) = 0.318$. The covariate $X_i$ is modeled as $X_i^* + 0.4Z_i$, where $X_i^*$ follows a multivariate standard normal distribution. This implies a correlation between $X_i$ and $Z_i$. The errors $(\epsilon, v)$ are bivariate normally distributed with mean 0, variance 1, and covariance $\frac{4}{9}$. Consequently, the treatment variable $W_i$ is endogenous.

To assess the performance of our D-RSMD estimators, we conduct 5,000 Monte Carlo replications for each scenario, where $q_X \in \{3, 30\}$ and the sample size $n \in \{3000, 5000\}$. Our benchmark scenario involves 3000 observations and 30 control variables, which mirrors our empirical example. For instance, when examining the Oregon Health Insurance Experiment, we segment the dataset into three age-based groups, each containing around 6,000 observations and 21 covariates.

We report and compare simulation results for several estimators in our study. These include the D-RSMD estimator proposed in this paper (referred to as DRSMD-Lasso in Section (ref)), where we use the Lasso method for estimating nuisance parameters. We also consider the R-SMD estimator by AS2021 using Lasso, the R-GMM estimator combining Robinson Transformation with GMM (GMM-Lasso in Section (ref)), a GMM estimator treating $f_{0,1}(X_i)$ as linear in $X_i$, and a GMM (Oracle) estimator utilizing the true $f_{0,1}(X_i)$. We present results for the D-RSMD estimator with one and two instruments, along with results for the R-SMD estimator under both conditions. GMM-type estimators require at least two instruments to estimate two parameters. We list the available instrument sets used for estimation in the third column of each table.

All nuisance parameters for the D-RSMD estimator in both the simulation and empirical application sections are estimated using the Lasso method with cross-validation. We employ a maximum polynomial degree of 5 for the nuisance parameters. This involves generating 5-degree polynomials for each covariate and using the Lasso method with cross-validation to select relevant controls and their respective polynomials, resulting in predicted conditional expectations. This approach provides a non-parametric estimation similar to the sieve method by selecting polynomial degrees. Cross-validation is utilized to select the penalty level and mitigate the risk of overfitting. For RSMD and RGMM estimators with $q_X = 30$, we also adopt the same Lasso method for polynomial selection to ensure comparability across estimators.

table[table omitted — 3,465 chars of source]

We present the results in Table (ref). Firstly, focusing on the first panel with $q_X = 3$, we observe that D-RSMD using $Z_1$ achieves the lowest MAD and Med.SE compared to all estimators. The RRs are close to 5% for $\theta_{w0}$ and slightly oversized for $\theta_{wx0}$. Using the instrument $Z_2$ (not in the DGP) for D-RSMD results in higher MAD and Med.SE compared to utilizing the correct instrument. The D-RSMD estimator also performs well when employing $(Z_1, Z_1 X_{1})$ and $(Z_2, Z_2 X_{1})$. Comparing D-RSMD estimators, those using $Z_1$ or the pair of instruments including the correct instrument tend to yield lower MADs and Med.SEs, which is reasonable because working with the correct variable increases estimation precision. The Med.Bias column indicates that for $\theta_{w0}$, D-RSMD estimators using $(Z_1, Z_1 X_{1})$ and $(Z_2, Z_2 X_{1})$ exhibit higher bias than other D-RSMD estimators, suggesting that using $Z_1$ leads to lower bias in this DGP. This property of D-RSMD also holds when there are 30 controls in the model. However, using one instrument may not always bring a lower bias. This is shown in the appendix when $Z_2$ is the instrument used in the DGP.

The second panel of Table (ref) presents results for $q_X = 30$. Similar to the scenario with $q_X = 3$, among D-RSMD estimators, using $Z_1$ as the instrument yields better results, as evidenced by lower Med.Bias, MAD, and Med.SE values. Comparing the results obtained with $Z_1$ versus $Z_2$, we notice a slight distortion in the RR when $Z_1$ is employed. However, this size distortion diminishes when using a sample size of 5,000 observations per replication, as shown in the last panel of the table. Comparing D-RSMD with RSMD underscores the advantage of D-RSMD in producing less biased estimates, particularly when employing the Lasso method for nuisance parameter estimation.

Furthermore, comparing D-RSMD results with the GMM (Oracle) estimator reveals that D-RSMD estimates using $Z_1$ closely approximate the GMM (Oracle) estimates for $\theta_{w0}$ in terms of MAD, Med.SE, and RR. For $\theta_{wx0}$, D-RSMD results with $Z_1$ outperform the GMM (Oracle) estimator using $(Z_1, Z_1 X_{1})$, exhibiting lower MAD and Med.SE values. These findings suggest that D-RSMD with $Z_1$ achieves satisfactory performance due to its utilization of an infinite number of unconditional moments, equivalent to the conditional moment, unlike the GMM (Oracle) estimator, which relies solely on two unconditional moments.

Analyzing outcomes from instrument sets $(Z_1, Z_1 X_{1})$ and $(Z_1, Z_2)$ reveals notable differences between D-RSMD and GMM (Oracle) estimators. Specifically, for GMM (Oracle) estimators, $(Z_1, Z_1 X_{1})$ emerges as the preferred choice, exhibiting the lowest Med.Bias for both $\theta_{w0}$ and $\theta_{wx0}$. Conversely, GMM estimators using $(Z_1, Z_2)$ demonstrate scattered results and numerous extreme cases, with the RR approaching 0 for $\theta_{wx0}$, highlighting the challenges posed by instrument correlation within $(Z_1, Z_2)$ for GMM type estimators.

In contrast, the choice of instrument sets has a smaller impact on D-RSMD results compared to its effect on GMM (Oracle) estimators. This observation suggests that D-RSMD exhibits more stable performance when utilizing at least one valid instrument or an instrument closely related to the valid instrument in the instrument set, compared with GMM-type estimators.

Empirical Application

In this section, we investigate the heterogeneous treatment effects of Medicaid enrollment from a Randomized Controlled Trial (RCT) conducted in Oregon. In early 2008, Oregon expanded Medicaid enrollment through a lottery system for low-income, uninsured adults. This lottery provided random selection, enabling researchers to analyze the effects of health insurance on various outcomes, including medical, financial, and labour market impacts. This experiment has been extensively studied in various articles, such as Baicker2014 and Finkelstein2012. Detailed information and datasets related to this RCT are available on the public website.\footnote{See \url{https://www.nber.org/research/data/oregon-health-insurance-experiment-data}} In this paper, we utilize a dataset derived from the original datasets.\footnote{Special thanks to A. Colin Cameron for providing the dataset and bringing the Oregon Health Insurance Experiment to my attention through his and Pravin K. Trivedi's book in 2022. The dataset is also accessible at \url{https://www.stata-press.com/data/mus2.html}} The dataset comprises 18,572 observations.

Many studies have focused on the homogeneous treatment effects of Medicaid. Finkelstein2012 assumed homogeneity in their analysis but also explored the potential for heterogeneous treatment effects through regression analysis. Estimates from 2SLS rule out the possibility of heterogeneous treatment effects. However, this may be due to the misspecification of the first stage or problems with the generated instrument variables.

To address this problem, we employ the D-RSMD estimator to estimate heterogeneous treatment effects. We compare the results between the new and traditional estimation methods, both with and without the generated instrument variable, to obtain reliable inference.

Here, the heterogeneity arises from interaction terms involving indicators ($X_{1}$) for household income above 50% of the federal poverty line in 2008 (income), household income (hhincome), TANF (cash welfare assistance to low-income families), or cigarette smoking level (smoke). The lottery serves as the instrumental variable. We adopt similar controls to those used by Baicker2014 and Finkelstein2012, including household controls, lottery and survey wave indicators, and individual characteristics. The dependent variables include the current employment indicator, constructed from three indicators: hours of employment (employment), total out-of-pocket spending on medical care (out-of-pocket cost), and whether the individual currently owes money to a healthcare provider (debt for health). All dependent variables are derived from a mail survey conducted between July 2009 and March 2010, approximately one year after treatment.

Given that health conditions, employment status, and other outcomes are closely linked to individuals' ages, we identify age as a key source of heterogeneity. We derive an age group variable from the year of birth data. First, we calculate the age of each individual, resulting in ages ranging from 21 to 64 years. Next, we partition the original dataset into three or five subsets based on age to create the age-group variable. Each subsample contains approximately 5900 to 6900 observations. Individuals within the same age group share more similarities. In Section (ref), we present results categorized into three age groups (e.g., 21 to 35, 35 to 50, and 50 to 64). Here, the `agegroup` variable represents a categorical variable with three values, mirroring the age groups.

To provide a clearer comparison, we contrast the D-RSMD estimator with traditional estimators, such as the GMM estimator. The D-RSMD estimator is designed to cope with a nonparametric first stage and a partially linear second stage. The framework for the D-RSMD estimator is outlined as follows:

align*[align* omitted — 271 chars of source]

The framework for GMM estimators includes a linear first stage within the indicator function as well as a linear second stage for the dependent variable.

align*[align* omitted — 290 chars of source]

The model for GMM-Lasso estimators includes a linear first stage for the indicator function and a partially linear second stage for the dependent variable.

align*[align* omitted — 288 chars of source]

In Table (ref), we present preliminary results for three estimators of $\theta$. D-RSMD stands out as the sole estimator utilizing the lottery for estimating multiple parameters. We report D-RSMD results in two columns: using only the lottery ($Z_1$) and incorporating the lottery and its interaction term $(Z_1, Z_1 X_{1})$. This setup allows for a direct comparison of results across the two instrument sets and three estimators.

Using the DRSMD-Lasso estimator with the lottery as the sole instrument ($Z_1$), we obtain statistically significant estimates for Medicaid and the interaction between Medicaid and income at a 5% significance level. Similarly, we observe statistically significant estimates for income and the interaction term between Medicaid and age at a 15% significance level with the lottery. These findings corroborate the conclusions drawn in Section (ref), suggesting heterogeneous treatment effects in Medicaid related to age and income when employing the valid instrument alone. Conversely, analyzing outcomes using the GMM estimator or GMM-Lasso with the instrument set $(Z_1, Z_1 X_{1})$ would erroneously lead to the conclusion of no heterogeneity in treatment effects, underscoring the problems associated with using generated instruments or moments.

table[table omitted — 1,497 chars of source]

Results

In this section, we examine the potential heterogeneous treatment effects of Medicaid on health-related debt, with a focus on household income as a source of heterogeneity. We partition the dataset into three subsets based on age to facilitate intuitive interpretation. The complete table presenting all treatment effects is provided at the end of this section. Additional results for robustness checks are in Appendix Section (ref).

We consider two sets of instruments: the lottery (used exclusively for D-RSMD) and the pair of the lottery along with its interaction with $X_1$ (indicating income greater than 50% of the poverty line). To illustrate, the model for debt is specified as:

eqnarray*[eqnarray* omitted — 119 chars of source]

Table (ref) presents results for three age groups: 21–35, 35–50, and 50–64 years old. For the age group 35–60, the total number of observations is 6693. Heterogeneous treatment effects are observed among individuals. The estimates obtained by D-RSMD (using the lottery as an instrument or the lottery and its interaction with $X_1$ as instruments) are very similar to each other and to the GMM estimator (using the lottery and its interaction with $X_1$ as instruments). This indicates a strong and reliable interaction between the lottery and $X_1$. Notably, only the D-RSMD estimation procedure reveals statistically significant heterogeneous effects. The age group exhibits significant results for both $\theta_{w0}$ and $\theta_{wx0}$ when employing our new estimator with the valid lottery instrument, suggesting that the effect of Medicaid on debt for health varies depending on income level, thus indicating heterogeneity.

For the age group 35–60, the interpretation of the D-RSMD estimate (using $Z_1$ or $(Z_1, Z_1 X_{1})$) for $\theta_{w0}$ is that for people with an income 50% below the federal poverty line, Medicaid enrollment decreases their probability of owing money to health providers by 23.2 log points on average, ceteris paribus. For people with an income above 50% of the federal poverty line, Medicaid enrollment reduces their probability of owing money, but not by that much. This is reasonable because enrollment in Medicaid is not that critical to reducing debt for individuals with higher incomes compared with those with lower incomes.

Now focus on the $\theta_{wx0}$. Comparing the results of D-RSMD between the two instrument sets ($Z_1$ or $(Z_1, Z_1 X_{1})$), we find that the estimates are close, but the standard errors using $Z_1$ are substantially lower, suggesting that using one valid instrument provides a different result for statistical significance.

For the age group 21–35, all estimators generate similar results using $(Z_1, Z_1 X_{1})$. The coefficients for $\theta_{w0}$ are not statistically significant, and the ones for $\theta_{wx0}$ are statistically significant. It suggests that when households' incomes are above the 50% federal poverty line, individuals with Medicaid will be less likely to owe money to their health providers. It also suggests that Medicaid helps people with higher income levels more than it helps people with lower incomes. Using only the lottery as the instrument, our new procedure generates the opposite results. The interpretation is that for people with lower incomes, Medicaid enrollment decreases their probability of owing money to health providers by 19.9 log points on average, holding other variables constant. For people with higher incomes, the effect of Medicaid decreases.

Based on the D-RSMD results using only the lottery, we do not find support in the data to say that there are heterogeneous treatment effects for individuals between 21 and 35 years old. Furthermore, we discover that the results for individuals aged 21–35 differ significantly between the two instrument sets. It suggests that the generated interaction $Lottery \times X_{1}$ is invalid.

Next, we re-estimate a homogeneous model. In a homogeneous model, the interaction term is not included in the regression model. The traditional estimation method only requires one instrument to estimate one parameter. When we examine the homogeneous treatment effects using both instrument sets in Table (ref), all estimators yield similar results to the DRSMD-Lasso method using only the lottery as the instrument in the heterogeneous treatment effects panel. This suggests that using a valid instrument to estimate both parameters produces more reliable outcomes. In practical applications where the homogeneity of the model is uncertain, the DRSMD-Lasso method proves to be highly valuable. Its ability to provide reliable estimates under both homogeneous and heterogeneous conditions, coupled with its smaller standard errors compared to traditional methods, makes it a preferred choice. In particular, the smaller standard errors enhance the significance of results obtained through t-tests.

The estimates and standard errors of average treatment effects in Table (ref) are calculated based on Table (ref). We conduct the robustness check in the Appendix.

table[table omitted — 2,514 chars of source]
table[table omitted — 1,531 chars of source]