Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
83,126 characters · 14 sections · 53 citation commands
Estimating Marginal Treatment Effects under Unobserved Group Heterogeneity
\onehalfspacing
Assessing heterogeneity in treatment effects is an important issue for precise treatment evaluation. The marginal treatment effect (MTE) framework (heckman1999local, heckman2005structural) has been increasingly popular in the literature as it provides us with rich information about the treatment heterogeneity not only in terms of observed individual characteristics, but also in terms of unobserved individual cost of the treatment. Moreover, once the MTE is estimated, it can be used to build other treatment parameters such as the average treatment effect (ATE) and the local average treatment effect (LATE). For recent developments on the MTE approach, see, for example, cornelissen2016late, lee2018identifying, mogstad2018using, and mogstad2018identification.
While the conventional treatment evaluation methods can address heterogeneity across observable groups of individuals, many applications may exhibit “unobserved” group-wise heterogeneity in treatment effects for various reasons. For example, the presence of multiple treatment eligibility criteria may create unobserved groups. As a typical application, consider evaluating the causal effect of college education. Since schools typically offer a variety of admissions options such as entrance exams and sports referrals, this process classifies individuals into several groups, and the admission criteria to which each individual has applied is typically unknown to researchers. Such differences in admissions requirements may result in heterogeneous treatment effects of college education. Another potential reason for the presence of unobserved group heterogeneity is that the population may be composed of groups with different preference patterns. For instance, consider estimating the causal effect of foster care for abused children, as in doyle2007child. Here, the treatment variable of interest is whether the child is put into foster care by the child protection investigator. The author discusses a possibility that the child protection investigators may have different preference patterns that place relative emphasis on child protection. These examples suggest unobserved group patterns in the treatment choice process that may lead to some heterogeneity in treatment effects.
In this paper, we study endogenous treatment effect models in which individuals are grouped into latent subpopulations, where the presence of the latent groups is accounted for by a finite mixture model. Finite mixture approaches have been successfully used in various fields to analyze data from heterogeneous subpopulations (mclachlan2004finite). For example, many empirical studies on economics employ finite mixture approaches to address unobserved group heterogeneity (e.g., keane1997career; cameron1998life). However, the use of finite mixture models in treatment evaluation has been considered only in a few specific applications (e.g., harris2009impact; munkin2010disentangling; deb2018heterogeneous; samoilenko2018using). Compared with these studies, our modeling approach is applicable to various contexts and formally builds on rubin1974estimating's (rubin1974estimating) causal model by directly extending it to finite mixture models.
For this model, we develop identification and estimation procedure for the MTE parameters that can be unique to each latent group. The proposed group-wise MTE is a novel framework in the literature, which should be informative for understanding the heterogeneous nature of treatment effects by capturing both group-level and individual-level unobserved heterogeneity simultaneously. Importantly, as we discuss below, the presence of unobserved group heterogeneity threatens the validity of the conventional instrumental variables (IVs)-based causal inference methods, such as the conventional MTE approach and the two-stage least squares approach (2SLS) for estimating the LATE. Specifically, we demonstrate that the presence of unobserved heterogeneous groups may invalidate the monotonicity condition (e.g., imbens1994identification,heckman2018unordered).\footnote{ Recently, there has been an increasing number of studies that deal with situations where the conventional monotonicity condition would not be satisfied (e.g., lee2018identifying, mogstad2020causal, mogstad2020policy, mountjoy2019community, hoshino2021treatment). Among them, our model setup is closely related to the one in mogstad2020causal, mogstad2020policy in that these papers explicitly allow for heterogeneity in the treatment choice action associated with multiple IVs. } This result implies that the conventional MTE parameter may not even be well-defined under unobserved group-wise heterogeneity.
Our identification strategy builds on the method of local IV by heckman1999local. Our main identification result requires three key conditions. The first condition is that there exists a valid group-specific continuous IV. As in the standard IV estimation, each group-specific IV must satisfy that the IV is independent of unobserved variables and that it is a determinant of the treatment. The second condition is the exogeneity of group membership; that is, group membership is conditionally independent of the unobserved variables affecting treatment choices. Since the second condition may be demanding in some applications, we also provide supplementary identification results when group membership is endogenous. The third condition is the identifiability of the “first-stage” treatment choice equation as in the conventional MTE analysis. Since nonparametrically identifying finite-mixture binary response models is extremely challenging, restricting our attention to finite-mixture Probit models, we provide sufficient conditions for the identification of the treatment choice model.\footnote{ Several recent studies consider nonparametric identification of finite mixture models. For example, bonhomme2016non present an identification result based on repeated measurements data satisfying some independence property, and kitamura2018nonparametric develop nonparametric identification for a regression model with additive error. However, to the best of our knowledge, no existing results are applicable to our case. Nonparametric identification analysis on finite-mixture binary response models would be of great interest, but it still remains an open research question. }
Based on our constructive identification results, we propose a two-step semiparametric estimator for the group-wise MTEs. In the first step, we estimate a finite-mixture treatment choice model using a parametric maximum likelihood (ML) method. In the second step, the MTE parameters are estimated using a series approximation method. Under certain regularity conditions, we show that the proposed MTE estimator is consistent and asymptotically normally distributed.
As an empirical illustration, we investigate the effects of college education on annual income using the data for labor in Japan. We focus on heterogeneity caused by the following two latent groups: the first comprises individuals whose college enrollment decisions are mainly affected by regional educational characteristics (group 1), while the second is a group of individuals who are mainly affected by at-home study environment (group 2). More specifically, we employ regional characteristics, such as the local college enrollment rate and the rate of workforce participation for high school graduates, as the primary IVs for group 1. For group 2, we create variables that measure the quality of study environment at home, and use these as the IVs. Our empirical results indicate that for group 1, the treatment effect of college education is significantly positive if the (unobserved) cost of going on to a college is small. In contrast, we cannot find such heterogeneity for group 2.
\paragraph{Organization of the paper.}
Section (ref) introduces our model and presents our main identification result for the group-wise MTE. In Section (ref), we discuss the estimation procedure for the MTE parameters and prove its asymptotic properties. Section (ref) provides two additional discussions: first, on the identification of MTE when group membership is endogenous and, second, on the identification of LATE when only binary IVs are available. Section (ref) gives the results of Monte Carlo experiments. Section (ref) presents the empirical illustration, and Section (ref) concludes the paper.
In this section, we introduce our treatment effect model that allows for the presence of an unknown mixture of multiple subpopulations. Throughout the paper, we assume that the number of groups is finite and known, which is denoted as $S \in \mathbb{N}$. Each individual belongs to only one of the $S$ groups, and the group the individual belongs to, which we denote by $s \in \{1, \dots, S\}$, is a latent variable unknown to us. Our goal is to measure the causal effect of a treatment variable $D \in \{0, 1\}$ on an outcome variable $Y \in \mathbb{R}$ for each group separately. Let $Y^{(d)}$ be the potential outcome when $D = d$. Then, the observed outcome can be written as $Y = D Y^{(1)} + (1 - D) Y^{(0)}$. Suppose that the potential outcome equation is given by
where $X \in \mathbb{R}^{\mathrm{dim}(X)}$ is a vector of observed covariates, $\epsilon^{(d)} \in \mathbb{R}$ is an unobserved error term, and $\mu^{(d)}$ is an unknown structural function. This model specification is fairly general in that the functional form of $\mu^{(d)}$ is fully unrestricted and the distribution of the treatment effect $Y^{(1)} - Y^{(0)}$ can be heterogeneous across different groups.
Based on the latent index framework by heckman1999local,heckman2005structural, we characterize our treatment choice model as follows:
where $\mathbf{1}\{\cdot\}$ is an indicator function that takes one if the argument inside is true and zero otherwise, for $j \in \{ 1, \ldots , S \}$, $Z_j \in \mathbb{R}^{\mathrm{dim}(Z_j)}$ is a vector of IVs that may contain elements of $X$, $\epsilon_j^D \in \mathbb{R}$ is an unobserved continuous random variable, and $\mu_j^D$ is an unknown function. We allow for arbitrary dependence between $\epsilon_j^D$'s. Assume that for all $j$, the error $\epsilon_j^D$ is independent of $Z_j$'s conditional on $X$. Moreover, we require that each $Z_j$ includes at least one group-specific continuous variable to ensure that the function $\mu_j^D(Z_j)$ does not degenerate to a constant after conditioning the values of $(X, Z_1, \ldots, Z_{j-1}, Z_{j+1}, \ldots, Z_S)$.
To proceed, let $F_j(\cdot | X)$ be the conditional cumulative distribution function (CDF) of $\epsilon_j^D$ given $X$. Further, let $P_j \coloneqq F_j( \mu^D_j(Z_j) | X )$ and $V_j \coloneqq F_j ( \epsilon_j^D | X )$. By construction, each $V_j$ is distributed as $\text{Uniform}[0,1]$ conditional on $X$. Using these definitions, we can rewrite (ref) as follows: $D = \mathbf{1}\left\{P_j \ge V_j \right\}$ if $s = j$.
For the treatment choice model in (ref), we can interpret its meaning in several ways. The first interpretation is that there are actually multiple different treatment eligibility rules prescribed by policy makers. In the example of college enrollment, there are typically several different types of admissions processes for each school, for example, paper-based entrance exams, sports referrals, and so on. Such a situation would correspond to this first type of interpretation. Another interpretation is that there are several types of treatment preference patterns. For example, consider again $D = 1$ if an individual goes to college and $D = 0$ otherwise. Suppose that a common instrumental variable $Z$ is the introduction of a physical education requirement in colleges along with mandatory augmented athletics facilities.\footnote{ This example scenario is borrowed from heckman2005structural, Subsection 6.3. } When we specify the functional form of $\mu^D_j$ as in Remark (ref), we imagine that some people dislike physical education ($\gamma_{z1} < 0$) while others like it ($\gamma_{z2} > 0$). In this situation, we can view the treatment choice model (ref) as a binary response model with a discrete random coefficient.
Our main identification results are based on the following assumptions.
Assumption (ref)(i) is an exclusion restriction requiring that the IVs are conditionally independent of all unobserved random variables including the latent group membership. Assumption (ref)(ii) is somewhat demanding in that we require prior knowledge as to which variables may be relevant/irrelevant to the membership of each group. A similar assumption can be found in the econometrics literature on finite mixture models (e.g., compiani2016using). In practice, which IVs belong to which group needs to be determined on a case-by-case basis, using background knowledge and/or some economic theoretical framework underlying the presumed group structure in the data.\footnote{ Note that the groups are distinguished only by their susceptibility to different IVs. We are generally unable to know if these groups correctly represent the groups originally presumed by researchers. Therefore, we need to make a comprehensive interpretation of the estimated groups by combining our prior knowledge of the data and the estimation results for the IVs. To obtain a sharper interpretation for the groups, some researchers consider a model in which the group membership probability is a function of individual characteristics, similar to our model in (ref). Such a model is useful, but it does not allow us to know exactly who belongs to which group. This is a common limitation in the finite mixture framework. } If researchers do not have such information, simply comparing the information criterion values of the models with different IVs might be practically useful (see also the results in Table (ref) in Subsection (ref)). Note that in the absence of group-specific covariates, nonparametric identification of the treatment choice model would be in general infeasible, since when conditioned on the covariates, the observable choice probability is just a mixture of Bernoulli distributions, which is another Bernoulli distribution. The continuity of group-specific IV is necessary to nonparametrically identify the MTE parameter. As will be shown in Theorem (ref), the MTE parameter for group $j$ can be identified through a partial derivative with respect to $P_j$ conditioned on $(X, \{P_h\}_{h \neq j})$, which cannot be well-defined if all the IVs $Z_j$ are discrete.\footnote{ Even when all group-specific IVs are discrete, if additional parametric functional form restrictions are imposed, it would be possible to identify MTE, in a similar manner to brinch2017beyond. Investigation of such an approach is left for future research. } The continuity assumption is also utilized to establish the identification of the finite-mixture Probit model, where the partial derivative of the choice probability with respect to $Z_j$ plays a key role in our identification strategy (see Appendix (ref)).
The conditional independence in Assumption (ref)(i) excludes some types of endogenous group formation by imposing that group membership $s$ does not affect the conditional distribution of $\epsilon^D_j$ given $X$.\footnote{ Assumption (ref)(i) differs from the no-additive-interaction assumption of unobserved confounders for identifying the ATE in the IV estimation. The assumption sates that, in our treatment choice context,
for $\bm{z}, \bm{z}' \in \text{supp}[\bm{Z} | X]$, as in Assumption 5(a) of wang2018bounded. Under Assumptions (ref)(i) and (ref), we observe that $\mathbb E[ D | \bm{Z} = \bm{z}, X, s = j] = \Pr[ \mu_j^D(z_j) \ge \epsilon_j^D | X ]$ and $\mathbb E[ D | \bm{Z} = \bm{z}, X] = \sum_{h=1}^S \Pr[ \mu_h^D(z_h) \ge \epsilon_h^D | X ] \cdot \pi_h$, which implies that (ref) does not hold unless $S = 1$. As a result, their identification results do not recover the ATE in our situation. This is because $s$ interacts nonlinearly with $\bm{Z}$ in the finite mixture treatment choice equation (ref). In addition, note that the no-additive-interaction assumption for the outcome equation (Assumption 5(b) of wang2018bounded) is not satisfied in our situation, as $s$ can be arbitrarily correlated with $Y^{(1)} - Y^{(0)}$. } In other words, the expected “potential” treatment status when in group $j$ does not vary with the actual group membership $s$. More specifically, in the above-mentioned example of college enrollment with multiple eligibility rules, this assumption can be interpreted as that, conditional on the individual characteristics $X$, the probability of enrolling in a college for each group is independent of the admission process the individuals actually take. Thus, if unobserved confounders affect both group membership and treatment choice probability (e.g., unobservable individual talent), Assumption (ref)(i) may not hold. As will be shown in Subsection (ref) (see also Appendix (ref)), the assumption can be dropped at the cost of additional analytical complexity, and we can still establish some identification results of MTE parameters. In Assumption (ref)(ii), we assume homogeneous membership probability for each group, which is commonly used in the literature on finite mixture models. This assumption is made only for simplicity, and the theorem shown below still holds without modifications even when the membership probability is a function of $X$. Moreover, we will show in Subsection (ref) that the group-wise MTE can be identified even when the membership probability depends on other covariates besides $X$.
To interpret these assumptions in a practical scenario, consider our empirical application in Section (ref), in which $Y$ is annual income, $D$ indicates college enrollment, and $X$ includes demographic variables such as age and sex. Suppose that we have two latent groups with different treatment preference patterns: the individuals in group 1 determine whether to go to college or not mainly by their neighborhood's educational characteristics and those in group 2 decide it by the quality of own study environment. Specifically, $Z_1$ includes variables such as local college enrollment rate, and $Z_2$ includes, for example, the number of books in their home when they were 15. Assumption (ref) is satisfied if these IVs are not direct determinants of annual income and if they are also irrelevant to the other unobserved determinants and the preference type conditional on the observed demographic variables. Assumption (ref) requires that the preference type is unrelated to both the demographic variables and the unobserved factors of college enrollment.
A directed acyclic graph (DAG) in Figure (ref) is helpful for further understanding Assumptions (ref) and (ref). The solid and dashed arrows indicate the effects of the observed and unobserved variables, respectively. Assumption (ref)(i) rules out the direct pathways between $Z_j$ and $(\epsilon^{(d)}, \epsilon_j^D, s)$. The arrow from $Z_j$ to $D$ is necessary for Assumption (ref)(ii). Under Assumption (ref)(i), there is no direct pathway between $s$ and $\epsilon_j^D$. In this figure, for simplicity, we suppress $X$, which can be associated with all variables except for $s$ (Assumption (ref)(ii)).
To identify the treatment effects of interest, we first need to identify the treatment choice model (ref). Although there has been a long history of research on the identification of finite mixture models, only few studies address binary outcome regression models (e.g., follmann1991identifiability; butler1997consistency). Moreover, these previous results are typically based on the availability of repeated measurements data, which cannot be directly applied to our situation. Therefore, in Appendix (ref), we present a new identification result for a finite-mixture treatment choice Probit model. Since our major focus is on the identification of treatment effect parameters, the details are omitted here, and we hereinafter treat $P_j$'s and $\pi_j$'s as known objects.
The MTE parameter specific to group $j$ is defined as
where $m_j^{(d)}(x,p) \coloneqq \mathbb E [ Y^{(d)} | X = x, s = j, V_j =p]$ is the marginal treatment response (MTR) function specific to group $j$ and $d \in \{0, 1\}$.\footnote{ Note that we can consider other types of MTE parameters than (ref). For example, an MTE that uses $P_j$ instead of conditioning on $X$ has recently gained attention because of its good computational properties (cf. zhou2019marginal), but it is not pursued in this study considering the page limitation. } This group-wise MTE parameter captures the unobserved heterogeneity in the treatment effects with respect to both $s$ and $V_j$. Similar to the conventional MTE framework, we can infer whether there is individual unobserved treatment selectivity by examining how $\text{MTE}_j(x,p)$ varies with $p$. Specifically, if $\text{MTE}_j(x,p)$ is declining in $p$, it indicates that individuals who belong to group $j$ and are more likely to choose $D = 1$ tend to obtain larger gains from the treatment. The case of no unobserved heterogeneity corresponds to a constant MTE function (cf. carneiro2011estimating,mogstad2018identification). Importantly, our MTE parameter allows us to examine this group-wisely.
Once the group-wise MTEs are identified for all $x$ and $p$, we can recover many other treatment parameters. For example, $\text{CATE}_j(x) = \int_0^1 \text{MTE}_j(x,v)\mathrm{d}v$, where $\text{CATE}_j(x) = \mathbb E [ Y^{(1)} - Y^{(0)} | X = x, s = j]$ is the group-wise conditional average treatment effect. Furthermore, the group-wise ATE: $\text{ATE}_j = \mathbb E [ Y^{(1)} - Y^{(0)} | s = j]$ can be obtained by $\text{ATE}_j = \int \text{CATE}_j(x)f_X(x)\mathrm{d}x$, where $f_X$ is the marginal density function of $X$ (note that Assumption (ref)(ii) implies $f_X(\cdot | s = j) = f_X(\cdot)$). Then, the ATE for the whole population is simply given by $\text{ATE} = \sum_{j=1}^S \pi_j \text{ATE}_j$. We can also identify the so-called policy relevant treatment effect (PRTE); see Remark (ref) below and Appendix (ref).
Below, we show that the MTR functions can be identified through the partial derivatives of the following functions:
where $\mathbf P = (P_1, \ldots, P_S)$ and $\mathbf p = (p_1, \ldots, p_S)$. Note that $\psi_d(x, \mathbf{p})$ can be directly identified from the data on $\text{supp}[X, \mathbf P | D = d]$, where $\text{supp}[X, \mathbf P | D = d]$ denotes the joint support of $(X, \mathbf P)$ given $D = d$.
As shown in the proof of Theorem (ref), we can find the equality $\psi_1(x, \mathbf{p}) = \sum_{j = 1}^S \pi_j \int_0^{p_j}m_j^{(1)}(x,v) \mathrm{d} v$. Then, the first result is straightforward by partially differentiating both sides of this with respect to $p_j$. The second result can be obtained similarly. Here, Assumption (ref)(ii) plays an essential role in enabling us to shift $P_j$ only with the other components of $\mathbf{P}$ being fixed. Notice that the results in Theorem (ref) give us a testable implication for the validity of Assumption (ref)(ii), as kindly pointed out by a referee. That is, for $\mathbf{p}$ and $\mathbf{p}'$, where $\mathbf{p}'$ differs from $\mathbf{p}$ only in terms of group $k$'s specific IV, the equality $\partial \psi_d(x, \mathbf{p})/\partial p_j = \partial \psi_d(x, \mathbf{p}')/ \partial p_j$ may not hold for $j \neq k$ without this assumption. However, to verify this equality from data, we need to estimate the partial derivatives of $\psi_d$ nonparametrically, which should be difficult in practice because of the curse of dimensionality (and even if the equality is confirmed, it only gives a necessary condition for Assumption (ref)(ii)). Note also that since our proposed MTE estimator is defined directly based on Assumption (ref)(ii), it does not contain any information to verify $\partial \psi_d(x, \mathbf{p})/\partial p_j = \partial \psi_d(x, \mathbf{p}')/ \partial p_j$. To assess the impact of the violation of Assumption (ref)(ii), we conducted a small numerical analysis. This demonstrates that our MTE estimator exhibits poor performance in the absence of group-specific IVs, which is consistent with our theory. For more details, see Appendix (ref).
This section considers the estimation of the group-wise MTE parameters using an independent and identically distributed (IID) sample $\{(Y_i, D_i, X_i, \mathbf Z_i): 1 \le i \le n\}$. Throughout this section, Assumptions (ref) and (ref) are assumed to hold.
\paragraph{First-stage: ML estimation} We consider the following parametric model specification:
We assume that $\epsilon_j^D$ is independent of $(X, \mathbf Z)$, and the CDF $F_j$ of $\epsilon_j^D$ is a known function such as the standard normal CDF. Define $\gamma = \left(\gamma_1^\top, \ldots, \gamma_S^\top \right)^\top$ and $\pi = (\pi_1, \ldots, \pi_S)^\top$. Then, the conditional likelihood function for an observation $i$ when $s_i = j$ is given by $L_i(\gamma | s_i = j) \coloneqq F_j( Z_{ji}^\top \gamma_j )^{D_i} [ 1 - F_j( Z_{ji}^\top \gamma_j )]^{1 - D_i}$. Thus, the ML estimator for $\gamma$ and $\pi$ can be obtained by
where $\Gamma \subset \mathbb{R}^{\sum_{j = 1}^S \dim(Z_j)}$ and $\mathcal{C}_S \coloneqq \{\widetilde \pi \in (0, 1)^S: \sum_{j = 1}^S \widetilde \pi_j = 1\}$ are the parameter spaces. We then obtain the estimator of $P_j = F_j( Z_j^\top \gamma_{j} )$ as $\widehat P_j = F_j ( Z_j^\top \widehat \gamma_{n, j} )$. In the numerical studies below, we use the expectation-maximization (EM) algorithm to solve this ML problem following the literature on finite mixture models (e.g., dempster1977maximum; mclachlan2004finite; train2008algorithms).
\paragraph{Second-stage: series estimation}
For the potential outcome equation, we assume the following linear model for convenience, which is a popular setup in the literature (see, e.g., carneiro2009estimating):
Here, the coefficient $\beta_j^{(d)}$ may depend on the group membership $j$ to account for potentially heterogeneous effects of $X$. Assume that $X$ is conditional-mean independent of $(\epsilon^{(d)}, \epsilon_j^D, s)$ for all $d$ and $j$. Then, we have
with $m_j^{(d)}(x,p) = x^\top \beta_j^{(d)} + \mathbb E [\epsilon^{(d)} | s = j, V_j = p]$. By the same argument as in the proof of Theorem (ref), we can show that there exist univariate functions $g_j^{(1)}$ and $g_j^{(0)}$ for each $j$ satisfying
Then, letting $\nabla g_j^{(d)}(p) \coloneqq \partial g_j^{(d)}(p) / \partial p$, we observe that $\nabla g_j^{(1)}(p) = \mathbb E [\epsilon^{(1)} | s = j, V_j = p]$ and $\nabla g_j^{(0)}(p) = - \mathbb E [\epsilon^{(0)} | s = j, V_j = p]$. Hence, it follows that
Further, by the law of iterated expectations,
Hence, we obtain the following partially linear additive regression models:
where $\mathbb E [\varepsilon^{(d)} | X, \mathbf P] = 0$ by the definition of $\psi_d$ for both $d \in \{0, 1\}$. To estimate the coefficients $\beta_j^{(d)}$'s and the functions $g_j^{(d)}$'s, we use the series (sieve) approximation method such that $g_j^{(d)}(p) \approx b_K(p)^\top \alpha_j^{(d)}$ with a $K \times 1$ vector of basis functions $b_K(p) = (b_{1K}(p), \dots, b_{KK}(p))^\top$ and corresponding coefficients $\alpha_j^{(d)}$.
To proceed, letting $d_{SXK} \coloneqq S(\text{dim}(X) + K)$, define $\underset{d_{SXK}\times 1}{\theta^{(d)}} \coloneqq (\beta_1^{(d)\top}, \ldots , \beta_S^{(d)\top}, \alpha_1^{(d)\top}, \ldots , \alpha_S^{(d)\top} )^\top$ for $d \in \{0, 1\}$, and
Then, we can approximate the regression models (ref) and (ref), respectively, by
which implies that we can estimate $\theta^{(d)}$ by
where $A^{-}$ is a generalized inverse of a matrix $A$. Note, however, that $\widetilde \theta_n^{(d)}$'s are infeasible since $P_j$'s and $\pi_j$'s are unknown in practice. Then, define $\widehat R_K^{(d)}$ analogously as above but replacing $P_j$'s and $\pi_j$'s with their estimators obtained in the first step. The feasible estimators can be obtained by
The feasible estimator of $g_j^{(d)}(p)$ is given by $\widehat g_j^{(d)}(p) = b_K(p)^\top \widehat \alpha_{n,j}^{(d)}$, and the infeasible estimator is given by $\widetilde g_j^{(d)}(p) = b_K(p)^\top \widetilde \alpha_{n,j}^{(d)}$. Letting $\nabla b_K(p) \coloneqq \partial b_K(p)/\partial p$, we can also estimate $\nabla g_j^{(d)}(p)$ by $\nabla \widehat g_j^{(d)}(p) \coloneqq \nabla b_K(p)^\top \widehat \alpha_{n,j}^{(d)}$ and $\nabla \widetilde g_j^{(d)}(p) \coloneqq \nabla b_K(p)^\top \widetilde \alpha_{n,j}^{(d)}$. Thus, we obtain the estimators of the MTR functions as follows:
Finally, the feasible estimator of the MTE is given by $\widehat{\text{MTE}}_j(x,p) = \widehat m_j^{(1)}(x, p) - \widehat m_j^{(0)}(x, p)$, and the infeasible estimator is given by $\widetilde{\text{MTE}}_j(x,p) = \widetilde m_j^{(1)}(x, p) - \widetilde m_j^{(0)}(x, p)$.
This section presents the asymptotic properties of the proposed MTE estimators for the model given by equations (ref) and (ref). In the following, for a vector or matrix $A$, we denote its Frobenius norm as $\| A \| = \sqrt{\text{tr}\{ A^\top A \}}$ where $\text{tr}\{ \cdot \}$ is the trace. For a square matrix $A$, $\lambda_{\text{min}}(A)$ and $\lambda_{\text{max}}(A)$ denote the smallest and the largest eigenvalues of $A$, respectively.
The IID sampling condition in Assumption (ref) is standard in the literature. For Assumption (ref)(iii), if the parameters $\gamma$ and $\pi$ are identifiable, the $\sqrt{n}$-consistency of the ML estimators is straightforward; see Appendix (ref) for a special case where $\epsilon_j^D$'s are distributed as the standard normal. Note that combining Assumptions (ref)(i)-(iii) imply $\max_{1 \le i \le n} |\widehat P_{ji} - P_{ji}| = O_P(n^{-1/2})$ for all $j$. In Assumption (ref), we assume the (conditional-mean) independence between the observables and the unobservables.
Let $\nabla^a g_j(p) \coloneqq \partial^a g_j(p) / (\partial p)^a$ and $\nabla^a b_K(p) \coloneqq (\partial^a b_{1K}(p) / (\partial p)^a, \dots, \partial^a b_{KK}(p) / (\partial p)^a)^\top$ for a non-negative integer $a$. Further, define $\Psi_K^{(d)} \coloneqq \mathbb E \left[ R_K^{(d)}R_K^{(d)\top} \right]$ and $\Sigma_K^{(d)} \coloneqq \mathbb E \left[ (\xi_K^{(d)})^2 R_K^{(d)}R_K^{(d)\top} \right]$, where $\xi_K^{(d)} \coloneqq e^{(d)} + B_K^{(d)}$, and $e^{(d)}$ and $B_K^{(d)}$ are unobserved error terms in the series regressions whose definitions are given in (ref) in Appendix (ref).
Assumption (ref)(i) imposes smoothness conditions on the function $g_j^{(d)}$. These conditions are standard in the literature on series approximation methods. For example, Lemma 2 in holland2017penalized shows that Assumption (ref)(i) is satisfied by B-splines of order $k$ for $k - 2 \ge r$ such that $\mu_0 = r$ and $\mu_1 = r - 1$. This assumption requires $K$ to increase to infinity for unbiased estimation while Assumption (ref)(ii) requires that $K$ should not diverge too quickly. It is well known that $\zeta_l(K) = O(K^{(1/2) + l})$ for B-splines (e.g., newey1997convergence). Thus, when B-spline basis functions are employed, Assumption (ref)(ii) is satisfied if $K^5 / n \to 0$. Assumption (ref) ensures that the matrices $\Psi_K^{(d)}$ and $\Sigma_K^{(d)}$ are positive definite for all $K$ so that their inverses exist. Assumption (ref) is used to derive the asymptotic normality of our estimator in a convenient way.
Finally, we introduce the selection matrices $\underset{\dim(X) \times d_{SXK}}{\mathbb{S}_{X, j}}$ and $\underset{K \times d_{SXK}}{\mathbb{S}_{K, j}}$ such that $\mathbb{S}_{X, j} \theta^{(d)} = \beta_j^{(d)}$ and $\mathbb{S}_{K, j} \theta^{(d)} = \alpha_j^{(d)}$ for each $d$ and $j$. The next theorem gives the asymptotic normality for the infeasible estimator.
As shown in Lemma (ref), $\sigma_{K,j}^{(d)}(p)$ corresponds to the asymptotic standard deviation of the MTR estimator for $D = d$. Further, $cov_{K}(p)$ is the asymptotic covariance between the MTR estimators for $D = 1$ and $D = 0$, which is supposed to be non-zero in our case. This non-zero covariance originates from replacing the unobserved membership indicator $\mathbf{1} \{ s = j \}$ with the membership probability $\pi_j$.
As a corollary of Theorem (ref), the asymptotic properties of the feasible MTE estimator can be derived relatively easily with the following additional assumption.
chen2018optimal show that this assumption holds true for B-splines and wavelets.
To proceed, let $\mathcal{C}(\mathcal{D})$ be the set of uniformly bounded continuous functions defined on $\mathcal{D}$. Further, let $T$ be a generic random vector where $\text{supp}[T]$ is compact and $\dim(T)$ is finite. We define the linear operator $\mathcal{P}_{n,K,j}^{(d)}$ that maps a given function $q \in \mathcal{C}(\text{supp}[T])$ to the sieve space defined by $b_K$ as follows:
where $\Psi_{nK}^{(d)} \coloneqq n^{-1} \sum_{i=1}^n R_{i, K}^{(d)} R_{i, K}^{(d) \top}$. The functional form of $q$ may implicitly depend on $n$. The operator norm of $\mathcal{P}_{n,K,j}^{(d)}$ is defined as
which is typically of order $O_P(1)$ for splines and wavelets under some regularity conditions.\footnote{ A set of easy-to-check conditions ensuring this is to verify the conditions in Lemma 7.1 in chen2015optimal and show that $\mathbb{S}_{K, j} \left[ \Psi_{nK}^{(d)} \right]^{-1}$ is stochastically bounded in the $\ell$-infinity norm. See also Appendix B.5 of hoshino2021treatment. }
As shown above, the feasible MTE estimator has the same asymptotic distribution as the infeasible estimator. This is because the asymptotic distributions of the MTE estimators are determined only by the second-stage series estimator of the nonparametric component $\nabla g_j^{(d)}(p)$ whose convergence rate is slower than the first-stage ML estimator and the series estimator of the parametric component $\beta_j^{(d)}$.
The standard errors of the MTE estimators can be straightforwardly computed by replacing unknown terms in $\sigma_{K,j}^{(d)}(p)$ and $cov_K(p)$ with their empirical counterparts: $\widehat \Psi_{nK}^{(d)} \coloneqq \frac{1}{n} \sum_{i=1}^n \widehat R_{i, K}^{(d)} \widehat R_{i, K}^{(d) \top}$, $\widehat \Sigma_{nK}^{(d)} \coloneqq \frac{1}{n} \sum_{i=1}^n (\widehat \xi_{i, K}^{(d)})^2 \widehat R_{i, K}^{(d)} \widehat R_{i, K}^{(d) \top}$, and $\widehat C_{nK} \coloneqq \frac{1}{n} \sum_{i=1}^n \widehat \xi_{i, K}^{(0)} \widehat \xi_{i, K}^{(1)} \widehat R_{i, K}^{(0)} \widehat R_{i, K}^{(1) \top}$, where $\widehat \xi_{i, K}^{(1)} \coloneqq D_i Y_i - \widehat R_{i, K}^{(1)\top} \widehat \theta_n^{(1)}$ and $\widehat \xi_{i, K}^{(0)} \coloneqq (1 - D_i) Y_i - \widehat R_{i, K}^{(0)\top} \widehat \theta_n^{(0)}$.
We consider relaxing Assumption (ref) by allowing the membership variable $s$ to be dependent of both $\epsilon_j^D$ and a vector of covariates $W$. For simplicity, we focus on the case of $S = 2$ only. Suppose that $s$ and $D$ are determined by the following model:
where $U$ is an unobserved continuous random variable distributed as $\text{Uniform}[0,1]$ conditional on $X$, and $\pi$ is an unknown function that takes values on $[0,1]$. The other components are the same as those in (ref). We assume that $W$ is conditionally independent of $U$ given $X$. Then, $\Pr(s = 1|X, W) = \pi(W)$. In addition, we assume that $W$ contains at least one continuous variable that is not included in $(X, Z_1, Z_2)$.
In this setup, we re-define the MTE parameter and the MTR function, respectively, as follows:
Further, letting $Q \coloneqq \pi(W)$, we define
For identification of these functions, we first need to establish the identification of all model parameters in (ref). In Appendix (ref), we present a supplementary identification result for an endogenous finite-mixture treatment choice model where the error terms are assumed to be jointly normal. For notational simplicity, we denote the cross-partial derivatives with respect to $(q, p_j)$ by $\nabla_{qp_j}$; for instance, $\nabla_{qp_1} \psi_1(x, q, p_1, p_2) = \partial^2 \psi_1(x, q, p_1, p_2) / (\partial q \partial p_1)$.
We introduce the following identification conditions that can be visually understood using the DAG in Figure (ref).
The following theorem shows that the group-wise MTE parameter $\text{MTE}_j(x, p)$ defined in (ref) can be recovered by the weighted average of $\text{MTE}_j(x, q, p)$ in (ref).
In some empirical situations, only discrete IVs are available, and many are binary. In this subsection, we focus on a situation where only binary IVs are available and show that some LATE parameters are still identifiable by the Wald estimand. For expositional simplicity, the condition $X = x$ is suppressed throughout this subsection.
Let $Z$ be an $S \times 1$ vector of binary instruments $Z = (Z_1, \dots, Z_S) \in \mathcal{Z}$, where $\mathcal{Z} \coloneqq \text{supp}[Z]$. In this subsection, we allow $Z$ to contain some overlapping elements. In an extreme case, when there is only one instrument common for all groups, we have $\mathcal{Z} = \{ \mathbf{0}_S, \mathbf{1}_S \}$, where $\mathbf{0}_S$ and $\mathbf{1}_S$ are $S \times 1$ vectors of zeros and ones, respectively. Let $Y^{(d, z)}$ be the potential outcome when $D = d$ and $Z = z$. Similarly, denote the potential treatment status when $s = j$ and $Z = z$ as $D_j^{(z)}$.
Assumption (ref)(i) can be violated when the instruments affect the outcome directly or when the group-$j'$-specific instrument $Z_{j'}$ affects the treatment status of the individuals in group $j$ ($j \neq j'$). Note that, under this condition, it holds that $D = D_j^{(z_j)}$ conditional on $s = j$ and $Z = z$. Assumption (ref)(ii) is essentially the same as Assumption (ref)(i). Assumption (ref)(iii) is similar to the monotonicity condition in imbens1994identification. Following the literature, the individuals with $D_j^{(1)} = D_j^{(0)} = 1$ can be referred to as $Z_j$-always-takers; those with $D_j^{(1)} = D_j^{(0)} = 0$ are $Z_j$-never-takers; those with $D_j^{(1)} > D_j^{(0)}$ are $Z_j$-compliers; and those with $D_j^{(1)} < D_j^{(0)}$ are $Z_j$-defiers. Hence, the condition (iii) ensures that for all $j$, there are $Z_j$-compliers but no $Z_j$-defiers in group $j$.
We conduct a set of Monte Carlo experiments to evaluate the finite sample performance of our estimators. Here, we provide the main simulation results only, and some supplementary experiments are provided in Appendix (ref).
Setting $S = 2$ with the membership probabilities $(\pi_1, \pi_2) = (0.6, 0.4)$, the treatment variable is generated by $D = \mathbf{1}\{Z_j^\top \gamma_j \ge \epsilon_j^D\}$ for $s = j$, where $Z_j = (1, X_1, \zeta_j)^\top$, $X_1 \sim N(0, 1)$, $\zeta_j \sim N(0, 1)$, and $\epsilon_j^D \sim N(0, 1)$ for both $j \in \{ 1, 2 \}$. We set $\gamma_1 = (\gamma_{11}, \gamma_{12}, \gamma_{13})^\top = (0, -0.5, 0.5)^\top$ and $\gamma_2 = (\gamma_{21}, \gamma_{22}, \gamma_{23})^\top = (0, 0.5, -0.5)^\top$. The potential outcomes are generated by $Y^{(d)} = \sum_{j \in \{1,2\}} \mathbf{1}\{s = j\} X^\top \beta_j^{(d)} + \epsilon^{(d)}$, where $X = (1, X_1)^\top$, $\epsilon^{(0)} = \sum_{j \in \{1,2\}}\mathbf{1}\{s = j\}V_j + \eta^{(0)}$, $\epsilon^{(1)} = \sum_{j \in \{1,2\}}\mathbf{1}\{s = j\} V_j^2 + \eta^{(1)}$, and $\eta^{(d)} \sim N(0, 0.5^2)$ for both $d \in \{ 0, 1 \}$. Here, $V_j = \Phi(\epsilon_j^D)$, and $\Phi$ denotes the standard normal CDF. The coefficients are set to $\beta_1^{(0)} = (-1, 1)^\top$, $\beta_2^{(0)} = (1, 2)^\top$, $\beta_1^{(1)} = (1, -1)^\top$, and $\beta_2^{(1)} = (2, 1)^\top$. For each setup, we considered two sample sizes, $n \in \{ 1000, 4000 \}$.
For the first-stage estimation of the finite mixture Probit model, we use the EM algorithm. The second-stage MTE estimation is carried out using both infeasible and feasible estimators. We employ B-splines of order $3$ for the basis functions. The number of inner knots of B-splines, say $\widetilde K$, is set to $\widetilde K = 1$ when $n = 1000$ and $\widetilde K \in \{1,2\}$ when $n = 4000$. To stabilize the series regression, ridge regression with penalty $10 n^{-1}$ is also considered for comparison. The simulation results reported below are based on 1,000 Monte Carlo replications.
Table (ref)(a) shows the bias and root mean squared error (RMSE) of estimating the group-wise MTE for both groups with $x = 0.5$ and $v \in \{0.2, 0.4, 0.6, 0.8\}$ (labelled respectively as MTE1.1, and MTE1.2, and so on for group 1, and similarly labelled for group 2). Overall, the performance of the feasible estimator is almost the same as that of the infeasible estimator. This is consistent with our asymptotic theory. The bias of our estimator is satisfactorily small, except for cases close to the boundary. The RMSE quickly decreases as the sample size increases, as long as the number of basis terms remain unchanged, as expected. Theoretically, we need to employ a larger number of basis terms as the sample size increases, however using $\widetilde K = 2$ seems too flexible and the increase in the variance is rather problematic even for $n=4000$ (of course, this result is more or less specific to our choice of the functional form for MTE). The RMSE for group 1 tends to be smaller than that for group 2, probably because of the difference in their group sizes. Introducing a penalty term can improve the overall RMSE; hence, we recommend employing ridge regression in practice with a moderate sample size.
Table (ref)(b) presents the simulation results for the ML estimation of the finite mixture Probit model. For all estimation parameters, the bias is satisfactorily small, especially when $n = 4000$. The RMSE is approximately halved when the sample size increases from $1000$ to $4000$, implying $\sqrt{n}$-consistency of our ML estimator.
In this empirical analysis, we investigate the effects of college education on income in the Japanese labor market. There are two sources used for data collection. The primary data source is the Japanese Life Course Panel Survey 2007 (wave 1),\footnote{ Acknowledgement: The data for this secondary analysis, "Japanese Life Course Panel Surveys, wave 1, 2007, of the Institute of Social Science, The University of Tokyo," was provided by the Social Science Japan Data Archive, Center for Social Research and Data Archives, Institute of Social Science, The University of Tokyo. } which includes detailed information on Japanese workers aged 20 to 40, including their working condition, annual income, education level, and family member characteristics. The outcome variable $Y$ of interest is the respondent's annual income, and the treatment variable is defined as follows: $D = 1$ if the respondent has a college degree (including junior colleges and technical colleges) or higher, $D = 0$ otherwise. The second data source is the School Basic Survey conducted by the Ministry of Education, Japan. Using this dataset, we collect information on regional university enrollment statistics for each prefecture where the respondents were living at the age of 15 to create IVs for the treatment. Table (ref) below shows the list of variables used in this analysis. After excluding observations with missing data and those with zero income, the analysis is performed on 3,318 individuals.
In this analysis, we employ a finite mixture Probit model with two latent groups in which group 1 is composed of individuals whose college enrollment decisions are mainly affected by the regional educational characteristics, and group 2 comprises individuals who are mainly affected by their home educational environment. For the IVs that should be included in $Z_1$ and $Z_2$, we identify them by running a stepwise AIC-based variable selection. In doing this, we assume that $Z_1$ contains at least cap, runiv, and rwork and that $Z_2$ contains at least homeenv and books. Note that we do not rule out the possibility that these variables are common IVs (in fact, cap is identified as a common IV).\footnote{ We estimated the models with $S = 3$ or larger using the same AIC-based model selection procedure. However, the EM algorithm for such models did not converge within tolerable parameter values (the results are not reported here due to word limit constraints, but are available on request). This would indicate that the two-group model is appropriate for our dataset. }
In Table (ref), we present the result of our mixture model and that of the standard Probit model without mixture for comparison. First of all, our treatment choice model has a significantly smaller AIC value than the model with no mixture. The estimation results indicate that approximately 80% of our observations are classified into group 1, which is composed of individuals who are influenced not only by regional characteristics of university enrollment, but also by their family characteristics such as their parents' education level and the number of siblings. The remaining 20% belong to group 2, which is composed of individuals whose home study environment is major determinant of the college entrance decisions.
Following the suggestions from the Monte Carlo results in Appendix (ref), we employ the third-order B-splines with $\widetilde K = 1$ and use the ridge regression with the penalty equal to $10n^{-1}$ for the estimation of the MTEs.\footnote{ Note that introducing a sufficiently small penalty term does not alter our asymptotic results. }
Figure (ref) shows our main results. We find that for those who belong to group 1, the effect of a college degree on income is not significant if they have higher unobserved costs of pursuing higher education, while it is significantly positive if the cost is moderate. This result indicates the presence of a certain treatment selectivity for returns in group 1. Recall that regional educational characteristics and family characteristics are the main factors affecting the college enrollment status of the members in this group. Thus, we can expect that their personal willingness or resistance toward higher education is somewhat heterogeneous within the group, depending on their surrounding environment. Such heterogeneity may also affect the magnitude of the educational returns, leading to the nonlinear shape of the MTE curve for group 1, as depicted in the figure. In contrast, for the members in group 2, those who are influenced by their own study environment, the MTE curve is relatively flat and moderately positively significant in almost the entire region of $p$. This implies that their unobserved costs are less dramatically related to educational returns than those in group 1.
This paper considered identification and estimation of MTE when the data is composed of a mixture of latent groups. We developed a general treatment effect model with unobserved group heterogeneity by extending the Rubin's causal model to finite mixture models. We proved that the MTE for each latent group can be separately identified under the availability of group-specific continuous IVs. Based on our constructive identification result, we proposed the two-step semiparametric procedure for estimating the group-wise MTEs and established its asymptotic properties. An empirical application to the estimation of economic returns to college education indicates the usefulness of the proposed model.
The authors thank the editor, two anonymous referees, Toru Kitagawa, Yasushi Kondo, Ryo Okui, Myung Hwan Seo, Katsumi Shimotsu, Hisatoshi Tanaka, Yuta Toyama, and seminar participants at Hitotsubashi University, Seoul National University, and Waseda University for their valuable comments. This work was supported by JSPS KAKENHI Grant Numbers 15K17039, 19H01473, and 20K01597.