Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
44,479 characters · 6 sections · 81 citation commands
Informational Content of Factor Structures in Simultaneous Binary Response Models
{
} {\bf Keywords:} Factor Structures, Discrete Choice, Causal Effects.\\
{\bf JEL:} C14, C31, C35.
\setcounter{footnote}{0}
\setcounter{equation}{0}
Factor models see widespread and increasing use in various areas of econometrics. This type of structure has been employed in a variety of settings in cross sectional, panel and time series models, and have proven to be a flexible way to model the behavior of and relationship between unobserved components of econometric models. The basic idea behind factor models is to assume that the dependence across the unobservables is generated by a low-dimensional set of mutually independent random factors. The applied and theoretical research employing factor structures in econometrics is extensive. In particular, these models are often used in the treatment effect literature as a way to identify the joint distribution of potential outcomes from the marginal distributions, and then recover the distribution of treatment effects from this joint distribution.\footnote{See also AH07Handbook for an extensive discussion of factor structures and prior studies using these models in the context of treatment effect estimation.} Factor models have been used in a number of different contexts in applied microeconomics. These include, among others, earnings dynamics AC89,BR10, estimation of returns to schooling and work experiences ahmr17, as well as cognitive and non-cognitive skill production technology CHS10. hv071,hv072 provide various additional references. All of these papers, with the notable exception of CHS10, rely on linear factor models where the unobservables are assumed to be written as the sum of a linear combination of mutually independent factors and an idiosyncratic shock.
In this paper we bring together the literature on factor models with the literature on the identification and estimation of binary response models (kleinspady93,lewbel_1,ParkPhillips00,blundellpowell), in particular triangular binary choice models (chesher3,vytlacilyildiz,shaikhvytlacil11,HV13), by exploring the {\em informational content} of factor structures in this class of models.\footnote{See also recent work by LSZ20, who study the identification of a triangular linear model assuming that the disturbances are related through a factor model.} Focusing on this class can be well motivated from both an empirical and theoretical perspective. From the former, many treatment effect models fit into this framework as treatment is typically a binary and endogenous variable in the system, whose effect on outcomes is often a parameter the econometrician wishes to conduct inference on. From a theoretical perspective, inference on this type of system can be complicated, if not impossible without strong parametric assumptions, which may not be reflected in the observed data. Imposing no restriction on the structure of endogeneity often fails to achieve identification of parameter, or at best only do so in sparse regions of the data, thus making inference impractical in practice. In this context, modeling the endogeneity between the selection and the outcome by a factor structure may be a useful “in-between” setting, which, at the very least, can be used to gauge the sensitivity of the parametric approach to their stringent assumptions.
We start our analysis by imposing a particular factor structure to the two unobservables in our system of binary equations described in further detail in the next section, and explore the informational content of this assumption. We assume that the unobservables from the treatment equation ($V$) and the outcome equation ($U$) are related through the following factor model:
where $\Pi$ is an unobserved random variable assumed to be distributed independently of $V$ and $\gamma_0$ is a scalar parameter. This structure generalizes the canonical case where the unobservables $(U,V)$ are jointly normally distributed, for which this relationship always holds. Our main finding is that there is indeed informational content of factor structures in the sense that, in contrast to prior literature - notably vytlacilyildiz - one no longer requires an additional “non-standard” exclusion restriction, nor the strong support conditions on the covariates entering the outcome equation that are generally needed for identification in these models. Our identification results are constructive and translate directly into a rank based estimator of the coefficient associated with the binary endogenous variable, which we provide and study in a supplement to this paper.
While an appealing feature of the structure considered in Equation ((ref)) is that it is a natural extension of the bivariate Probit specification that has often been considered in the literature, this model does impose significant restrictions on the nature of the dependence between the unobservables $U$ and $V$. In the paper we extend this baseline specification by considering a linear factor structure of the form:
where $(W, \eta_1,\eta_2)$ are mutually independent unobserved random variables. We study the informational content of this extended factor structure in the context of triangular binary choice models and establish identification, assuming access to at least two continuous noisy measurements of the unobserved factor $W$. This setup has been used in a number of applications, in particular in labor economics. In these applications, the unobserved factor is typically interpreted as latent individual ability, about which several continuous noisy measurements are available from the data. This is the case of, for instance, CHH03, CHS10, HHV18 and ahmr17, who use components of the Armed Services Vocational Aptitude Battery test as measurements of cognitive ability.
The rest of the paper is organized as follows. In Section (ref) we formally describe the triangular system with our factor structure, and discuss our main identification results for the parameters of interest in this model. Section (ref) explores identification in more general factor structure models which involve multiple idiosyncratic errors, in a context where one has access to two continuous noisy measurements of the common unobserved factor. Section (ref) concludes. We prove Theorems (ref) and (ref) in Sections (ref) and (ref), respectively. In Section (ref), we establish the sharp identified set of $\alpha_0$ when the support condition for point-identification is violated in the one-factor model and the necessary and sufficient condition for point-identification in the two-factor model with two continuous measurements of the common factor. The corresponding results are proved in Sections (ref) and (ref). The Supplementary Material studies the asymptotic properties of a rank-based estimator for $\alpha_0$ and explores its finite sample properties through some Monte Carlo simulation exercises.
Notation: throughout the paper we write $\mathbf{1}\{A\}$ to denote the usual indicator function that takes value 1 if event $A$ happens, and 0 otherwise. We also denote by $d(U)$ and $d(U|V)$ the lengths of the support of random variable $U$, and the conditional support of $U$ given $V$, respectively.
In this section we consider the identification of the following triangular binary model:
where $Z\equiv(Z_1,Z_2)$ and $(U,V)$ is a pair of random shocks. $Z_2$ and $Z_3$ provide the exclusion restrictions in the model, and the distribution of $(Z_2,Z_3)$ is required to be nondegenerate conditional on $Z_1'\lambda_0+Z_3'\beta_0$. We further assume that the error terms $U$ and $V$ are jointly independent of $(Z_1,Z_2,Z_3)$. The endogeneity of $Y_2$ in ((ref)) arises when $U$ and $V$ are not independent.
The above model, or minor variations of it, have often been considered in the recent literature. See for example, vytlacilyildiz, abhausmankhan, vellaetal, vuongxu, khannekipelov2 and references therein. A key parameter of interest\footnote{As is always the case in models with binary outcomes, both the interpretation and the usefulness of regression coefficients warrant explanation. In the model considered here the coefficient on the treatment variable and the coefficients on exogenous variables in the binary outcome equation enable us to construct “equivalence classes" to answer important policy questions. For example consider the case where the dummy endogenous variable is job training, the exogenous regressor is years experience and the outcome variable is employment status. Knowing all coefficients would be informative on how many additional years of experience would be needed to compensate for a lack of training so the probability of being employed stays the same.} in our paper as is in much of the literature is $\alpha_0$. In this paper we provide conditions under which the parameters of interest are point-identified. As such, our analysis complements alternative partial-identification approaches that have been proposed in the context of triangular binary models. See, in particular, chiburis10, shaikhvytlacil11, and mourifie15.\footnote{In Section (ref) in the supplement, we establish the sharp identified set of $\alpha_0$ when the support condition for point-identification is violated. This result highlights that, except for the fact that the sign of $\alpha_0$ is identified, we generally cannot say much about the value of $|\alpha_0|$. Related work by shaikhvytlacil11 also provides partial identification results for a triangular binary model. That the bounds for $\alpha_0$ are generally tighter in their analysis reflects the identifying power of the additional support restrictions that they impose.} As discussed in the aforementioned papers, the parameter $\alpha_0$ is difficult, if not impossible to point identify and estimate without imposing parametric restrictions on the unobserved variables in the model, $(U,V)$.
The difficulty of identifying $\alpha_0$ in semi-parametric “distribution-free" models, and the sensitivity of its identification to misspecification in parametric models is what motivates the factor structure we add in this paper to the above model. Specifically, to allow for endogeneity in the form of possible non-zero correlation between $U$ and $V$, we augment the model with the following equation:
where $\Pi$ is an unobserved random variable, assumed to be distributed independently of $(V,Z_1,Z_2,Z_3)$, and $\gamma_0$ is an additional unknown scalar parameter. Importantly, this type of factor structure always holds when the residuals of both equations are jointly normally distributed. Furthermore, this specification corresponds to the type of structure used in Independent Component Analysis (ICA), where $V$ and $\Pi$ are two mutually independent factors. This method has found many applications in various fields, including signal processing and image extraction; applications in economics include e.g., Hyvarinenetal, Monetaetal and Gourierouxetal. While, in contrast to the ICA literature, the factors and the factor loadings are not the main objects of interest in our analysis, this dimension-reducing structure plays a key role in our identification results.
Our aim is to first explore identification of the parameters $(\alpha_0, \delta_0, \gamma_0,\beta_0,\lambda_0)$ under standard nonparametric regularity conditions on $(V,\Pi)$. Note that the parameter $\delta_0$ in the selection equation can be identified up to scale in various ways. See, for example, kleinspady93 and MRC, among others. We then impose the usual condition that one of $\delta_0$'s coordinates is equal to one to fix the scale. For simplicity, for the rest of the paper, we denote $X \equiv Z'\delta_0$ and assume $X$ is observed. We further define $X_1 \equiv Z_1'\lambda_0 + Z_3'\beta_0$. However, we cannot identify $\lambda_0$ and $\beta_0$ beforehand. We propose instead to identify them along with $\alpha_0$.
Our main identification result is based on the Assumptions {\bf A1}-{\bf A4} we state below:
Before turning to our main identification result, a couple of remarks are in order.
We now turn to our main identification result, Theorem (ref), which concludes that under our stated conditions and our factor structure we can attain point identification of the vector of parameters $\theta_0$.
An important takeaway from this result, which we discuss further in Subsection (ref) below, is that imposing the factor structure ($\ref{eq:factor1}$) yields point-identification under weaker support conditions when compared to the existing literature, and does not require a second exclusion restriction either. In particular, our model delivers point-identification of the parameters of interest even in situations where all of the regressors from the outcome equation are discrete. This indicates that, from the selection equation combined with the factor structure that we impose here, we can overturn the non-identification result of bierenshartog which would apply to the outcome equation alone.
The proof of Theorem (ref), which is reported in Section (ref) in the Supplementary Appendix, relies on the fact that, for two observations $(Z_1,Z_3,X)$ and $(\tilde{Z}_1, \tilde{Z}_3, \tilde{X})$,
where $f_V(\cdot)$ is the pdf. of $V$, which is identified over the support of $X$, and $P^{ij}(z_1,z_3,x) \equiv Prob(Y_1 = i,Y_2 = j|Z_1 =z_1,Z_3 = z_3,X=x)$ ($\partial_x P^{ij}(z_1,z_3,x) $) denote the choice probability (partial derivative of the $ij$-choice probability with respect to the third argument), which are both identified from the data.
We now discuss in detail how our setup and main identification result relates to the existing literature.
In a related work, HV13 consider the identification of a generalized bivariate Probit model.\footnote{See also recent work by han_lee who study semiparametric estimation and inference in the framework considered by HV13.} Our linear factor structure and the one-parameter copula model considered in HV13 are not nested by each other. First, note that based on the factor structure, we can recover $F_\Pi$, the distribution of $\Pi$, as a function of $(F_U,F_V,\gamma_0)$ by deconvolution. We can then write the copula of $(U,V)$ as $$F_{U,V}(F_U^{-1}(u),F_V^{-1}(v)) = \int_{-\infty}^{F_V^{-1}(v)}F_\Pi(F_U^{-1}(u) - \gamma_0 w; F_U,F_V,\gamma_0)f_V(w)dw = C(u,v;F_U,F_V,\gamma_0).$$ The copula depends not only on $\gamma_0$ but also on two infinite dimensional parameters $(F_U,F_V)$. Thus, unlike HV13, our factor structure cannot be characterized by a one-parameter copula. In addition, in order to achieve identification, HV13 first nonparametrically identify the two marginals by assuming the existence of a full support regressor that is common to both equations.\footnote{HV13 establish their identification of the coefficient on the endogeneous regressor (Theorems 4.2 and 5.1) under the assumption that the marginal distributions $F_\varepsilon$ and $F_\nu$ are known. Then, they verify this condition by showing the identification of these two marginal distributions using large support common regressors.} In contrast, our approach does not rely on the existence of such a regressor. Under the factor structure assumed in our analysis, we bypass the nonparametric identification of the marginals as a whole and directly consider the identification of the structural parameters. It follows that our model cannot be nested by the one-parameter copula model considered by HV13. On the other hand, there exist one-parameter copula models that cannot be decomposed into linear factor structures.\footnote{For instance, suppose that $(U,V)$ has a Gaussian copula with correlation $\rho$, and that the marginal distributions of $U$ and $V$ are uniform $[0,1]$. It then follows that, denoting by $\Phi(.)$ the standard normal cdf., $\left(\Phi^{-1}(U), \Phi^{-1}(V)\right)$ is bivariate normal with correlation $\rho$, which in turn yields the following non-linear relationship between $U$ and $V$: $U = \Phi\left( \rho \Phi^{-1}(V) + W\right)$, where $W$ is normally distributed and independent from $V$.} This implies that our model does not nest HV13 either.
Our analysis also relates to vytlacilyildiz and vuongxu, who consider the identification of $\alpha_0$ in a triangular binary model. Our identification result, however, differs from theirs in important ways. Namely, denote $X = Z'\delta_0 = Z_1'\delta_{1,0} + Z_2'\delta_{2,0}$. Then, Assumption A4 implies that we can find a pair of observations $(z_1,z_2,z_3)$ and $(\tilde{z}_1,\tilde{z}_2,\tilde{z}_3)$ such that
In contrast, using our notation, vytlacilyildiz require that one can find a pair of observations $(z_1,z_2,z_3)$ and $(\tilde{z}_1,\tilde{z}_2,\tilde{z}_3)$ such that $z'\delta_{0} = \tilde{z}'\delta_0$ and
vuongxu do not assume the existence of $Z_3$. In our binary outcome setup, the functions $h(0,x,\tau)$ and $h(1,x,\tau)$ defined in vuongxu are equal to $1\{x + F_{-U}^{-1}(\tau) \geq 0\}$ and $1\{x + \alpha+F_{-U}^{-1}(\tau) \geq 0\}$, respectively, where $x = z_1'\lambda_0$ and $F_{-U}$ is the CDF of $-U$. Then, vuongxu requires that we can find $z_1$ and $\tilde{z}_1$ in the support of $Z_1$ so that for any $\tau_1,\tau_2$, if $1\{\tilde{z}_1'\lambda_0 + F_{-U}^{-1}(\tau_1) \geq 0\} = 1\{\tilde{z}_1\lambda_0 + F_{-U}^{-1}(\tau_2) \geq 0\}$, then $1\{z_1'\lambda_0 + \alpha_0 + F_{-U}^{-1}(\tau_1) \geq 0\} = 1\{z_1\lambda_0 + \alpha_0 + F_{-U}^{-1}(\tau_2) \geq 0\}$. Provided that the support of $U$ nests the supports of $Z_1'\lambda_0$ and $Z_1'\lambda_0+\alpha_0$, vuongxu is then equivalent to:\footnote{To see this, note that if, say, $z_1'\lambda_0 +\alpha_0 > \tilde{z}_1'\lambda_0$, then we can find $\tau_1,\tau_2$ such that $-z_1'\lambda_0 - \alpha_0 \leq F_{-U}^{-1}(\tau_1) < -\tilde{z}_1'\lambda_0$ and $F_{-U}^{-1}(\tau_2) < -z_1'\lambda - \alpha_0 < -\tilde{z}_1'\lambda_0$. This violates the above requirement, and thus, shows that vuongxu implies (ref). On the other hand, if $z_1'\lambda_0 +\alpha_0 = \tilde{z}_1'\lambda_0$, then vuongxu holds trivially. }
Several remarks are in order. First, note that sufficient support conditions for the restrictions (ref)--(ref) are $d(Z_1'\lambda_0 + Z_3'\beta_0 - Z'\delta_{0}\gamma_0) \geq |\alpha_0|$, $d(Z_1'\lambda_0 + Z_3'\beta_0|Z'\delta_0)\geq |\alpha_0|$, and $d(Z_1'\lambda_0 |Z'\delta_0)\geq |\alpha_0|$ with a positive probability, respectively, where $d(\cdot)$ denotes the “length" of its argument. These three support conditions are such that $$d(Z_1'\lambda_0 + Z_3'\beta_0 - Z'\delta_{0}\gamma_0) \geq d(Z_1'\lambda_0 + Z_3'\beta_0|Z'\delta_0) \geq d(Z_1'\lambda_0|Z'\delta_0),$$ where the first and second inequalities are strict if $Z_2$ and $Z_3$ have at least one continuous component, respectively. Importantly, we show in Section (ref) of the Supplement that for a version of the triangular binary model with univariate $Z_2$ and $Z_3$ and no common regressor $Z_1$, the support condition $d(Z_1'\lambda_0 + Z_3'\beta_0|Z'\delta_0)\geq |\alpha_0|$ is actually also necessary to the identification of the model without factor structure. This implies that by imposing our factor structure, one can identify values of $\alpha_0$ in a region that cannot be identified in the model considered by vytlacilyildiz. Such region is characterized in Section (ref) of the Supplement.
Second, it directly follows from these support conditions that, in the presence of a factor model and in contrast to both vytlacilyildiz and vuongxu, variation in $Z_2$ helps in the identification of $\alpha_0$. In that sense, the factor model allows to restore the intuition from standard IV approaches in linear models that variation in the instrument $Z_2$ is critical to the identification of the parameters of the outcome equation. Related to this, the support of $Z_2$ plays an important role in our identification analysis. In particular, if $Z_2$ is discrete, our identification strategy requires sufficient variation in the variables in the outcome equation, namely $Z_1$ and $Z_3$. In this case, our support requirement is equivalent to that assumed by vytlacilyildiz.
Third, another important aspect of Assumption A4 is that it does not impose any constraint on the variables from the outcome equation. Specifically, consider a case where the outcome equation does not contain a variable that is excluded from the selection equation (i.e., $\beta_0=0$), the regressor that is common to both equations, $Z_1$, is scalar and binary, and where $\lambda_0 = 1$. In this case, one can show that the identifying support conditions associated with vytlacilyildiz (ref) and vuongxu (ref) generally fail to hold, except for a finite set of values $\alpha_0 \in \{-1,0,1\}$. In contrast, our support restriction (ref) holds under more general conditions: without any restriction on $\alpha_0$ if one element of $Z_2$ is continuous with large support, and on a continuum of possible values for $\alpha_0$ if one element of $Z_2$ is continuous with bounded support. In that sense, the factor structure replaces the need for a continuous component in $(Z_1,Z_3)$ in the outcome equation.
Finally, at a high level, our identification strategy shares similarities with the Local Instrumental Variable (LIV) approach that has been proposed by hv05 and further discussed by cl09. In particular, our identifying restriction (ref) can be alternatively derived from a local IV strategy applied to a potential outcomes model characterized by $Y_1(y_2)={\bf 1}\{Z_1'\lambda_0+Z_3'\beta_0+\alpha_0 y_2-U>0\}$, with treatment given by $Y_2= {\bf 1}\{Z'\delta_0-V>0\}$. In contrast to the LIV literature though, we focus in our analysis on the structural parameter $\alpha_0$ rather than on the marginal treatment effects. Our identification result shows that, by leveraging the identifying power of the factor structure, one can identify $\alpha_0$ under weaker support restrictions than in the prior literature. In particular, our strategy makes it possible to use variation in $X = Z'\delta_0$ to identify $\alpha_0$, even when all the components of $Z_1$ and $Z_3$ are discrete.\footnote{An alternative approach to identifying this parameter can be found in lewbel_1. In his approach a second equation to model the endogenous variable is not needed, nor is the factor structure we impose. However, he imposes a strong support condition on a variable like $Z_3$ requiring that it exceeds the length of the unobservable $U$.}
\setcounter{equation}{0} Up until now we have proposed identification and estimation results for a triangular system with a particular factor structure. A disadvantage of this structure is that it only includes one idiosyncratic shock ($\Pi$). We consider below an extension that addresses this limitation.
Namely, we consider the following model:
where $X_1 = Z_1'\lambda_0 + Z_3'\beta_0$, $X = Z'\delta_0$, $U = \gamma_0 W + \eta_1$, $V = W + \eta_2$, and $(W, \eta_1,\eta_2)$ are mutually independent. In this setup, $W$ can be interpreted as an unobserved confounder that satisfies the matching-on-unobservables condition $(Y_1(0),Y_1(1)) \perp\!\!\!\perp Y_2 | W,X,X_1$ AH07Handbook. Recall that, following the arguments in Section (ref) above, we assume that $X$ is observed. In addition, we assume two auxiliary continuous measurements
where $(W, \eta_1,\eta_2,\eta_3,\eta_4)$ are mutually independent, and $\nu_0 \neq 0.$\footnote{In practice, the continuous measurements might also depend on some observable characteristics. Our analysis goes through in this case after residualizing $Y_3$ and $Y_4$.}
Our identification result is based on the following assumptions:
We now discuss these assumptions, before turning to the identification result. First, Assumption B0 is similar to Assumptions A1 and A2. We only need one of the idiosyncratic errors in the continuous measurements to be independent of the covariates because the other one is used to identify the distribution of the common factor $W$ only. Second, as we assume in Assumption B1 that $\gamma_0 \neq 0$ and $X$ has full support, the support condition $$d(Z_1'\lambda_0 + Z_3'\beta_0 - \gamma_0 X) \geq |\alpha_0|.$$ holds automatically. The full support condition of $X$ is necessary to identify the density of $V$, which is further used to identify the distribution of $\eta_2$. Assumption B1 reinforces this condition by supposing that $X$ has full support conditional on $Z_1$ and $Z_3$, which is needed to identify the parameters from the outcome equation in a second step. Since $X = Z'\delta_0$ with $Z=(Z_1,Z_2)$, this is in turn equivalent to $Z_2$ having full support conditional on $Z_1$ and $Z_3$. Third, Assumptions B2--B6 imply Assumptions 1 to 4 in SH13. In practice we add the condition that the characteristic function of $\eta_2$ does not vanish, which is used for the deconvolution arguments in the proof of Theorem (ref). We refer the reader to SH13 for more discussions of these assumptions.\footnote{Note that SH13 hold automatically in our model with $\nu_0 \neq 0.$}
The proof of Theorem (ref) can be found in Section (ref) of the Supplement. Several remarks are in order. First, while we allow for a more general factor structure on the unobservables $U$ and $V$, we also depart from our baseline specification by supposing that we have access to two continuous noisy measurements of the common factor $W$. This is a standard requirement in the nonparametric measurement error literature hs08. Besides, assuming access to a set of (selection-free) noisy measurements of the unobserved factors is also very standard in the evaluation literature. See, among many others, CHH03, HN07, hv071, and CHS10.
For instance, in applications in labor economics, the unobserved factor $W$ often captures individual ability. This would apply, for example, to the evaluation of the effect of employment while in college ($Y_2$) on college graduation ($Y_1$). In this example, natural candidates for $Z_2$ are local labor market variables, including average wages and unemployment rate, while candidates for $Z_1$ include, among others, eligibility to financial aid programs providing tuition subsidy to students who maintain a minimum level of academic achievement.\footnote{See ScottClaytonJHR11 for an evaluation of a program of this kind (PROMISE scholarship in West Virginia), and for a discussion of similar merit-based scholarship programs in place in other states.} In this context, cognitive skill measurements, such as the ASVAB test components that are available in the NLSY79 and NLSY97 surveys, are natural and often used candidates for the continuous measurements $(Y_3,Y_4)$ ahmr17.
Second, as is clear from the proof of Theorem (ref), the key purpose of the continuous measurements is to identify the distribution of the common factor $W$. While we assume in this section that the measurement equations are linear, it is possible to identify $\theta_0$ with a more general nonlinear system of continuous measurements, provided that the researcher has access to at least three such measurements. One can then combine Theorem 2 in CHS10 (Section 3.3, pp. 894-895), that yields identification of the distribution of $W$, with the proof of Theorem (ref) in order to show identification of $\theta_0$ for the case of nonlinear auxiliary measurements. Assuming access to a set of at least three measurements also makes it possible to relax the non-normality requirement imposed in Assumption B2.
Third, under the previous set of assumptions, the average treatment effect (ATE) is also identified. Key to this identification result is the full support condition on $X$ given $Z_1$ and $Z_3$ (Assumption B1). Note that the conditional ATE given $X_1=x_1$ is equal to $F_U(x_1 + \alpha_0) - F_U(x_1)$. In addition, $$P(Y_1 = 1,Y_2 =1|X_1=x_1, X=x) = F_{U,V}(x_1 + \alpha_0, x).$$ One can let $x \rightarrow \infty$ so that $$\lim_{x \rightarrow \infty}P(Y_1 = 1,Y_2 =1|X_1=x_1, X=x) = F_U(x_1+\alpha_0).$$ Similarly, $$\lim_{x \rightarrow -\infty}P(Y_1 = 1,Y_2 =0|X_1=x_1, X=x) = F_U(x_1).$$ This identifies the conditional and unconditional ATE.
Fourth, similar to the earlier discussions in Remark (ref) and Section (ref), Assumption B7 may still hold even when $Z_3$ is an empty set and $Z_1$ is discrete, since $W$ is assumed to have full support. In such a case, identification primarily relies on the factor structure and the variation of the covariates in the selection equation, rather than that in the outcome equation. In this respect, this identification result is similar in spirit to Theorem (ref) and different from the existing identification results in the literature for triangular binary models, e.g., vytlacilyildiz and vuongxu. More generally, in Section (ref) in the supplement we establish that the factor model provides identification restrictions that are not otherwise available.\footnote{Specifically, we consider a version of the model (ref), where we do not impose the factor structure and allow for an arbitrary (unknown to econometricians) dependence structure across the unobservables of the model. In this case, we show non-identification of $\alpha_0$ as long as $|\alpha_0| > b-a$, where $[a,b]$ denotes the conditional support of $X_1$ given $X$ and, consistent with our Assumption B1, $X$ has full support on the real line. However, by imposing the factor structure (and other conditions implied by B0--B7), Theorem (ref) shows that $\alpha_0$ is identified for this model even when $|\alpha_0|>b-a$.}
Finally, we can relax the rank invariance condition to rank similarity by replacing $Y_1 = 1\{X_1 + \alpha_0 Y_2 - U \geq 0\}$ by $Y_1 = 1\{X_1 + \alpha_0 Y_2 - U(Y_2) \geq 0\}$. We then require $U(y_2) = \gamma_0 W + \eta_1(y_2)$ for $y_2 = 0,1$. If Assumptions {\bf B0}--{\bf B7} hold with $\eta_1$ replaced by $(\eta_1(1),\eta_1(0))$ and $P(\eta_1(1) \leq e) = P(\eta_1(0) \leq e)$ for $e \in \Re$, then we can still identify $\theta_0$ by a similar argument as the proof of Theorem (ref).
In this paper, we explore the identifying power of linear factor structures in the context of simultaneous binary response models. We impose two alternative types of factor structures on the unobservables of the model. The first setup is a natural distribution-free extension of the bivariate Probit model, while the second model corresponds to a standard linear factor model with one common factor and two equation-specific idiosyncratic shocks. We establish that both factor models have identifying power in that they make it possible to relax some of the exclusion and support conditions typically required for identification in this class of models (vytlacilyildiz, vytlacilyildiz). Overall, our analysis adds to our understanding of the identifying power of factor models, beyond their well known usefulness to recover the joint distribution of potential outcomes from the marginal distributions.
The work here opens areas for future research. The factor structure we assume could prove useful in more general nonlinear models. For instance, non-triangular discrete systems have shown to be an effective way to model entry games in the empirical industrial organization literature- see, for example, tamer_2. However, as shown in khannekipelov2, identification of structural parameters in these models can be even more challenging than for the triangular model considered in this paper, and furthermore, as shown recently in khannekipelov3, conducting valid uniform interest in all these models is very difficult. It would be useful to determine if factor structures on the unobservables could alleviate this problem. We leave this open question to future work.