EconBase
← Back to paper

Identification and Estimation of Weakly Separable Models Without Monotonicity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

44,222 characters · 7 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Dummy Endogenous Variables in Weakly Separable Multiple Index Models without Monotonicity

center[center omitted — 45 chars of source]

{ We study the identification and estimation of treatment effect parameters in weakly separable models. In their seminal work, \citeasnoun{vytlacilyildiz} showed how to identify and estimate the average treatment effect of a dummy endogenous variable when the outcome is weakly separable in a single index. Their identification result builds on a monotonicity condition with respect to this single index. In comparison, we consider similar weakly separable models with multiple indices, and relax the monotonicity condition for identification. Unlike \citeasnoun{vytlacilyildiz}, we exploit the full information in the distribution of the outcome variable, instead of just its mean. Indeed, when the outcome distribution function is more informative than the mean, our method is applicable to more general settings than theirs; in particular we do not rely on their monotonicity assumption and at the same time we also allow for multiple indices. To illustrate the advantage of our approach, we provide examples of models where our approach can identify parameters of interest whereas existing methods would fail. These examples include models with multiple unobserved disturbance terms such as the Roy model and multinomial choice models with dummy endogenous variables, as well as potential outcome models with endogenous random coefficients. Our method is easy to implement and can be applied to a wide class of models. We establish standard asymptotic properties such as consistency and asymptotic normality.}

{ JEL Classification: C14, C31, C35 \\ Key Words Weak Separability, Treatment Effects, Monotonicity, Endogeneity}

\setcounter{equation}{0}

Introduction

Consider a weakly separable model with a binary endogenous variable:

eqnarray[eqnarray omitted — 108 chars of source]

where $( v_1(X,D),v_2(X,D),...v_J(X,D) ) \equiv v(X,D) $ is a $J$-vector of unknown linear or nonlinear indices in the outcome equation (1.1) and $D$ is a binary endogenous variable defined by the selection equation (1.2). Here $X\in \mathbb{R}^{d_{x}}$ and $Z\in \mathbb{R} ^{d_{z}}$ are vectors of observable exogenous variables, which may have overlapping elements. Similar to \citeasnoun{vytlacilyildiz} we require exclusion restrictions that there is some element in $Z$ excluded from $X$, and that we can vary $X$ after conditioning on $\theta(Z)$. In the system of equations above, $U$ is the unobservable random variable normalized to follow the uniform distribution $U(0,1)$ and the error term $\varepsilon$ in the outcome equation is allowed to be a random vector. We assume $\left( X,Z\right) $ are independent of $(\varepsilon ,U)$. Note that we allow $v(X,D)$ to be a vector of multiple indices, whereas the method in \citeasnoun{vytlacilyildiz} can only be applied when it is a single index.

Since \citeasnoun{vytlacilyildiz}, other important work has considered identification and estimation of related models, but under alternative conditions. Examples with binary endogenous variables include \citeasnoun{hanvytlacil}, \citeasnoun{vuongxu}, \citeasnoun{lewbelqe}, \citeasnoun{khanetal}. Work for models when the endogenous variable is continuous includes \citeasnoun{neweyimbens}, \citeasnoun{dhault2014} and \citeasnoun{torgo2014}. \citeasnoun{fengjun} shows how to identify nonseparable triangular models where the endogenous variable is discrete and has larger support than the instrument variable.\footnote{ All these papers focus on point identification. For partial identification of a model with a binary outcome, see \citeasnoun{shaikhvytlacil11} and \citeasnoun{mourifie15}.}

As in the conventional framework, two potential outcomes $ Y_{1}$ and $Y_{0}$ satisfy

equation*[equation* omitted — 71 chars of source]

We only observe $\left( Y,D,X,Z\right) $, where $Y=DY_{1}+(1-D)Y_{0}$. In this model, as in \citeasnoun{vytlacilyildiz}, we do not impose parametric distribution on the error term or a linear index structure. \citeasnoun{vytlacilyildiz} assumes that $v(X,D) \in \mathbb{R}$ is a single index, and

equation[equation omitted — 122 chars of source]

Unlike \citeasnoun{vytlacilyildiz}, we do not impose any monotonicity structure. That is because our approach exploits the full information in the distribution of the outcome variable, instead of just its mean. Indeed, when the outcome distribution function is more informative than the mean, our method is applicable to more general settings than theirs; in particular we do not rely on their monotonicity assumption and at the same time we also allow for multiple indices.

In Sections (ref) and (ref) we provide some examples in which such a monotonicity condition fails, but the average effect of the binary endogenous variable is still identified. In addition, we allow for a weakly separable model with multiple indices, that is, $v(X,D)=(v_{1}(X,D),v_{2}(X,D),...,v_{J}(X,D)) \in \mathbb{R}^J$.

We consider the identification and estimation of the average treatment effect of $D$ on $Y$, $E(Y_{1}|X\in A)$, $E(Y_{0}|X\in A)$ \ and $E(Y_{1}-Y_{0}|X\in A)$, for some set $A$, without the aforementioned monotonicity. Indeed, for the case with multiple indices $v(X,D) \in \mathbb{R}^J$, the monotonicity condition is no longer well defined.

\citeasnoun{vuongxu} established nonparametric identification of individual treatment effects in a fully nonseparable model that includes a binary endogenous regressor, without the nonlinear index structure. They assume $ \varepsilon $ is a scalar and $g$ is strictly increasing in $\varepsilon $. In their setting, monotonicity in the outcome equation provides the identifying restriction to extrapolate information from local treatment effects to population treatment effects.

\setcounter{equation}{0}

Identification

Generally speaking, our identification strategy will be based on the notion of matching\footnote{ See \citeasnoun{ahn-powell}, \citeasnoun{chenkhantang}, \citeasnoun{vytlacilyildiz}, and more recently \citeasnoun{auerbach} for examples of papers that attain identification through matching.}. Consider the identification of $E(Y_{1}|X=x)$ for some $x\in S_{1}$, where $S_{d}$ denotes the support of $X$ given $D=d\in \{0,1\}$. Note that because $ (\varepsilon ,U)\bot (X,Z)$,

eqnarray[eqnarray omitted — 202 chars of source]

where $P(z)\equiv E(D|Z=z)$. The only term that is not directly identifiable on the right-hand side of ((ref)) is

equation*[equation* omitted — 82 chars of source]

The main idea behind our approach follows that of \citeasnoun{vytlacilyildiz} , which is to find some $\tilde{x}\in S_{0}\ $such that

equation[equation omitted — 50 chars of source]

so that

eqnarray*[eqnarray* omitted — 170 chars of source]

{Unlike \citeasnoun{vytlacilyildiz}, we utilize the full distribution of }$Y$ {\ (rather than its first moment) while searching for such pairs of }$(x, \tilde{x})${\ in ((ref)). This allows us to relax the single-index and monotonicity conditions in \citeasnoun{vytlacilyildiz}}.

For any $p$ on the support of $P(Z)$ given $X=x$, and for all $y$ define

eqnarray[eqnarray omitted — 264 chars of source]

where

equation*[equation* omitted — 89 chars of source]

with $v(x,d)$ being a realized index at $X=x$ and the expectation in the definition of $F_{g|u}$ is with respect to the distribution of $\varepsilon $ given $U=u$. The last equality in ((ref)) holds because of independence between $(\varepsilon ,U)$ and $(X,Z)$. By construction, $h_{1}^{\ast }(x,y,p)$ is directly identified from the joint distribution of $(D,Y,X,Z)$ in the data-generating process. Furthermore, for any pair $p_{1}>p_{2}$, define:

equation*[equation* omitted — 143 chars of source]

Likewise, define

eqnarray*[eqnarray* omitted — 237 chars of source]

and let

equation*[equation* omitted — 143 chars of source]

{Let }$P_{x}${\ denote the support of }$P(Z)${\ given }$X=x${. It can be shown that for any }$x\in S_{1}${\ and }$\tilde{x}\in S_{0}${,} and any $y$,

equation[equation omitted — 158 chars of source]

{if and only if}

equation[equation omitted — 125 chars of source]

{Sufficiency is immediate from the definition of }$h_{1}${\ and }$h_{0}${. To see necessity, note that for all }$p>p^{\prime }${\ on }$P_{x}\cap P_{ \tilde{x}}${, }

equation*[equation* omitted — 242 chars of source]

{and}

equation*[equation* omitted — 270 chars of source]

{Thus ((ref)) and ((ref)) are equivalent.}

We collect the assumptions for identification as follows:

ASSUMPTION A-1: The distribution of $U$ is absolutely continuous with respect to Lebesgue measure.

ASSUMPTION A-2: The random vectors $(U,\varepsilon )$ and $(X,Z)$ are independent.

ASSUMPTION A-3: The random variable $g(v(X,1),\varepsilon )$ and $ g(v(X,0),\varepsilon )$ have finite first moments conditional on $U=u$ for all $u \in [0,1]$..

ASSUMPTION A-4: For any $(x,\tilde{x})\in S_{1}\times S_{0}$, $ F_{g|p}(y;v(x,1))=F_{g|p}(y;v(x,0))$ holds for all $y$ and $p\in P_{x}\cap P_{\tilde{x}}$ if and only if $v(x,1)=v(\tilde{x},0)$.

ASSUMPTION A-5: $\Pr (X\in S_{1})>0$ \ and $\Pr (X\in S_{0})>0$ .

Note that A-4 is weaker than Assumption 4 in \citeasnoun{vytlacilyildiz}. Specifically, to identify pairs $(x,\tilde{x})$ with $v(x,1)=v(\tilde{x},0) $ , \citeasnoun{vytlacilyildiz} relies on the assumption that $v(X,D) \in \mathbb{R}$ is a single index and that for any $ (x,\tilde x)\in (S_1\times S_0)$, $E\left[ g(v(x,1),\varepsilon )|U=u\right] =E\left[ g(v(\tilde{x},0),\varepsilon )|U=u\right] $ if and only if $ v(x,1)=v(\tilde{x},0) $. There are two shortcomings with this approach. First, it requires the condition (Assumption 4) that $E\left[ g(v(x,d),\varepsilon )|U=p\right] $ is a strictly monotonic function of $ v(x,d)$. Second, when $v(x,d)$ is a vector of multiple indices instead of a single index, their approach breaks down. In comparison, we achieve the same purpose by matching conditional distributions $F_{g|p}(\cdot ;v(x,1))$ and $\ F_{g|p}(\cdot ;v( \tilde{x},0))$. As we show in Section (ref), in several important applications, the outcome Y is either discrete (e.g. multinomial choices), or multi-dimensional with both discrete and continuous components (e.g., potential outcomes determined by a Roy model). In either cases, the latent index function $v(.)$ is vector-valued and the monotonicity condition in \citeasnoun{vytlacilyildiz} is not satisfied.

\setcounter{equation}{0}

Examples

{We now present several examples in which the latent indices are multi-dimensional. In the first and third example, the monotonicity condition in \citeasnoun{vytlacilyildiz} is not satisfied; in the second example, the identification requires a generalization of the monotonicity condition into an invertibility condition in higher dimensions.}

Example 1. (Heteroskaedastic shocks in outcome) Consider a triangular system where a continuous outcome is determined by double indices $v(X,D)\equiv (v_{1}(X,D),v_{2}(X,D))$:

equation*[equation* omitted — 108 chars of source]

The selection equation determining the actual treatment is the same as (1.2). In this case the concept of monotonicity in $v\in \mathbb{R}^{2}$ is not well-defined, so the procedure proposed in \citeasnoun{vytlacilyildiz} is not suitable here\footnote{ For this particular design, the approach proposed in \citeasnoun{vuongxu} should be valid. But it will not be for a slightly modified model, such as $ Y=v_1(X,D)+(e_2+v_2(X,D)*e_1)$, whereas ours will be.}. Nevertheless, we can apply the method in Section (ref)\ to identify the average treatment effect by using the distribution of outcome to find pairs of $x$ and $\tilde{x}$ such that $v(x,1)=v(\tilde{x},0)$. Assume the range of $v_2(\cdot)$ is positive. To see the necessity in Assumption A4, note that

eqnarray*[eqnarray* omitted — 167 chars of source]

for $d=0,1$. If the CDF of $\varepsilon $ is increasing over $\mathbb{R}$, then for all $y$ and $x \in S_1$ and $\tilde{x} \in S_0$,

equation*[equation* omitted — 60 chars of source]

if and only if

equation*[equation* omitted — 105 chars of source]

Differentiating with respect to $y$ yields

equation*[equation* omitted — 46 chars of source]

which in turn implies

equation*[equation* omitted — 54 chars of source]

The sufficiency in Assumption A-4 is straight-forward.

Example 2. (Multinomial potential outcome) Consider a triangular system where the outcome is multinomial. The multinomial response model has a long and rich history in both applied and theoretical econometrics. Recent examples in the semiparametric literature include \citeasnoun{lee95}, \citeasnoun{ahnetal2018}, \citeasnoun{shietal2018}, \citeasnoun{pakesporter2014}, \citeasnoun{KOT2019}. But unlike the work here, none of those papers allow for dummy endogenous variables or potential outcomes.

equation*[equation* omitted — 80 chars of source]

where

equation*[equation* omitted — 129 chars of source]

In this case the index $v\equiv (v_{j})_{j\leq J}$ and the errors $ \varepsilon \equiv (\varepsilon _{j})_{j\leq J}$ are both $J$-dimensional. The selection equation that determines $D$ is the same as (1.2). In this case, we can replace $1\{Y\leq y\}$ by $1\{Y=y\}$ in the definition of $ h_{1},h_{0},h_{1}^{\ast },h_{0}^{\ast }$ and $F_{g|u}(\cdot ;v)$. Then for $ d=0,1$ and $j\leq J$,\

eqnarray*[eqnarray* omitted — 238 chars of source]

By \citeasnoun{ruud2000} and \citeasnoun{ahnetal2018}, the mapping from $ v\in \mathbb{R}^{J}$ to $(F_{g|u}(j;v):j\leq J)\in \mathbb{R}^{J}$ is smooth and invertible provided that $\varepsilon \in \mathbb{R}^{J}$ has non-negative density everywhere. This implies Assumption A-4.

Example 3. (Potential outcome from the Roy model) Consider a treatment effect model with an endogenous binary treatment $D$ and with the potential outcome determined by a latent Roy model. The Roy model has also been studied extensively from both applied and theoretical perspectives. See for example the literature survey in \citeasnoun{heckvyt-handbook} and the seminal paper in \citeasnoun{heckmanhonoreb}.

Here the observed outcome consists of two pieces:\ \ a continuous measure $ Y=DY_{1}+(1-D)Y_{0}$ and a discrete indicator $W=DW_{1}+(1-D)W_{0}$ for $ d=0,1$. These potential outcomes are given by

equation*[equation* omitted — 114 chars of source]

where $a$ and $b$ index potential outcomes realized in different sectors, with

equation*[equation* omitted — 68 chars of source]

The binary endogenous treatment $D$ is determined as in the selection equation (1.2). For example, $D\in \{1,0\}$ indicates whether an individual participates in certain professional training program, $W_{d}\in \{a,b\}$ indicates the potential sector in which the individual is employed, $ y^*_{j,d}$ is the potential wage from sector $j$ under treatment $D=d$, and $ Y_{d}\in \mathbb{R}$ is the potential wage if the treatment status is $D=d$. As before, we maintain that $(X,Z)\bot (\varepsilon ,U)$.

The parameter of interest is

equation*[equation* omitted — 46 chars of source]

which by the independence $(X,Z)\bot (\varepsilon ,U)$ and an application of the law of total probability can be decomposed into directly identifiable quantities and a counterfactual quantity

eqnarray[eqnarray omitted — 195 chars of source]

Again, we seek to identify this counterfactual quantity by finding $\tilde{x} \in S_{0}$ such that

equation[equation omitted — 110 chars of source]

This would allow us to recover the right hand side of ((ref)) as

equation*[equation* omitted — 77 chars of source]

To find such a pair of $(x,\tilde{x})$, define $h_{d,W}(x,p,p^{\prime }),h_{d,W}^{\ast }(x,p)$ by replacing $1\{Y\leq y\}$ with $1\{W=a\}$ in the definition of $h_{d},h_{d}^{\ast }$ in Section (ref). Similarly, define $h_{d,Y}(x,y,p,p^{\prime }),h_{d,Y}^{\ast }(x,y,p)$ by replacing $1\{Y\leq y\}$ with $1\{Y\leq y,W=a\}$ in the definition of $ h_{d},h_{d}^{\ast }$ in Section (ref). Then

eqnarray*[eqnarray* omitted — 275 chars of source]

and $h_{d,W}(x,p_{1},p_{2})$ and $h_{d,Y}(x,y,p_{1},p_{2})$ are both identified over their respective domains by construction. Assume $ (\varepsilon _{a},\varepsilon _{b})$ is continuously distributed with positive density over $\mathbb{R}^{2}$ conditional on all $u$. Then the statement

eqnarray*[eqnarray* omitted — 303 chars of source]

holds true if and only if ((ref)) holds. Then matching $ h_{1,W}(x,p,p^{\prime })=h_{0,W}(\tilde{x},p,p^{\prime })$ ensures

equation[equation omitted — 95 chars of source]

while matching $h_{1,Y}(x,y,p,p^{\prime })=h_{0,Y}(\tilde{x},y,p,p^{\prime }) $ at the same time ensures that in addition to ((ref))

equation[equation omitted — 66 chars of source]

Combining ((ref)) and ((ref)) is equivalent to ((ref)).

\setcounter{equation}{0}

Extension

{The identification strategy we have used so far requires matching exogenous variables }$x,\tilde{x}${\ on }$S_{0},S_{1}$. In some cases, with the outcome being continuous, we can construct similar argument for identifying a counterfactual quantity in a treatment effect model by matching different elements on the support of continuous outcome. This approach was not investigated in \citeasnoun{vytlacilyildiz}, which focused on the use of first moment of outcome. The following example illustrates this point.

Example 4. (Potential outcome with random coefficients) Random coefficient models are prominent in both the theoretical and applied econometrics literature. They permit a flexible way to allow for conditional heteroscedasticity and unobserved heterogeneity. See, for example \citeasnoun{hsiaochapter} for a survey. Here we consider a treatment effect model where the potential outcome is determined through random coefficients:

equation*[equation* omitted — 109 chars of source]

and the binary endogenous treatment $D$ is determined as in the selection equation (1.2). The random intercepts $\alpha _{d}\in \mathbb{R}$ and the random vectors of coefficients $\beta _{d}$ are given by

equation*[equation* omitted — 117 chars of source]

where for any $x$ and $d \in \{0,1\}$., $(\bar{\alpha}_{d}(x),\bar{\beta} _{d}(x))\in \mathbb{R}^{K+1}$ is a vector of constant parameters while $\eta _{d}\in \mathbb{R}$ and $\varepsilon _{d}\in \mathbb{R}^{K}$ are unobservable noises.

As in \citeasnoun{vytlacilyildiz}, assume some elements in $Z$ in the selection equation are excluded from $X$. We allow the vector of unobservable terms $(\epsilon _{1},\epsilon _{0},\eta _{0},\eta _{1},U)$ to be arbitrarily correlated. We also assume that

equation[equation omitted — 129 chars of source]

with the marginal distribution of $U$ normalized to standard uniform, so that $\theta (Z)$ is directly identified as $P(Z)\equiv E(D|Z=z)$.

In this example our goal is to identify the conditional distribution of $ Y_{d}$ given $X=x$ for $d=0,1$. From this result we can identify parameters of interest such as average treatment effects, quantile treatment effects, etc. Let $G_{P|x}$ denote the conditional distribution of $P\equiv P(Z)$ given $X=x$, which is directly identifiable from the data-generating process. By construction,

equation*[equation* omitted — 92 chars of source]

where

eqnarray*[eqnarray* omitted — 148 chars of source]

The first term on the right-hand side is identified as

equation*[equation* omitted — 49 chars of source]

while the second term is counterfactual and can be written as

eqnarray*[eqnarray* omitted — 356 chars of source]

For any $p$ on the support of $P$ given $X=x$, define

eqnarray*[eqnarray* omitted — 435 chars of source]

where the second equality uses ((ref)). Likewise, under ((ref)) we have:

eqnarray*[eqnarray* omitted — 239 chars of source]

Assume\footnote{ This type of distributional equality assumption generalizes the exact equality of $\epsilon_1,\epsilon_0$ as can be found in for example \citeasnoun{vytlacilyildiz}. Distributional equality has been used to motivate the rank similarity condition imposed frequently in the econometrics literature- see for example \citeasnoun{chernozhukovhansen}, \citeasnoun{franlef}, \citeasnoun{dongshen}, \citeasnoun{chenstackhan}.}

equation[equation omitted — 164 chars of source]

Under ((ref)), we have

equation[equation omitted — 175 chars of source]

Suppose for each pair $(x,y)$ we can find $t(x,y)$ such that

equation*[equation* omitted — 134 chars of source]

Then by construction

eqnarray*[eqnarray* omitted — 317 chars of source]

because of ((ref)). Thus the counterfactual $\phi _{0}(x,y,p)$ would be identified as $h_{0}^{\ast }(x,t(x,y),p)$.

It remains to show that for each pair $(x,y)$ we can uniquely recover $ t(x,y) $ using quantities that are identifiable in the data-generating process. To do so, we define two auxiliary functions as follows: for $ p_{1}>p_{2}$ on the support of $P$ given $X=x$, let

eqnarray*[eqnarray* omitted — 235 chars of source]

and

eqnarray*[eqnarray* omitted — 235 chars of source]

Suppose $\eta _{d}+x^{\prime }\epsilon _{d}$ is continuously distributed over $\mathbb{R}$ for all values of $x$ conditional on all $u \in [0,1]$. Then for any fixed pair $(x,y)$ and $p_{1}<p_{2}$,

equation*[equation* omitted — 67 chars of source]

if and only if

equation*[equation* omitted — 134 chars of source]

To see this, suppose $t(x,y)>y-\bar{\alpha}_{1}(x)-x^{\prime }\bar{\beta} _{1}(x)+\bar{\alpha}_{0}(x)+x^{\prime }\bar{\beta}_{0}(x)$, then ((ref) ) implies that $h_{0}(x,t(x,y),p_{1},p_{2})>h_{1}(x,y,p_{1},p_{2})$. A symmetric argument establishes a similar statement with \textquotedblleft $>$ \textquotedblright \ replaced by \textquotedblleft $<$\textquotedblright . This establishes our desired result.

\setcounter{equation}{0}

Estimation

Here we outline estimation procedures from a random sample of the observed variables that are motivated by our identification results. We will first describe an estimation procedure for the parameter $E[Y_1]$ in the first three examples. Let $P_{x}$ to denote the support of $P(Z)$ given $X=x$, $ f_p(.|x)$ denote the density of $P(Z)$ given $X=x$, and

equation*[equation* omitted — 92 chars of source]

and for simplicity assume a strong overlap condition that

equation*[equation* omitted — 63 chars of source]

Define a measure of distance between $h_{1}(x_{1},\cdot )$ and $ h_{0}(x_{0},\cdot )$

eqnarray*[eqnarray* omitted — 273 chars of source]

where $w(y)$ is a chosen weight function. Consider the case when $ h_{1}(x,y,p_{1},p_{2})$, $h_{1}(x,y,p_{1},p_{2})$ and $P(z)$ are known. For any given $X_{i}$, let $\tilde X_{i}$ be such that

equation*[equation* omitted — 90 chars of source]

which, under Assumption A-4 in Section (ref), is equivalent to

equation*[equation* omitted — 46 chars of source]

Define

equation*[equation* omitted — 84 chars of source]

Note that the conditional expectation on the right-hand side is equal to $E[ Y | D=0, v( X , 0 ) = v( X_i , 1), P=P_i] $, which in turn equals $E[ Y_1 | D=0, X=X_i, P=P_i] $. Then, following the discussion above, we define the following estimator for $\Delta \equiv E[Y_{1}]$:

equation*[equation* omitted — 101 chars of source]

or a weighted version

equation*[equation* omitted — 200 chars of source]

Limiting distribution theory for each of these estimators follows from identical arguments in \citeasnoun{vytlacilyildiz}. Here we formally state the theorem for the first estimator:

theoremUnder Assumptions A-1 to A-5, and the additional assumption that $Y_1$ has positive and finite second moment, then we have \begin{equation*} \sqrt{n}(\hat \Delta-\Delta)\overset{d}{\rightarrow} \mathbb{N}(0,V) \end{equation*} where \begin{equation*} V=Var(E[Y_1|X,P,D])+E[PVar(Y_1|X,P,D=1)] \end{equation*}

Now we describe an estimation procedure for the distributional treatment effect in Example 4, where we had a model with random coefficients. In this case, the parameter of interest is for a chosen value of the scalar $y$,

equation*[equation* omitted — 57 chars of source]

First, for fixed values of $y$ and $p_{1}>p_{2}$, we propose to estimate $ t(x,y)$ as

equation*[equation* omitted — 106 chars of source]

and then average over values of $p_{1},p_{2}$:

equation*[equation* omitted — 102 chars of source]

An infeasible estimator for the parameter $\Delta_2(y)$, which assumes $t(x,y)$ is known, would be

equation*[equation* omitted — 141 chars of source]

In practice, for feasible estimation, one needs to replace $t(x,y)$ by its estimator $\hat{\tau}(x,y)$.

\setcounter{equation}{0}

Simulation Study

This section presents simulation evidence for the performance of the proposed estimation procedures described in Section (ref), for both the Average Treatment Effect and the Distributional Treatment Effect. We report results for both our proposed estimator and that in \citeasnoun{vytlacilyildiz}, for several designs. These include designs where the said monotonicity condition fails, and designs where the disturbance terms in the outcome equation are multidimensional.

Throughout all designs we model the treatment or dummy endogenous variable as

equation*[equation* omitted — 27 chars of source]

where $Z,U$ are independent standard normal. We experiment with the following designs for the outcome

description\begin{equation*} Y=X+0.5\cdot D +\epsilon \end{equation*} where $X$ is standard normal, $(\epsilon, U)$ are distributed bivariate normal, each with mean 0 and variance 1, with correlations of 0,0.25,0.5. • \begin{equation*} Y=X+0.5\cdot D+(X+D)\cdot \epsilon \end{equation*} where $X$ is distributed standard normal, $(\epsilon, U)$ are distributed bivariate normal, each with mean 0 and variance 1, with correlations of 0,0.25,0.5. • \begin{equation*} Y=(X+0.5 \cdot D+\epsilon)^2 \end{equation*} where $X$ is distributed standard normal, $(\epsilon, U)$ are distributed bivariate normal, each with mean 0 and variance 1, with correlations of 0,0.25,0.5.

We note that the monotonicity condition is satisfied in design 1 but fails in the other two designs. For each of these designs, we report results for estimating the parameter $E[Y_1]$, which denotes the expected value for potential outcome under treatment $D=1$. The two estimators used in the simulation study were the one proposed in Section (ref) and the method proposed in \citeasnoun{vytlacilyildiz}. The summary statistics, scaled by the true parameter value, Mean Bias, Median Bias, Root Mean Squared Error, (RMSE), and Median Absolute Deviation (MAD) were evaluated for sample sizes of 100, 200, 400 for 401 replications. Results for each of these designs are reported in Tables 1 to 3 respectively. In implementing our procedure we assumed the propensity score function is known, and conducted next stage estimation using a nonparametric kernel estimator with normal kernel function, and a bandwidth of $n^{-1/5}$. This rate reflects “undersmoothing" as there are two regressors, the propensity score and the regressor $X$. For the estimator in \citeasnoun{vytlacilyildiz}, which involved the derivative of conditional expectation functions as well, estimating these functions nonparametrically gave very unstable results so we report results for an infeasible version of their estimator, assuming such functions, as well as the propensity scores, are known.

To implement the second stage of our proposed procedure, in calculating the distance $\|h_1(x_i,\cdot)-h_0(x_0,\cdot)\|$ we used an evenly space grid of values for $y$, and selected $n/50$ grid points, with $n $ denoting the sample size.

The results indicate the desirable properties of our proposed procedure, generally agreeing with Theorem (ref). In all designs our estimator has small values for bias and RMSE, with the value of RMSE decreasing as the sample size grows. In contrast, the procedure based on \citeasnoun{vytlacilyildiz} only performs well in Design 1, with values of bias and RMSE comparable to those using our method. As in our procedure these values decrease with as the sample size grows, which is expected, as the monotonicity condition rely on is satisfied in these designs. In this case, their approach has smaller standard errors largely due to the relative simpler structure of the infeasible version, but their biases persist even when the sample size increases.

For designs 2 and 3, where monotonicity is violated, the procedure proposed in \citeasnoun{vytlacilyildiz} does not perform well. In design 2 in Table 2 both the bias and RMSE are generally increasing with the sample size. Results for their estimator are better in design 3, but the bias hardly converges with the sample size and is much larger compared to our estimator.

We also simulate data from a model with dummy endogenous variable and potential outcomes determined by random coefficients. It is important to note that for this design, the original matching idea in \citeasnoun{vytlacilyildiz} does not apply. This is because different values of $x$ lead to different distribution of the composite error $\eta_d + x^{\prime }\epsilon_d$. Our contribution in Section (ref) is to propose a new approach based on matching different values of outcome $y$, rather than the regressors $x$. Based on the counterfactual framework discussed in Section (ref), here the treatment variable $D$ is modeled as the same way as the dummy endogenous variable above. Similarly the regressor $X$ is standard normal. For both $Y_0,Y_1$ the random intercepts were modeled as constants (0 and 1, respectively) and the additive error terms were each standard normal. For the random slopes, the means were 1 and 2 respectively, and the additive error terms were also standard normal, independent of all other disturbance terms and each other. Here we use the procedure in Section (ref) to estimate the parameter $\Delta_2=P(Y_1<y)$, where in the simulation we set $y=1$. The same four summary statistics are reported for sample sizes 100,200,400, based on 401 replications. Results for this random coefficients design are reported in Table 4.

The estimator proposed in Section (ref) performs well; but the bias and RMSE are much small at 400 observations compared to 100 and 200 observations, indicating convergence at the parametric rate.

center[center omitted — 4,113 chars of source]

{ \setcounter{equation}{0} }

Conclusion

{ In this paper, we considered identification and estimation of nonseparable models with endogenous binary treatment. Existing approaches are based on a monotonicity condition, which is violated in models with multiple unobserved idiosyncratic shocks. Such models arise in many important empirical settings, including Roy models and multinomial choice models with dummy endogenous variables, as well as treatment effect models with random coefficients. We establish novel identification results for these models which are constructive and conducive to estimation procedures which are easy to compute and whose limiting distributional properties follow from standard large sample theorems. A simulation study indicates adequate finite sample performance of our proposed methods. }

{ This paper leaves open areas for future research. Our method requires the selection of the number and location of cutoff points, so a data driven method for selecting these would be useful. Furthermore, the relative efficiency of our proposed approach needs to be explored, perhaps by deriving efficiency bounds for these new classes of models. }