EconBase
← Back to paper

Double machine learning for causal inference in a multivariate sample selection model

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

26,710 characters · 0 sections · 1 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

{\bf Proof of Theorem 3 from Section (ref). Identification of LATES.}

{\bf Part 1. Case }$\mathbf{m = 1}\text{ }${\bf .}

We follow very closely the structure of the proof of Theorem 1 from Frolich and extend it to the case of MSSM.

Consider $(x, p, w, z) \in \mathrm{supp}(X_i, \bar{P}_i^{(z)}, W_i^{(D)}, Z_i)$. For brevity, denote $\tilde{X}_i = (X_i, \bar{P}_i^{(z)}, Z_i)$ and $\tilde{x} = (x, p, z)$. By using the law of total expectation, we obtain:

equation[equation omitted — 1,298 chars of source]

Note that if $W_i^{(D)} = 1$, then we observe $Y_{1i}$ for compliers. Similarly, if $W_i^{(D)} = 0$, then we observe $Y_{0i}$ for compliers. For always-takers and never-takers we always observe $Y_{1i}$ and $Y_{0i}$, respectively. By using these facts, we obtain:

equation[equation omitted — 1,799 chars of source]

Because of Assumptions 1F, 4F, and 5F, we may apply Lemma 2, which allows us to remove the conditioning on $W_i^{(D)}$ in the expectations and probabilities:

equation[equation omitted — 1,745 chars of source]

Because of Assumption 3F, we may divide both sides of the last expression by the conditional probability of complying:

equation[equation omitted — 520 chars of source]

To expand the denominator in the last expression, note that:

equation[equation omitted — 703 chars of source]

and:

equation[equation omitted — 699 chars of source]

Hence, by taking the difference, we obtain:

equation[equation omitted — 546 chars of source]

By plugging equation ((ref)) into the denominator of equation ((ref)), we get the expression for the conditional local average treatment effect:

equation[equation omitted — 799 chars of source]

where the conditional expectations used in this expression exist by Assumption 6F.

By taking an expectation and applying the Bayes' theorem, we obtain:

equation[equation omitted — 1,565 chars of source]

By expanding $\mathrm{CLATES}(X_i, \bar{P}_i^{(z)} \mid Z_i = z)$ via equation ((ref)), we finally establish the result for $m = 1$:

equation[equation omitted — 1,551 chars of source]

{\bf Part 2. Case }$\mathbf{m > 1}${\bf .}

We apply the law of total expectation:

equation[equation omitted — 611 chars of source]

Note that for any $z^* \in \{z^{(1)}, \dots, z^{(m)}\}$, we have:

equation[equation omitted — 904 chars of source]

By applying equation ((ref)) to the denominator of equation ((ref)), we obtain:

equation[equation omitted — 492 chars of source]

By plugging equation ((ref)) into equation ((ref)) and then this modified equation ((ref)) into equation ((ref)), we obtain:

equation[equation omitted — 1,368 chars of source]

Note that in the last expression, the numerator of the first factor cancels with the denominator of the second factor. Therefore, we finally obtain:

equation[equation omitted — 1,291 chars of source]

$\hfill\blacksquare$

{\bf Proof of Theorem 4 from Section (ref). Identification of LATE.}

{\bf Part 1. Case }$\mathbf{m = 1}${\bf .}

Consider $\left( x, p, w, z \right) \in \text{supp}\left( X_i, \bar{P}_i^{(z)}, W_i^{(D)}, Z_i \right)$. By using Lemma 3 and Assumption 8F, we obtain:

equation[equation omitted — 1,657 chars of source]

where the last equality is easy to derive by using the same steps as in the proof of Theorem 3.

By defining $\mathrm{CLATE}(x,p)$ and applying equations ((ref)), ((ref)), and Assumption 8F, we obtain:

equation[equation omitted — 1,405 chars of source]

Therefore, we have established that $\mathrm{CLATE}(x,p)$ equals $\mathrm{CLATES}(x,p \mid Z_i=z)$. In addition, we have simplified its expression. By using these results and the same approach as in the proof of Theorem 3, we obtain:

equation[equation omitted — 1,733 chars of source]

Note that if $D_i$ is not subject to non-random selection, then pre-last representation of LATE is preferable. However, for greater generality, we consider the last representation in which $D_i$ is conditioned on $Z_i=z$.

{\bf Part 2. Case }$\mathbf{m > 1}${\bf .}

Since $\sum\limits_{t=1}^{m} \mathbb{P}\left( Z_i = z^{(t)} \mid \tilde{Z}_i = 1, \mathrm{complier}_i = 1 \right) = 1$, the expression for $m>1$ is as follows:

equation[equation omitted — 556 chars of source]

By inserting equation ((ref)) into equation ((ref)), we obtain:

equation[equation omitted — 971 chars of source]

By plugging equation ((ref)) into equation ((ref)), we finally obtain:

equation[equation omitted — 571 chars of source]

$\hfill\blacksquare$

\ifSubfilesClassLoaded