Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
57,109 characters · 18 sections · 110 citation commands
Robustify and Tighten the Lee Bounds: A Sample Selection Model under Stochastic Monotonicity and Symmetry Assumptions
{Keywords: attrition, partial identification, randomised controlled trial, treatment effects}
Sample selection is an important challenge for identifying causal effects in empirical economics (see, e.g., Hansen2022econometrics). Earlier approaches model the selection mechanism as follows:
where $Y_i^*$ is the latent outcome and researchers observe $(Y_i, X_i, W_i)$; and assume that the covariate $W_i$ includes some variables which are not in $X_i$, meaning that researchers can access some “instrument" that only affects an individual's selection decision but not the outcome. A similar assumption is made by not only such single index semiparametric models but more general nonparametric models (e.g., Das_elat:2003). However, this exclusion restriction is not always satisfied in practice because such an instrument may not be easily accessible. To overcome this challenge, Lee:2009 developed a partial identification strategy that does not necessitate this restriction. The Lee:2009 bounds are especially used in field experimental studies, wherein the attrition problem often occurs (e.g., Baranov2020, Bursztyn2020, Muralidharan:2019). Another application includes Chen_Roth:2023.
Although popular, the Lee:2009 bounds have two important limitations. The first is the monotonicity assumption. The identification of the Lee:2009 bounds relies on the assumption of a monotone response to the treatment assignment; specifically, for every person, their outcome under treatment has to be observable if it would be observed in the absence of treatment. This assumption, while classic and seemingly natural in many contexts, is somewhat strong and its empirical validity can be unclear in several cases, because few economic theories strongly affirm the non-existence of any individuals deviating from behavioural models.
The second limitation is that the Lee:2009 bounds can be wide. For example, Mobarak:2023 wrote, “[c]ommon nonparametric bounds (e.g., Lee:2009 Lee:2009) are often wide and uninformative, as they trim the data from the extremes of the outcome distribution." Several studies provide a similar assessment (e.g., Barrow_Rouse:2018, Delius_Sterck:2024). Therefore, considering the Lee:2009 bounds are the tightest possible under his assumptions, it will be beneficial to consider an additional, reasonable restriction to help draw more informative policy implications.
This article aims to address these issues by introducing new choices of assumptions. We first relax the monotonicity assumption to its stochastic counterpart, which allows a certain stochastic deviation from the monotone response. This stochastic monotonicity assumption operates on the intuition that “the monotonicity approximately, but perhaps not exactly, holds." Such scenarios will frequently arise in empirical analysis. Moreover, embracing a stochastic variation to choice aligns more closely with observed experimental evidence (e.g., Tversky1969, Agranov_Ortoleva:2017). Afterwards, we introduce a nonparametric distributional assumption that helps tighten the bounds. Specifically, we propose using a symmetry assumption---which need not be satisfied exactly (Assumption (ref))---on the density function of $Y_i^*$ under treatment for always-takers, with its motivation explained graphically. This symmetry assumption restricts the admissible class of trimmed densities, enabling us to avoid Lee:2009's (Lee:2009) “extreme" trimming of the outcome distribution, thereby tightening the bounds over the standard Lee:2009 bounds.
Based on these assumptions, we develop partial identifications and propose simple estimators. Importantly, the bounds are root-$n$ consistently estimable, making the implementation practically feasible. The potential usefulness of the bounds is demonstrated using an empirical example.
This article proceeds as follows. The rest of this section reviews the related literature. In Section (ref), we discuss the identification of the sample selection model under the stochastic monotonicity assumption. Section (ref) introduces the symmetry assumption and establishes identification. Some numerical examples are also provided in these sections. Estimation and inference procedures are considered in Section (ref). The empirical illustration in Section (ref) highlights the usefulness of the obtained bounds. Section (ref) concludes this article.
The sample selection problem has been an important issue in economics (Heckman:1979). Lee:2009 is a seminal work that derives nonparametric bounds under the monotonicity assumption based on the insight of Horowitz_Manski:1995. Similar studies include Imai:2008 and Chen_Flores:2015. Based on a similar idea, Bartalotti_etal:2023 recently considered the identification of the marginal treatment effect in the presence of sample selection, which essentially extends Lee:2009's model.
This study is not the first to focus on the monotonicity assumption. Relatedly, Semenova:2023 explored a different kind of generalization of the Lee:2009 bounds, specifically identifying the bounds assuming that exact monotonicity holds conditional on the covariates. This approach will be very useful when the direction of selection effects systematically differs depending on the covariates. However, in some cases, researchers may not have access to such informative covariates. Besides, even conditional on the covariates, the requirement that the monotone response in selection has to hold for every such person may still be practically strong. Hence, when the predictive power of the covariates seems weak but one can assume that the fraction of those deviating from complete monotonicity is not so large, our approach can be a useful complementary tool.
Our relaxation of the exact monotonicity assumption using the information of the lower bound on a parameter (see Assumption (ref)) is similar to the techniques used in the sensitivity analysis literature, such as Conley_etal:2012, who treated the exogeneity condition in linear instrumental variable methods. The stochastic monotonicity assumption is also similar to the bounded variation assumption found in the difference-in-differences and regression discontinuity literature (e.g., Imbens_Wager:2019, Manski_Pepper:2018, Rambachan_Roth:2023). Note that, in our setting, the lower bound (denoted by $\vartheta_L$ in Assumption (ref)) can be easier for a researcher to set because $\vartheta_L$ is the bound on probability, and thus, has an intuitive meaning. Furthermore, what we are imposing is clear.
The width of the Lee:2009 bounds are also discussed in the literature. Beheghel_etal:2015 proposed an interesting idea to derive narrower nonparametric bounds on the treatment effect, whereas an instrument that only affects the participation decision is needed. Honore_Hu:2020, Honore_Hu:2022 derived much sharper bounds than Lee:2009 bounds under a fully structured semiparametric models. Although their results do not rely on the excluded instrument, their identification depends on the imposed structural assumptions, meaning that it is less robust to specification error than nonparametric models as Lee:2009. Intuitively, our approach utilising (nonparametric) distributional assumption is inbetween these two studies: Our procedure does not require the additional instruments but still builds on Lee:2009's insight.
Finally, our procedure can be applicable to various empirical settings in which sample selection can occur. A prominent example is the attrition problem in randomised controlled trials (Duflo_etal:2007, Section 6.4). To demonstrate such an application, we use the data from Muralidharan:2019's (Muralidharan:2019) study as an empirical example in Section (ref).
Here, we introduce a stochastic version of the monotonicity assumption and provide the partial identification under this assumption. Section (ref) reviews the general selection model of Lee:2009 and the monotonicity assumption. In Section (ref), we extend Lee:2009's (Lee:2009) results to establish the more robust bounds under the stochastic monotonicity.
Following Lee:2009, we consider the following general selection model:
where $Y_{1i}^*$ ($Y_{0i}^*$) is the potential outcome when an individual $i$ is assigned to a treatment (control) group; $D_i\in\{0,1\}$ is a binary indicator of treatment status that equals $1$ when the individual is treated; and $S_{1i}\in\{0,1\}$ and $S_{0i}\in\{0,1\}$ are the potential sample selection indicator under the treated and control states, respectively. $S_i$ equalis $1$ when the individual $i$ does not exhibit attrition, and $0$ otherwise.
We are interested in the average treatment effect for the always takers ($S_{0i}=S_{1i}=1$), or
Lee:2009 showed that $\tau$ can be partially identified under the following assumptions:
The basic idea of the Lee:2009 bounds is as follows. Under Assumptions (ref) and (ref), we have
The point identification fails due to the first term of the right-hand side, as it is not identified. Here, under Assumption (ref), we have the following decompsition:
where $q_0$ is defined in Theorem (ref), which is identified and in $[0,1]$ under Assumption (ref). By assuming that the smallest $q_0$ values of $Y_i$ are entirely attributed to the group with $S_{0i}=1, S_{1i}=1$, we can derive the lower bound as follows:
which is Lee:2009's lower bound. The upper bound can be obtained similarly.
Although it plays several important roles in the above transformation, the monotonicity assumption is somewhat restrictive. Intuitively speaking, it requires that every individual in the population obeys the behavioural rule that $S_{0i}=1\Longrightarrow S_{1i}=1$ with no exception. This may fail if the population consists of heterogeneous individuals. Further, as discussed in Section (ref) using an empirical example, several context-dependent concerns about its plausibility may exist.
However, if the monotonicity assumption is roughly true, would the Lee:2009 bounds successfully produce valid bounds? Unfortunately, a slight deviation from the exact monotonicity can lead to misleading conclusions, as illustrated in the next example:
Therefore, when the monotonicity does not exactly hold, the Lee:2009 bounds can be invalid.
When computing the Lee:2009 bounds, economists often operate under the premise that the monotonicity assumption is largely, though perhaps not perfectly, met. As highlighted by the above example, considering bounds that account for this nuanced understanding can enhance robustness. To align with this perspective, this subsection introduces a weaker and interpretable assumption that builds on the intuition that “the monotonicity mostly, but perhaps not exactly, holds" and then derives the bounds on the treatment effect $\tau$ under this assumption.
In particular, we consider the identification under the following stochastic version of the monotonicity:
Part (i) states that the monotonicity holds with a probability of no less than some known value $\vartheta_L$. This $\vartheta_L$ can be considered a representation of the validity of the monotonicity. It can be determined based on a researcher's institutional knowledge and beliefs. For example, when a researcher assumes that almost all participants satisfy the monotonicity assumption but anticipates that a small fraction of them perhaps deviates from the monotonic behaviour, setting $\vartheta_L=0.95$ (alternatively $\vartheta_L=0.90$ or $0.99$) is reasonable in a similar vein to the statistical significance level. This kind of bound on a parameter is common in various settings (Conley_etal:2012, Manski_Pepper:2018, Imbens_Wager:2019, Rambachan_Roth:2023), whereas its intuitive meaning is not always clear. However, the lower bound $\vartheta_L$ has a clear meaning and it is easy to map the researcher's knowledge to it.
The lower bound $\vartheta_L$ provides another intuitive relation. Under Assumption (ref), we have
Hence, under Assumption (ref)-(i), it follows that
That is, setting $\vartheta_L$ is equivalent to determining the upper bound on $\P{S_{1i}=1 | S_{0i}=0}$ once the identified parameters are known. This could be used to empirically validate Assumption (ref)-(i) by computing the empirical analogue of the upper bound and checking if $\P{S_{1i}=1 | S_{0i}=0}$ is not unreasonably small.
Part (ii) states that the attrition and potential outcomes are irrelevant for those whose outcomes are (potentially) observable without treatment. Lee:2009's monotonicity assumption automatically satisfies this. However, it is not explicitly treated in Lee:2009, and thus, we discuss this assumption a little more. Assumption (ref)-(ii) can be rewritten as
Importantly, $Y_{1i}^*, Y_{0i}^*$ can affect $S_{1i}$ only through $S_{0i}$ if $S_{0i}=1$. When we suppose that $S_{1i}$ still depends on $Y_{1i}^*, Y_{0i}^*$ after controlling $S_{0i}=1$, condition (ref) will fail. In contrast, this condition can be satisfied, for example, when
where $k_0$ and $k_1$ are some structural functions, and $\zeta_{0i}$ and $\zeta_{1i}$ are idiosyncratic errors. This behavioural structure includes the following examples:
As illustrated in Example (ref), Assumption (ref) accounts for scenarios where Assumption (ref) is largely reasonable, yet subject to potential deviations due to stochastic shocks affecting the decision-makers' tastes. Note that, other than this rather restrictive model's setup in Example (ref), such a stochastic flip of choice is widely acknowledged in decision theory, as evidenced by Tversky1969 and Agranov_Ortoleva:2017.
Under this stochastic monotonicity assumption, we have the following bounds.
The max operator and $\vartheta_F$ are due to the Fréchet inequality, which is equivalent to the inequality constraint $\P{S_{1i}=1|S_{0i}=0}\leq 1$; see also (ref). This ensures that $\P{S_{1i}=1|S_{0i}=1}\geq\vartheta_F$ under Assumption (ref). Therefore, the trimming is not performed at $y^1_{\vartheta_L q_0}$ but at $y^1_{\vartheta q_0}$.
When $\vartheta_L=1$, the bounds coincide with Lee:2009's (Lee:2009). As can be seen from Theorem (ref), the length between the bounds becomes wider as $\vartheta_L$ decreases as long as $\vartheta_L\geq\vartheta_F$. This is natural since the less information we have about the individual's selection rule, the greater the length of the bounds. With this monotonic relation between $\vartheta_L$ and the bounds, the bounds under Assumption (ref) can also be used to perform sensitivity analysis or examine the monotonicity assumption's identifying power. For example, when researchers are concerned about whether the monotone response assumption is reasonable, drawing the bounds with multiple $\vartheta_L$ will be helpful. If the bounds do not cover zero with somewhat smaller $\vartheta_L$, researchers can safely conclude that the treatment has a positive effect.
In the previous subsection, we saw a simple example where a small deviation from the monotonicity leads to misleading bounds. The stochastic monotonicity can produce more robust bounds:
While the bounds have been made more robust with respect to the deviation from monotonicity, this robustification may lead to wider bounds that could obscure policy implications. This is especially concerning given that the standard Lee:2009 bounds are already considered to be wide. The subsequent section explores a method to tighten these bounds while maintaining robustness against deviations from monotonicity.
The Lee:2009 bounds are often argued to be too wide and less informative even under the exact monotonicity assumption (e.g., Barrow_Rouse:2018, Delius_Sterck:2024, Mobarak:2023). To address this practical issue, we introduce an additional assumption to tighten the bounds. As this assumption does not require additional variables that satisfy the exclusion restriction, the following identification result can be useful in the absence of such variables. With such an additional excluded variable at hand, the procedure proposed in Beheghel_etal:2015 may be an alternative tool.
Before introducing the assumption, we again recall the idea behind the Lee:2009 bounds. This clarifies why the bounds can be wide and motivates our assumption as a natural extension. For ease of exposition, temporarily assume that the exact monotonicity holds. As noted in Section (ref), the point identification fails due to $\E{Y_{1i}^*|S_{0i}=S_{1i}=1}$. An important insight from Lee:2009 was the decomposition (ref), and we obtain Lee:2009's lower bound as
by assuming that the smallest $q_0$ values of $Y_i$ is entirely attributed to $Y_{1i}^*$ of the group with $S_{0i}= S_{1i}=1$. Importantly, this is equivalent to assuming that the conditional density of $Y_{1i}^*$ given $S_{0i}=S_{1i}=1$ can, after appropriate rescaling, have the form like the shaded area in Figure (ref), wherein the solid line represents the density of $Y_i$ conditional on $D_i=1, S_i=1$.
Is this form, right-truncated at $y^1_{q_0}$, empirically natural? In several cases, this may be counter-intuitive and perhaps makes practitioners suppose that “Lee:2009 Bounds are based on extreme assumptions about sample selection" (Delius_Sterck:2024).
In many empirical contexts, a “well-behaved" density is more probable. One promising candidate is the class of symmetric densities, especially since the distribution of several economically important variables, such as log wages and test scores, frequently shows approximate symmetry (see also Section S3 of the Online Appendix and Figure (ref)). Assuming the conditional density is symmetric, we can avoid the aforementioned right-truncated trimmed density, and the lower bound is attained when it has the form as in Figure (ref). Such a functional form may not be counter-intuitive and perhaps more natural in practice. Additionally, this additional information could tighten the bounds. Further, the symmetry immediately implies that the lower bound is given by
which can be easily estimated, bypassing the need for, for example, a first-stage nonparametric estimation. The following subsection formalises the symmetry assumption and provides identification.
Motivated by the previous discussion, we introduce the following symmetry assumption:
We also introduce the following tail smoothness condition, which is implicitly used in Figure (ref). This assumption is not essential to obtain valid bounds.
Intuitively, it says that the shaded area in Figure (ref) is below the conditional density $f_{y}^1$. Before discussing these assumptions, we provide the identification results:
Therefore, we can obtain tightened bounds.
Besides, as Theorem (ref) suggests, the symmetry assumption can be used in conjunction with the stochastic monotonicity assumption, as illustrated by the following example:
We now shift our focus to the assumptions. The symmetry assumption might provoke the most debate. Given the common use of symmetric distributions in economic modelling and econometrics (e.g., Powell:1986), this assumption is likely to be accepted among economists and serves as a reasonable starting point. Moreover, some economically important variables, such as log wages, seem to often exhibit symmetric-like distributions, as documented in the Online Appendix (Section S3). Furthermore, exact adherence to the symmetry assumption is not requisite. Indeed, the validity of the bounds, which is the most important part in practice, can still be ensured under a less restrictive condition; that is, the mean-median coincidence assumption:
Formally, Theorem (ref) holds by assuming this instead of Assumption (ref). Therefore, a much larger class of distributions, including asymmetric distributions, is allowed for $[\rotatebox[origin=c]{180}{$\nabla$}^{\mathrm{s}}, \nabla^{\mathrm{s}}]$ to be valid.
Further justification is contingent on the specific empirical contexts and the researchers' expertise. An illustrative discussion on this using a concrete empirical example is provided in Section (ref).
Note that the set of Assumptions (ref) to (ref) can be viewed as an additional layer for a layered analysis, which examines how the bounds differ with our imposed assumptions (see Manski_Nagin:1998). Therefore, reporting the bounds with and without the symmetry assumption is valuable, even amidst certain scepticism. This enables not just researchers but also policymakers to reassess the credibility of this assumption, and adjust their interpretation of the reported bounds accordingly.
Second, an empirically important result is Corollary (ref). For this result, Assumption (ref) plays a role. Meanwhile, this assumption is just a sufficient condition for $[\rotatebox[origin=c]{180}{$\nabla$}^{\mathrm{s}}, \nabla^{\mathrm{s}}]\subseteq [\rotatebox[origin=c]{180}{$\nabla$}, \nabla]$ to be true. For example, $\rotatebox[origin=c]{180}{$\nabla$}\leq\rotatebox[origin=c]{180}{$\nabla$}^{\mathrm{s}}$ is often satisfied when the trimmed density without the symmetry assumption (i.e., the shaded area of Figure (ref)) is left-skewed, as is the case in Figure (ref). The left-skewness often occurs when the density smoothly decreases in its left tail, due to Lee:2009's trimming procedure (Figure (ref)). Hence, in many cases, $[\rotatebox[origin=c]{180}{$\nabla$}^{\mathrm{s}}, \nabla^{\mathrm{s}}]$ may be tightened even when Assumption (ref) does not hold.
Finally, we can verify Assumption (ref) by drawing a density estimate using a standard estimation strategy, such as the kernel density estimator or the local polynomial density estimator (Cattaneo_etal:2020). This visual inspection is useful to check the sharpness of the bounds, or examine the potential for further narrowing the obtained bounds. Some discussion on this is provided in Section (ref) using an empirical example.
We consider the estimation and inference procedures. We focus on the lower bound obtained in Theorem (ref) to avoid redundancy. The upper bound is analogous and the procedures for the bounds obtained in Theorem (ref) are similar, which are provided in the Online Appendix. We will consider two types of estimators in the subsequent subsections.
We first consider when $\vartheta_L \geq \vartheta_F$ is known. The inequality $\P{S_{1i}=1|S_{0i}=1} \geq \vartheta_F$, or equivalently $\P{S_{1i}=1|S_{0i}=0}\leq 1$, is not a very strong requirement. Hence, $\vartheta_F$ may not be very large. Thus, $\vartheta_L \geq \vartheta_F$ can be satisfied in many applications, especially when $\vartheta_L$ is large, e.g., $\vartheta_L=0.95$. Lee:2009's exact monotonicity ($\vartheta_L=1$) is an leading example of this case.
Define $\bm{\beta} \coloneqq (\beta^L, q, \alpha, \eta)$ and $\bm{\beta}_0 \coloneqq (\beta^L_0, q_0, \alpha_0, \eta_0)$, where $\beta^L_0 = y^1_{\vartheta_L q_0/2}$ and $\eta_0 = \E{Y_i | D_i=0, S_i=1}$. Similarly to Lee:2009, defining
we can estimate $\bm{\beta}_0$ by $\min_{\bm{\beta}} (\sum_{i=1}^n g(\bm{\beta}))^\prime(\sum_{i=1}^n g(\bm{\beta}))$. Write the minimiser by $\hat{\bm{\beta}} = (\hat{\beta}^L, \hat{q}, \hat{\alpha}, \hat{\eta})$. Then, we can estimate the lower bound by $\widehat{\rotatebox[origin=c]{180}{$\nabla$}^{\mathrm{s}}_1} = \hat{\beta}^L - \hat{\eta}$. We have the next asymptotic normality.
We can define $\widehat{\nabla^{\mathrm{s}}_1} $ similarly and have $\sqrt{n}\left(\widehat{\nabla^{\mathrm{s}}_1} - \nabla^\mathrm{s}\right) \to_d \mathcal{N}(0, \Omega_U + \Omega_C)$, where
The probabilities and conditional variance can be estimated by their sample analogues. The conditional density $f^1_y$ can be consistently estimated by the standard kernel method, and standard bandwidth selectors can be used by assuming additional smoothness assumption on $f^1_y$ (see, e.g., Jones_etal:1996). Then, the asymptotic variances can be consistently estimated by the continuous mapping theorem. Therefore, similarly to Lee:2009, we can perform inference based on Imbens_Manski:2004's (Imbens_Manski:2004) confidence interval (CI).
Next, we consider the estimation and inference procedure when a researcher does not know which is larger, $\vartheta_L$ and $\vartheta_F$. Define $\bm{\gamma} \coloneqq (\gamma^L, \gamma^F, q, \alpha, \eta)$ and $\bm{\gamma}_0 \coloneqq (\gamma^L_0, \gamma^F_0, q_0, \alpha_0, \eta_0)$, where $\gamma^L_0 = y^1_{\vartheta_L q_0/2}$ and $\gamma^F_0 = y^1_{(1 + q_0(1-1/\alpha_0))/2}$. Defining
we can estimate $\bm{\gamma}_0$ by $\min_{\bm{\gamma}} (\sum_{i=1}^n \Tilde{g}(\bm{\gamma}))^\prime(\sum_{i=1}^n \Tilde{g}(\bm{\gamma}))$. We represent the minimiser by $\hat{\bm{\gamma}} = (\hat{\gamma}^L, \hat{\gamma}^F, \hat{q}, \hat{\alpha}, \hat{\eta})$. We can estimate $\rotatebox[origin=c]{180}{$\nabla$}^\mathrm{s}$ by the following:
Now, we have the following lemma:
Then, we can apply the procedure developed by Chernozhukov_etal:2013 for inference. Conceptually, Chernozhukov_etal:2013's (Chernozhukov_etal:2013) CI is different from Imbens_Manski:2004's (Imbens_Manski:2004) in that the former is the CI for the identified region, i.e., the bounds $[\rotatebox[origin=c]{180}{$\nabla$}^\mathrm{s}, \nabla^\mathrm{s}]$, while the latter is for the true parameter, $\tau$. Therefore, the CI based on Chernozhukov_etal:2013 can be conservative for $\tau$, although valid.
We illustrate the usefulness of the bounds based on the stochastic monotonicity and symmetry assumptions using the data from Muralidharan:2019. The authors used a randomised controlled trial to examine the effect of the after-school education programme, called Mindspark intervention, on math and Hindi test scores.
The Mindspark intervention consists of two sessions, technology-led and group-based instructions, while the instruction was mainly provided in the formar session (Muralidharan:2019). This technology-led instruction is a computer-based interactive instruction. The content provided in this session is automatically personalised for each student, and this adaptation is performed dynamically based on the beginning assessment and every subsequent activity completed.
The educational attainment is assessed by the baseline and endline tests. These tests are designed to capture a broad range of students' achievements. The test questions ranged in difficulty from “very easy" to “grade-appropriate" levels (Muralidharan:2019).
In this intervention, the lottery winners (314 students, $D_i=1$) were assigned to the treatment group, and the losers (305 students, $D_i=0$) were assigned to the control group. $S_{di}$ denotes whether student $i$ takes the endline test when $D_i=d$. Note that we define the outcome in interest ($Y_{1i}^*$ and $Y_{0i}^*$) as the difference between the endline and baseline test scores for ease of interpretation, while the authors used a different indicator when computing the Lee:2009 bounds (Muralidharan:2019).
One problem with the Mindspark intervention was attrition. $15.3\%$ and $10.5\%$ dropped out from the treatment and control groups, respectively. To address the potential endogenous sample selection, Muralidharan:2019 computed Lee:2009 bounds. They assumed that the monotonicity in the selection, or $S_{0i}\geq S_{1i}$ holds almost surely; this direction is opposite to that in the previous sections.
We first compute the Lee:2009 bounds based on this Muralidharan:2019's (Muralidharan:2019) assumption (i.e., $\P{S_{0i}=1 | S_{1i}=1}=1$). The estimated bounds and Imbens_Manski:2004's (Imbens_Manski:2004) CIs are reported in rows (1) and (3) of Table (ref). The CI for the math scores is $(0.153, 0.627)$, and thus, the difference in test scores between the endline and baseline for always-takers is positive. In contrast, the CI for the Hindi scores is $(-0.019, 0.423)$ and covers zero. This may imply that the possibility of no effect cannot be rejected.
However, this may perhaps be due to the “extreme" trimming of the Lee:2009 bounds. To address this point, we estimate bounds under the symmetry assumption, whose validity is discussed in the subsequent subsection. Now, $1 = \vartheta_L \geq \vartheta_F$ is satisfied, and thus, we can perform the estimation/inference procedure outlined in Section (ref). The estimation results are shown in rows (2) and (4) of Table (ref). The obtained bounds are sharper than the Lee:2009 bounds. The lengths of the bounds for math and Hindi scores are around $70\%$ shorter. The CIs are also tightened; their lengths are approximately $29\%$ and $36\%$ shorter than those of the Lee:2009 bounds. Consequently, the CI for the Hindi scores does not cover zero, suggesting a positive treatment effect under monotonicity and symmetry. These results showcase the possible usefulness of the new bounds obtained in Theorem (ref).
So far, we have considered bounds when the monotonicity exactly holds. However, a possible concern is the validity of this monotonicity assumption. Recall that Muralidharan:2019 assumed that if a treated student took the end-line test, they must have taken the test if they were in the control group. This assumption might be justified because the untreated students “were told that they would be provided free access" to the programme after the end of the experiment if they took the end-line test (Muralidharan:2019). However, it would not be surprising if some students with $S_{1i}=1$ dropped out when not in treatment. For example, students who did not win the lottery might have sought alternative educational resources (e.g., textbooks, other tutoring services) and then might no longer be interested in participating in the Mindspark programme by the time of the endline test. Alternatively, not participating in the Mindspark intervention could cause parents to overlook the endline test, potentially leading to scheduling conflicts (e.g., travel, shopping, assisting with parents' work) on the test day, thereby preventing their children from attending.
Hence, we compute the bounds under stochastic monotonicity. Considering the incentive provided to the lottery losers, we assume $\vartheta_L=0.95$. The results are summarised in Table (ref).
The results for the math scores are reported in rows (1)--(3). The treatment effect is still positive even if the monotonicity assumption does not exactly hold, which suggests the robustness of the positive effect of the Mindspark intervention on the math scores. The results for Hindi are in contrast. Although the bounds and CIs in rows (5)--(6) are much tighter than those without symmetry in row (4), the CIs for the Hindi scores under stochastic monotonicity cover zero. This means that the treatment effect can be zero if we allow the monotonicity to be only slightly violated. Therefore, the validity of the exact monotonicity will be important to assess whether the Mindspark intervention improved the students' Hindi scores. This illustrates the importance of relaxing monotonicity and examining the sensitivity of the results under the exact monotonicity. The stochastic monotonicity assumption is useful for such an objective.
Below, we provide an illustrative discussion of the validity of our assumptions.
In the context of the Mindspark intervention, the mean-median coincidence assumption could be reasonable. As noted before, the main instruction provided via Mindspark software is automatically personalised. Besides, the tests include a wide range of questions from easy to difficult, designed to measure the various improvements of students. Considering these points, we may assume that the improvement in test scores is similar across students, and the observed differences may stem from stochastic reasons, such as compatibility with the questions. Then, we may assume $Y_{1i}^* = t + \varepsilon_i$, where $t\in\mathbb{R}$ and $\E{\varepsilon_i} = \mathrm{Median}[\varepsilon_i] = 0$ such as the Gaussian error, which implies Assumption (ref). This certain homogeneity assumption is consistent with Muralidharan:2019, wherein they found limited evidence of heterogeneity in students' progress by their initial learning level, gender, and socioeconomic status.
We assess the sharpness of the lower bound on Hindi scores under symmetry (Table (ref) (5)--(6)), which may be the most intriguing aspect, by visually verifying Assumption (ref). Let $F_0$ and $f_y^0$ be the distribution and density functions, respectively, of $Y_i$ conditional on $D_i=0$ and $S_i=1$; and $y^0_r = F_0^{-1}(r)$. In Figure (ref), we show the kernel density estimates of $f_y^0$, 99% point-wise robust bias-corrected (RBC) CIs of calonico2018effect, and their folded ones at $\hat{y}^0_{1-\hat{\vartheta}\hat{q}_0 /2}$, which is indicated by a dotted vertical line. We find no strong evidence suggesting Assumption (ref) is refuted; the folded line is always below the original density estimates, and CIs do not indicate that the former is above the latter. Hence, the sharpness may not be rejected.
When the upper bound in (ref) is too small, it may suggest the implausibility of Assumption (ref)-(i). As the direction of the monotonicity is opposite in our example, we compute the empirical upper bound on $\P{S_{0i} = 1 | S_{1i} = 0}$, finding it to be $0.56$, which is not unreasonably small. Assumption (ref)-(ii) may be justified by the premise that most of the students with $S_{1i}=1$ are motivated and would likely take the endline test to participate in the program if they lose the lottery, and that the attrition resulting from the previously mentioned taste shock or overlapping schedules can be regarded as a random shock.
This study considered the partial identification of Lee:2009's general sample selection model under a stochastic version of the monotonicity and symmetry assumptions. The former assumption was introduced to robustify the Lee:2009 bounds to the small deviation from the exact monotonicity, while the latter assumption can be used to tighten the Lee:2009 bounds, which may be empirically wide. The obtained bounds are root-$n$ consistently estimable, making it easy to perform estimation and inference in practice.
Extensive literature has applied the results of Lee:2009 or Horowitz_Manski:1995. Our proposed stochastic monotonicity assumption can be applied in many such contexts. As an important example, we treat Bartalotti_etal:2023 in the Online Appendix. Furthermore, the simplicity of such an extension will justify theoretical studies based on the monotonicity assumption.