Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
44,821 characters · 11 sections · 24 citation commands
Extreme Quantile Treatment Effects under Endogeneity: Evaluating Policy Effects for the Most Vulnerable Individuals
The quantile treatment effect (QTE) is a widely adopted concept in empirical research for quantifying heterogeneous treatment effects. Various methods have been proposed to identify the QTE in the presence of endogeneity. Popular approaches include instrumental variables designs abadie2002instrumental,chernozhukov2005iv, changes-in-changes designs athey2006identification, and regression discontinuity designs frandsen2012quantile, among others. Their identification results are often accompanied by corresponding methods for estimation and inference.
However, existing methods of estimation and inference are predominantly developed for intermediate quantiles, leaving a gap in the literature regarding estimation and inference for extreme quantiles, such as those at the very low or very high ends of the distribution. This is a notable limitation from practical perspectives, as policymakers are often particularly interested in the effectiveness of policies for extreme subpopulations, such as individuals living in extreme poverty or those with severe health conditions.
Two studies address this issue in specific contexts: zhang2018extremal focuses on the observed unconfoundedness design; and sasaki2024extreme focus on changes-in-changes design. However, extreme QTEs under other endogeneity designs remain unexamined. This paper aims to fill this gap by introducing a novel method for the estimation and inference of extreme QTEs across a broad class of endogeneity scenarios and research designs.
Consider situations where the distributions, $F_1$ and $F_0$, of the potential outcomes under treatment and no treatment, respectively, are identifiable. In these cases, the QTE is quantified by $F_1^{-1}(q) - F_0^{-1}(q) $ for $q \in (0,1)$, where $F_j^{-1}$ denotes the left inverse of $F_j$ for each $j \in \{0,1\}$. Our proposed method is applicable in situations where $F_1$ and $F_0$ can be identified and estimated by $\hat{F}_1$ and $\hat{F}_0$, respectively, with $r_n(\hat{F}_j(\cdot) - F_j(\cdot)), \, j \in \{0,1\} $, weakly converging to a Gaussian process at some rate $r_n$. All the aforementioned examples, along with many others, satisfy this requirement, making our proposed method applicable to a broad class of designs.
Instead of directly using $\hat{F}_1^{-1}(q) - \hat{F}_0^{-1}(q)$ as the estimator, we first estimate the tail indices of $F_1$ and $F_0$, and then obtain estimates of the QTE through Pareto approximation, which is valid under the regular variation condition. This approach enables us to robustly estimate the QTE even in regions where there are few or even no observations in the tails. Subsequently, we propose a method of subsampling inference for the extreme QTE based on this estimation strategy.
We demonstrate the applicability of our method through two leading examples under endogeneity: the instrumental variables design abadie2003semiparametric and the regression discontinuity design frandsen2012quantile. Our simulation studies reveal that the proposed method performs robustly in both scenarios.
We present an empirical application of the proposed method to evaluate the effects of job training provided by the Job Training Partnership Act (JTPA). Previous literature has highlighted the positive QTE of this program in the intermediate quantiles. However, our study of the extremely low quantiles reveals rather negative treatment effects, which are statistically significant. This finding suggests that the program may have adverse outcomes for the most disadvantaged individuals.
Literature: There is a substantial body of literature in statistics and econometrics addressing extreme quantiles -- see the book by de2006extreme. chernozhukov2005extremal proposes methods for estimation and inference in quantile regressions, focusing on the tails of the outcome distribution. chernozhukov2011inference introduce a subsampling approach for inference in these settings. For quantile treatment effects (QTEs), d2018extremal consider a sample selection model, exploiting tail independence in selection. zhang2018extremal analyzes estimation and inference for extreme QTEs within a broad class of observational data characterized by observational unconfoundedness. To our knowledge, only one paper examines extreme QTEs under endogeneity: sasaki2024extreme investigate this within the changes-in-changes (CIC) framework. Their method capitalizes on the unique features of the identifying formula in the CIC, but it does not generalize to other models with endogeneity, such as those using instrumental variables or regression discontinuity designs, which are the focus of this paper. See chernozhukov2017extremal for a survey.
Let $Y_0$ and $Y_1$ denote the potential outcomes without and with treatment, respectively. Let $F_{j}\left( \cdot \right)$ denote a suitable (conditional) CDF of $Y_{j}$ of interest for each $j \in \{0,1\}$. Suppose that $\beta_j\left( \cdot \right)$ identifies $F_j\left( \cdot \right)$, i.e.,
In Examples (ref) and (ref) below, we provide a couple of examples of $\beta_j$ drawn from the literature on identification. For each of these examples, there is a plug-in analog estimator $\hat\beta_j$ of $\beta_j$.
For intermediate quantiles $q$, the $q$-th quantile treatment effect,
can be estimated by the sample counterpart $\hat\beta_{1}^{-1}\left( q \right) - \hat\beta_{0}^{-1}\left( q \right), $ where $\hat\beta_{j}^{-1}(\cdot)$ denotes the left inverse of $\hat\beta_{j}(\cdot)$. On the other hand, for extreme quantiles where $ q $ is close to zero or one, this na\"ive estimator behaves poorly.
Our goal is estimate the extreme quantile $F_{j}^{-1}\left( q \right)$ for $ q $ close to zero or one. Without loss of generality, we consider the right tail, that is, $q \rightarrow 1$. The extreme quantiles at the left tail can be analyzed by symmetric arguments.
By imposing regularly varying tail on $F_{j}$, we show that we can robustly estimate the extreme quantiles $F_{j}^{-1}\left(q \right)$. Consequently, the extreme QTE is estimated precisely. To this goal, we make the following assumption on the estimator $\hat\beta_{j}(\cdot)$ of $\beta_{j}\left( \cdot \right) = F_{j}\left( \cdot \right).$
Assumption (ref) can be and has been shown in the literature to be satisfied under a variety of settings, as illustrated in the two examples below.
Given our focus on the tail, we also need the following additional condition to regularize the tail of the CDF estimator $\hat{\beta}_j(\cdot)$.
Assumption (ref) is arguably mild. Compared with Assumption (ref), it essentially assumes that the normalized t-statistic is uniformly bounded in probability. We present this assumption in its current form for generality, and primitive conditions can be derived case-by-case.
We now move on to learn about the tail features of $F_{j}\left( \cdot \right) $. To this end, we assume that $F_{j}\left( \cdot \right) $ is regularly varying (RV) at infinity. We first provide an intuitive illustration to be followed by a formal theory. Specifically, suppose that $F_j(\cdot)$ satisfies
for any $y>0$, where the constant $\alpha _{j}$ is called the Pareto exponent, which characterizes the tail heaviness of $F_{j}\left( \cdot \right) $. The RV condition allows for the approximation
for all values $y>y_{\min,j}$ for sufficiently large $y_{\min,j}$. In practice, $y_{\min,j}$ plays the role of a tuning parameter and we will discuss the choice of $y_{\min,j} $ later.
Since $\hat{\beta}_{j}\left( \cdot \right) $ consistently estimates $F_{j}\left( \cdot \right) $, we can estimate $\alpha _{j}$ by fitting $\hat{\beta}_{j}\left( \cdot \right) $ with the Pareto distribution. By (ref) and (ref), we have
for $y>y_{\min,j}$. Integrating both sides with respect to $y$ motivates the estimator
of the Pareto exponent $\alpha _{j}$, where $w\left( \cdot \right) $ is some weighting function. This weighting function serves as regularization so that integrands in the numerator and the denominator are integrable.
It is natural to choose $w_j\left( \cdot \right) $ as the true density of the Pareto distribution with exponent $\alpha _{j}$, that is, $w_j\left( y\right) ={y^{-\alpha _{j}-1}}/{y_{\min,j}^{-\alpha _{j}}}.$ However, this is infeasible as $\alpha _{j}$ is unknown. As an alternative, we can choose any $\omega >0$ and set
In this case, the denominator of (ref) reduces to
To derive the asymptotic distribution of $\hat{\alpha}_{j}$, a second-order approximation is inevitable besides (ref). Specifically, we impose the following condition to formally characterize (ref).
First, we remark that this assumption is mild and satisfied by all the commonly used families of heavy-tailed distributions such as Student-t, F, Cauchy, etc. In particular, for the standard Student-t distribution with $v$ degrees of freedom, it holds with $\alpha _{j}=v$ and $\rho _{j}=2$. See, for example, de2006extreme.
Second, Assumption (ref) is used to control the leading bias of the estimator (ref). It is, therefore, analogous to the condition of bounded second (or higher) order derivatives in the context of kernel estimation. The parameters, $d_j$ and $\rho_j$, measure deviation of $F_j(\cdot)$ away from the reference Pareto distribution with exponent $\alpha_j$.
To obtain limiting distribution for the estimator $\hat{\alpha}_j$ around a pseudo true value, it suffices to choose $y_{\min,j}$ according to the following assumption.
Assumption (ref) concerns limiting behaviors of the regularization parameter $y_{\min,j}$ as $n \rightarrow \infty$. It controls the bias due to the deviation of $F_j$ from the benchmark Pareto distribution, which converges at the rate of $y^{-\rho_j}$ as $y \rightarrow \infty$.
The following theorem establishes the asymptotic distribution of $\hat{\alpha}_j$.
A proof is found in Appendix (ref).
Theorem (ref) derives the asymptotic normality of $\hat{\alpha}_{j}$. The $A_{jn}$ and $B_{jn}$ terms respectively characterize the variance and the bias. One can control the balance between the bias and variance via the tuning parameter $y_{\min,j}$. A larger $y_{\min,j}$ leads to smaller bias and larger variance. The theoretically optimal choice $y_{\min,j}^{\ast}$ requires $A_{jn}^{-1/2}\times B_{jn} \sim 1$. If we further assume that the variance $\Xi_j(y, y)$ of $Z_j(y)$ is approximately proportional to $y^{-\kappa_j}$ for some $\kappa_j>0$ in the limit as $y \rightarrow \infty$, then we can derive the optimal choice as $y_{\min,j}^{\ast} \sim r_{n}^{1/\left( \alpha _{j}-\kappa_{j}/2-1/2+\rho_{j}\right) }$. Since the second-order parameter $\rho _{j}$ is challenging to estimate in practice and the bias should vanish faster for the purpose of inference, we recommend setting $y_{\min,j}^{\ast} /r_{n}^{1/\left(\alpha_{j}-\kappa _{j}/2-1/2+\rho _{j}\right) }\rightarrow \infty $ so that $A_{jn}^{-1/2}\times B_{jn}=o(1)$. In other words, $y_{\min,j}^{\ast}$ is an undersmoothing choice. With this said, such a rule must be subject to a feasible choice of $y_{\min,j}$ satisfying Assumption (ref), i.e., $r_{n}^{1/\left(\alpha_{j}-\kappa _{j}/2-1/2+\rho _{j}\right) } \prec y_{\min,j} \prec \tilde{r}_n^{1/(\alpha_j - \kappa_{j}/2)}$, which further requires that $\rho_j>1/2$. Note that a smaller $\rho_j$ (closer to zero) implies a larger deviation from the benchmark Pareto distribution. The implicit requirement, $\rho_j>1/2$, for a feasible choice of $y_{\min,j}$, therefore, rules out huge deviations away (i.e., $\rho_j$ close to zero) from the benchmark Pareto distribution.\footnote{Student-t distribution entails $\rho_j = 2$ and hence satisfies this restriction.}
To conduct valid inference, we still need to know the constant $c_{j}$ in $A_{jn}$ as specified in Assumption (ref), which can be consistently estimated by
However, the asymptotic property of such estimator is unknown, and its finite sample performance might not be satisfactory. In addition, the asymptotic variance of the quantile treatment effect estimator becomes more complicated, as shown in the Corollary (ref) below. Therefore, we propose an alternative inference method based on subsampling and establishes its asymptotic validity in Section (ref). Conveniently, this method does not require estimating $c_j$ or other second-order parameters.
Given the estimator of $\alpha_j$ and its asymptotic behavior, we now proceed to estimate the extreme QTE. Section (ref) presents our extreme QTE estimator and derives its asymptotic distribution. However, the asymptotic distribution involves complicated variance expressions that vary across different applications (e.g., between Examples (ref) and (ref)). To overcome these challenges, Section (ref) proposes to conduct inference for the QTE based on subsampling as a general recipe in practice. A theoretical guarantee will be provided for this proposal.
Our extreme QTE estimator can be written as
The corollary below establishes its asymptotic distribution. Define
for $j=0,1,$ where
A proof is found in Appendix (ref).
The QTE involves the estimators $\hat{F}^{-1}_1(\cdot)$ and $\hat{F}^{-1}_0(\cdot)$, which converge at different rates respectively denoted by $\lambda_{1n}$ and $\lambda_{0n}$. Therefore, the convergence rate of $\widehat{QTE}$ is determined by the slower one, i.e., $\underline{\lambda }_{n}=\min \{\lambda _{1n},\lambda _{0n}\}$. Also, the asymptotic variance $\Omega _{Q}$ involves $\lim_{n\rightarrow\infty} \underline{\lambda }_{n}/\lambda_{1n}$ and $\lim_{n\rightarrow\infty} \underline{\lambda }_{n}/\lambda_{0n}$, which are bounded between zero and one, with the larger one being exactly one. In addition, it involves the covariance between $\hat{\alpha}_1$ and $\hat{\alpha}_0$, whose expression is complicated. We present the expression in the proof for brevity because we will not use it in practice thanks to the subsampling method presented in the following subsection.
The asymptotic distribution of the QTE estimator is complex, making direct inference based on the analytically estimated $\Omega _{Q}$ challenging to implement. Additionally, the analytic expression for this distribution can vary across different contexts (e.g., between Examples (ref) and (ref)), and no universal approach can be recommended for all applications. Moreover, previous studies have demonstrated that the bootstrap method performs poorly in estimating tail features hall1990using, bickel2008choice. To address these challenges, we propose using subsampling as an alternative and establish its asymptotic validity.
Specifically, consider a sequence of subsample sizes $b=b_{n}$ that grows with $b/n\rightarrow 0$ as $n\rightarrow \infty $. Let $B_{n}=\binom{n}{b}$ denote the total possible number of subsamples of size $b$. For a given $b$ and $t\in \{1,...,b_{n}\}$, let $S_{t}\subset \{1,...,n\}$ be one of the $B_{n}$ subsamples of the individual indices with $|S_{t}|=b$, and define the estimator (ref) based on the $t$-th subsample as $\hat{\alpha}_{j}^{t}$ for $j=0,1$. Furthermore, denote the corresponding QTE estimator based on the $t$-th subsample as
Note that the estimator $\hat{\beta}_{j}\left( \cdot \right) $ for the counterfactual CDF still uses the original full sample.
Given $B_{n}$ subsampling estimates, we propose to use the empirical CDF
to approximate the CDF of $\widehat{QTE}\left( q\right) $, denoted by $L^{\ast }\left( \cdot \right) $. The following corollary provides a theoretical guarantee that this subsampling approximation works asymptotically.
A proof is found in Appendix (ref).
In light of this theoretical result, we shall use the subsampling in the subsequent numerical and empirical analyses.
In this section, we use simulated data to analyze the finite-sample performance of our proposed method. We revisit Examples (ref) and (ref) for simulation designs, and present them in Sections (ref) and (ref), respectively.
Recall Example (ref) from Section (ref). In its setting, we generate independent copies of $(Y, D, Z, X' )'$ as follows. An individual is an always taker, a complier, or a never taker, indicated by $\mathbbm{1}_A$, $\mathbbm{1}_C$, and $\mathbbm{1}_N$, respectively. Each type emerges with the probability of $1/3$. We generate $10$-dimensional exogenous covariates $ X \sim \mathcal{N}(0,I_{10}), $ where $I_{k}$ denotes the $k\times k$ identity matrix. We in turn generate the instrument $Z \sim \text{Bernoulli}(\Lambda(X'\gamma))$, where $\Lambda$ is the logistic link function defined by $\Lambda(u) = \exp^u / (1+\exp^u)$ and $\gamma = (0.1,...,0.1)'$. The treatment selection is generated by
The potential outcomes are given by
where $TE_A=2$, $TE_C=1$, and $TE_N=0$ are the treatment effects for always takers, compliers, and never takers, respectively, and $t_{0,A}$, $t_{1,A}$, $t_{0,A}$, $t_{1,A}$, $t_{0,A}$ and $t_{1,A}$ are independently drawn from the Student-t distribution with 10 degrees of freedom. Finally, the observed outcome is generated by $$ Y = (1-D)Y_0 + D Y_1. $$ In this manner we generate $n \in \{2500,5000,10000\}$ independent copies of $(Y,D,Z,X')'$.
Since lower extreme quantiles are often of policy interest, we focus on the QTE at $q \in \{0.01,...,0.05\}$. We thus flip the sign of $Y$, and obtain the negative treatment effects. The tuning parameter $y_{\min,j}$ is set as the 97.5-th percentile of $\hat\beta_j$. We compute simulation statistics, such as the bias, standard deviation, root mean square error, and the 95% coverage frequency based on 10,000 Monte Carlo iterations. Table (ref) summarizes the simulation results.
Observe that the RMSE diminishes as the sample size $n$ increases for each quantile. Also, observe that the coverage frequency approach gets closer to the nominal probability of 95% as the sample size $n$ increases for each quantile. As the sample size increase, the coverage become more accurate for the extreme quantiles.
Recall Example (ref) from Section (ref). In its setting, we generate independent copies of $(Y,D,R)'$ as follows. As in the previous subsection, an individual is an always taker, a complier, or a never taker, indicated by $\mathbbm{1}_A$, $\mathbbm{1}_C$, and $\mathbbm{1}_N$, respectively. Each type emerges with the probability of $1/3$. We generate the running variable $ R \sim \mathcal{N}(0,1). $ We in turn generate the treatment selection by
where $\widetilde D(R)\sim \text{Bernoulli}(1/3 + \mathbbm{1}\{R>0\}/3)$. The potential outcomes are given by
where $TE_A=2$, $TE_C=1$, and $TE_N=0$ are the treatment effects for always takers, compliers, and never takers, respectively, and $t_{0,A}$, $t_{1,A}$, $t_{0,A}$, $t_{1,A}$, $t_{0,A}$ and $t_{1,A}$ are independently drawn from the Student-t distribution with 10 degrees of freedom. Finally, the observed outcome is generated by $$ Y = (1-D)Y_0 + D Y_1. $$ In this manner, we generate $n \in \{2500,5000,10000\}$ independent copies of $(Y,D,R)'$.
As in the previous example, we focus on the QTE at $q \in \{0.01,...,0.05\}$. We thus flip the sign of $Y$, and obtain the negative treatment effects. The tuning parameter $y_{\min,j}$ is set as the 97.5-th percentile of $\hat\beta_j$. For the limits of the conditional expectation function, we use the one-sided Epanechnikov kernel with the rule-of-thumb bandwidth $h = \hat\sigma_R n^{-1/5}$, where $\hat\sigma_R^2$ denotes the sample variance of $R$. We compute simulation statistics, such as the bias, standard deviation, root mean square error, and the 95% coverage frequency based on 10,000 Monte Carlo iterations. Table (ref) summarizes the simulation results.
The results are similar to those presented in the previous subsection. Specifically, the RMSE diminishes as the sample size $n$ increases for each quantile. Also, observe that the coverage frequency approach gets closer to the nominal probability of 95% as the sample size $n$ increases for each quantile. As the sample size increase, the coverage become more accurate for the extreme quantiles.
In this section, we present an empirical application of our proposed method of estimation and inference for extreme quantile treatment effects. We revisit the causal inference for the job training provided by the JTPA on participants' earnings and employment outcomes studied by numerous researchers HIST1998. Related to our approach, abadie2002instrumental used the method (Abadie's Kappa) presented in Example (ref) to obtain selection-adjusted quantile regression estimates. We use the same data set as in their empirical application. Instead of following their estimation strategy, however, we use our proposed method of estimation and inference for extreme quantiles.
The data come from the National JTPA Study, which includes information on individuals who were randomly assigned to receive the JTPA training and those who were not. Not all the assigned individuals took the treatment, and hence the compliance is imperfect. Nonetheless, the random assignment serves as a natural instrument for the endogenous treatment. We are interested in the treatment effects for the outcome of log wages. In our analysis, we use the same set of covariates, as well as the outcome, treatment, and instrument as in the original study.
Figure (ref) plots estimates, $\hat\beta_0$ and $\hat\beta_1$, of the distributions, $F_0$ and $F_1$, of the potential outcomes, $Y_0$ and $Y_1$, respectively. Note that, unlike the empirical CDFs, these estimated CDFs do not need to be non-decreasing. Observe that the estimate $\hat\beta_1$ dominates the estimate $\hat\beta_0$ for $q > 0.10$, implying positive quantile treatment effect estimates in these intermediate quantiles. In contrast, the estimate $\hat\beta_0$ dominates the estimate $\hat\beta_1$ for $q < 0.05$, implying negative quantile treatment effect estimates in these extremely low quantiles. However, the existing econometric methods of estimation and inference focusing on intermediate quantiles may not be applicable to analyzing the latter feature of these estimates. Our proposed method can fill this gap.
Before conducting the QTE estimation, we first evaluate whether the Pareto-type tail condition (Assumption (ref)) is a suitable assumption for this dataset. Figure (ref) replicates the estimates, $\hat\beta_0$ (top) and $\hat\beta_1$ (bottom), as reported in Figure (ref). Additionally, Figure (ref) also displays the Pareto fits of the corresponding distributions (depicted by gray lines), which are based on our estimates of the tail exponents, $\hat\alpha_0$ and $\hat\alpha_1$, for $\alpha_0$ and $\alpha_1$, respectively. By comparing the black and gray lines in each of the top and bottom panels, we can observe that the fit appears to be reasonably consistent.
We now estimate the extreme QTE using our proposed method. Figure (ref) shows estimates and 95% confidence intervals of the quantile treatment effects $QTE(q)$ for the extremely low quantiles $q \leq 0.025$. Observe that the treatment effect is significantly negative for each of $q \in [0.019, 0.025]$. Hence, the job training program may exacerbate the labor outcomes for extremely disadvantaged individuals. For even lower quantiles $q \in [0.002, 0.017]$, the point estimates are negative but the 95% confidence intervals contain zero.
Policymakers are frequently concerned with understanding the impact of interventions on the most vulnerable populations, particularly those at the extreme lower end of the economic spectrum. Their objectives often include improving the economic well-being of the most disadvantaged individuals, who are typically represented by the lowest quantiles in a distribution.
However, traditional econometric methods have predominantly focused on intermediate quantiles, which, while useful, may not fully capture the nuances or the extent of impact on the extreme ends of the distribution. These methods, therefore, have inherent limitations when it comes to assessing treatment effects on the lowest quantiles, potentially overlooking critical insights that are essential for crafting effective policy interventions for the most at-risk groups.
In this empirical application, we demonstrate that an analysis centered on extremely low quantiles can yield surprising and sometimes counterintuitive results. Specifically, our study reveals that treatment effects, which policymakers might assume to be uniformly positive, could in fact be negative for these disadvantaged groups. This finding underscores the importance of employing methodologies that specifically target these extreme quantiles to ensure that policy decisions do not inadvertently exacerbate the conditions of those they are intended to help.
In this paper, we introduce a new method for the estimation and inference of extreme QTEs, which are crucial in understanding the impacts of interventions on the most vulnerable or advantaged individuals within a population. The proposed method is versatile and can be applied to a wide range of empirical research designs that have been developed for causal inference in the presence of endogeneity. Leading examples are the instrumental variables design and the regression discontinuity design, which we focused on in this paper. Our method leverages regular variation and subsampling. The combination of these two features ensures robust performance even in the sparsest regions of the distribution, where traditional methods might fail or produce unreliable results. Our theoretical analysis is complemented by extensive simulation studies, which demonstrate that our method performs well under the scenarios of both the instrumental variables and regression discontinuity designs.
To illustrate the practical utility of our method, we apply it to a real-world policy evaluation: the assessment of job training programs provided by the JTPA. Previous studies using traditional instrumental variables estimation methods of the QTEs have largely focused on intermediate quantiles, finding positive treatment effects for individuals at the middle of the distribution. However, our method reveals a different picture when we extend the analysis to the extreme lower quantiles, representing the most disadvantaged individuals. Specifically, we find significantly negative QTEs for these individuals, suggesting that the job training program may not only fail to benefit the most vulnerable but might even have adverse effects on them. This finding contrasts sharply with the prevailing narrative in the literature, highlighting the importance of examining extreme quantiles to fully understand the impact of policy interventions across the entire distribution of outcomes.
Finally, we conclude by underscoring the theoretical contributions of this paper. Most existing methods for investigating tail features, including extreme quantiles, rely on hill1975's (hill1975) estimator or its numerous variants. See Hill2010,Hill2015, among many others. These methods are effective when the sample is drawn directly from the target distribution, $F_j(\cdot)$, as is the case in the changes-in-changes design, where sasaki2024extreme applied these approaches. However, such straightforward methods are inadequate for other scenarios involving endogeneity, such as the instrumental variables design and the regression discontinuity design, where each observation cannot be uniquely associated with a specific distribution $F_j$, as illustrated in Examples (ref) and (ref). This limitation necessitated the development of a novel theoretical framework to study the tails of the distribution $F_j$ without direct observation of samples from $F_j$. To our knowledge, this aspect of theoretical development is unprecedented in the literature, and it was driven by the need to address the challenges posed by specific empirical approaches that are widely used.