Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
55,049 characters · 16 sections · 59 citation commands
Causal Interpretation of Regressions With Ranks
\begingroup \footnote{The author would like to thank Raj Chetty for the helpful discussion.} \addtocounter{footnote}{-1} \endgroup
Rank transformations are often applied to variables without natural units, such as test scores krueger1999experimental), to enhance cross-outcome comparability, or variables with heavy tails, such as income chetty2014land,deutscher2023measuring, mogstad2024inference, to avoid excessive influences of extreme observations. In both cases, it was found that the rank transformation tends to reduce sensitivity to specifications dahl2008association. A recent survey by chetverikov2023inference found 119 papers that contain rank regressions in the top Economics journals published since 2013.
Despite the popularity of rank transformations in applied research, the causal interpretation of regression coefficients remains unclear when dependent or independent variables are transformed into ranks. For example, chetty2020race obtained an estimate of the relative mobility, defined as the coefficient of the parents' income rank in the regression of the child's income rank, of $0.35$ using the full US sample that pools all races. This was interpreted as “a 10 percentile increase in parents’ rank is associated with a 3.5 percentile increase in children’s rank on average”. It is known that this coefficient is equivalent to the Spearman's rank correlation coefficient spearman1904proof between the parents' and child's incomes, assuming that the income distribution is continuous chetverikov2023inference. However, it is unclear if the Spearman's correlation measures an interventional effect in any thought experiment where parents' incomes are randomized or exposed to exogeneous shocks. Ideally, we would want to interpret the coefficient through the distributions of the child's counterfactual incomes had the parents' income been a certain amount.
For regressions with ranks only in one side, even the effective estimands have yet been understood. For example, krueger1999experimental regressed the percentile ranks of test scores on the assignment of class types (i.e., a small class, a regular class, or a regular class with teacher aid) along with other controls using the Tennessee STAR data boyd2007project. Through the OLS estimate he found that “the gap in average performance is about 5 percentile points in kindergarten, 8.6 points in first grade, and 5–6 points in second and third grade.” While this suggests a positive effect of smaller class sizes, it is harder to interpret the effect size. Ideally, we would want to express this coefficient through the distributions of potential test scores had the student been assigned into each class type. The interpretation can be further complicated for more sophisticated econometric methods such as the two-stage least squares, which is also applied to analyze the effect of small classes on STAR data to address non-compliance krueger1999experimental.
In this paper, we derive the effective estimands as well as their causal interpretations for a suite of econometric methods including the ordinary least squares (OLS), two-stage least squares (2SLS), difference-in-differences (DiD), and regression discontinuity designs (RDD) under a variety of standard assumptions in the econometrics literature. Unlike chetverikov2023inference, we do not discuss statistical inference for these estimators and many of their results can be directly applied here. Moreover, we focus on continuous outcomes and do not study the effect of tie-breaking for discrete outcomes as in chetverikov2023inference and chetverikov2024csranks. Our paper also differs from the literature that introduces inference methods for ranks of true parameters from multiple heterogeneous units or populations bazylik2021finite,mogstad2024inference,chetverikov2024csranks in that we focus on the causal interpretations of common econometric methods with sample ranks under the superpopulation framework.
We start by introducing a primitive causal estimand which we call the Rank Average Treatment Effect (rank-ATE). The estimand had been studied in the statistics and machine learning literature, though in very different contexts. Akin to the standard ATE, it compares the distributions of outcomes under two treatment conditions, though not through the mean difference. We prove that the rank-ATE shares many desirable properties with the ATE and they have the same signs in a variety of scenarios.
For a randomized experiment with a binary treatment, we prove that the OLS with outcomes ranked based on the entire sample identifies the rank-ATE. This result remains if extra covariates that are not affected by the treatment are included. More surprisingly, we show that the effective estimand remains the same if the outcomes are ranked based solely on the treatment (or control) units. We call this property reference robustness and show that it is only satisfied by the OLS but not other aforementioned econometric methods. When the assignment is confounded under the selection on observables (or strong ignorability) assumption, we show that the effective estimand is a weighted average of covariate-specific rank-ATE, in a similar spirit to angrist1995estimating and borusyak2024negative for linear regressions.
Moving beyond the binary treatment, we consider general multi-valued or continuous treatments and allow for transformation of the treatment variable in the regression. This includes the rank-rank regression dahl2008association, chetty2014land, chetverikov2023inference and special cases of binscatter regression cattaneo2024binscatter as special cases. When the treatment is randomized, we prove that the effective OLS estimand is a weighted average of rank-ATEs that compare the potential outcomes at each level of treatment. Since the weights depend on the scale of the transformed treatment, the estimates for different treatment variables are generally incomparable. We propose a data-driven normalization procedure that restrict the estimand in the same range regardless of the choice of variables.
For other econometric methods, the effective estimands and causal interpretations are more delicate. To avoid mathematical complications, we focus on the case where the outcomes are transformed into ranks and the treatment variable is binary. We briefly summarize our findings below.
We conclude the paper by discussing a limitation of the rank-ATE: unlike the standard ATE, it cannot be expressed as an expected individual-level treatment effect. chen2023logs argued that the latter is a desirable property for the purpose of causal interpretations. Motivated by this concern, we propose an alternative partially identified estimand. The sharp identified set can be estimated by the standard plug-in methods fan2010sharp or the robust dual methods ji2023model.
Suppose we have $n$ units and, for unit $i\in \{1, \ldots, n\}$, $Y_i$ denotes the outcome of interest, $W_i$ denotes the treatment variable which can be binary or continuous, and $X_i\in \mathbb{R}^{d}$ denotes the vector of covariates. Let $R^{{\scriptscriptstyle Y}}_i$ denote the rank of $Y_i$ among $\{Y_1, \ldots, Y_n\}$, i.e.,
Let $\hat{F}_{Y,n}$ be the empirical cumulative distribution function (CDF) of $\{Y_1, \ldots, Y_n\}$, i.e., \[\hat{F}_{Y,n}(y) = \frac{1}{n}\sum_{j=1}^{n}I(Y_j \le y).\] Then we can rewrite $R^{{\scriptscriptstyle Y}}_i$ as
Throughout the note we assume that $(Y_i, W_i, X_i)$ are i.i.d. where $F_{Y}$ denotes the marginal distribution of $Y_i$, $F_{Y\mid w}$ the conditional distribution of $Y_i$ given $W_i = w$, and $F_{Y\mid w, x}$ the conditional distribution of $Y_i$ given $W_i = w$ and $X_i = x$. When no confusion can arise, we will use $(Y, W, X)$ to denote a generic independent draw from the same distribution as $(Y_i, W_i, X_i)$. Throughout we adopt the standard probability notation $O(\cdot), o(\cdot), O_{\mathbb{P}}(\cdot),$ $o_{\mathbb{P}}(\cdot)$.
An immediate yet crucial implication of the celebrated Dvoretzky–Kiefer–Wolfowitz inequality dvoretzky1956asymptotic, massart1990tight is that
This suggests, and we will make it rigorous, that $R^{{\scriptscriptstyle Y}}_i$ can be replaced by $F_Y(Y_i)$ when the goal is to prove consistency of a parameter estimate. For inferential purposes like hypothesis testing or confidence intervals, more sophiscated techniques would be involved chetverikov2023inference.
Let $F_1$ and $F_0$ be the distributions of a scalar outcome under two treatment conditions. Most point identified treatment effects are functionals of $(F_1, F_0)$ that measure the distributional discrepancy through certain summaries.\footnote{Causal effects as functionals of the joint distribution of potential outcomes are only partially identified without further assumptions since the potential outcomes are never observed simultaneously fan2010sharp,ji2023model.} For example, the standard ATE is the mean difference and the logarithm of the ratio ATE (or the Poisson regression estimand silva2006log,wooldridge2010econometric) is the difference of log-means. Here, we define the rank-ATE $\tau_{\mathrm{r}}(F_1, F_0)$ as follows:
The rank-ATE always takes values in $[-1/2, 1/2]$. Clearly, $\tau_{\mathrm{r}}(F_1, F_0)$ measures the probability that a randomly drawn sample from $F_1$ is larger than or equal to a randomly drawn sample from $F_0$. Unlike the ATE that requires the existence of finite first moments, the rank-ATE is always well-defined. An equivalent expression for $\tau_{\mathrm{r}}(F_1, F_0)$ is given by
We illustrate this estimand through the following examples.
The rank-ATE has the following nice properties.
We call the last property partial additivity because it only holds when the second argument of the first two terms is in the class $\{\zeta F_1 + (1 - \zeta)F_0: \zeta \in [0, 1]\}$. Note that the ATE satisfies all but invariance, which is a desirable property especially when the variables do not have a natural unit or it is unclear what transformation is most scientifically meaningful roth2023parallel. The first three properties can be proved directly by definition and the partial additivity is proved in Appendix (ref).
The rank-ATE has been studied in different contexts. In particular, it is the effective estimand of the Mann-Whitney test, also known as the Wilcoxon rank-sum test vermeulen2015increasing. It was also shown to be equivalent to the Area Under the Receiver Operating Characteristic curve (AUROC, also known as AUC), one of the most important measures in evaluating classification algorithms hanley1982meaning.
In this section, we study the case where $W_i$ is binary. Let $(n_1, n_0)$ denote the size of the treated and control groups, respectively, $\pi = \mathbb{P}(W_i = 1)$ the marginal treatment intensity, and $\pi(z) = \mathbb{P}(W_i = 1\mid X_i = z)$ the propensity score. Furthermore, we let $(Y_i(1), Y_i(0))$ denote the pair of potential outcomes for unit $i$. Throughout the section we assume the stable unit treatment value assumption (SUTVA, rubin1974estimating):
Let $F_{Y(1)}$ and $F_{Y(0)}$ denote the marginal distributions of $Y(1)$ and $Y(0)$, respectively. Clearly,
We further assume that $Y(1)$ and $Y(0)$ are continuous variables:
Even if $Y(1)$ and $Y(0)$ are discrete, one can perturb $Y(1)$ and $Y(0)$ by a random noise uniformly distributed on $[-\epsilon, \epsilon]$ for some sufficiently small $\epsilon$. This is equivalent to random tie-breaking in practice considering that all variables are discretized by rounding.
As a warm-up, we consider the rank-OLS estimator without covariates, defined as follows:
Above, the ranks are normalized by $n$ for ease of interpretation, as evidenced later, and the subscript indicates that no covariate is used. We study the property of $\hat{\beta}_{{\scriptscriptstyle\mathrm{nox}}}$ when assignments are completely random as in randomized experiments or quasi-experiments:
Then we can prove the following result that the rank-OLS identifies the rank-ATE.
An intriguing question is how the estimand changes with the reference population for ranking. In pricinple, one could rank each outcome among all units or among treated or control units. Specifically, define two other ranks as follows: \[R^{{\scriptscriptstyle Y}}_{i,w} = \sum_{j:W_j=w}I(Y_j\le Y_i), \quad w = 0, 1,\] and consider the corresponding rank-OLS: \[\hat{\beta}_{{\scriptscriptstyle\mathrm{nox}}, w} = \operatorname*{argmin} \frac{1}{n}\sum_{i=1}^{n}\left(\frac{R^{{\scriptscriptstyle Y}}_{i,w}}{n_w} - \beta_0 - \beta W_i\right)^2, \quad w = 0, 1.\] Note that we use different normalizations because the range of $R^{{\scriptscriptstyle Y}}_{i,1}$ and $R^{{\scriptscriptstyle Y}}_{i,0}$ are different. Surprisingly, both $\hat{\beta}_{{\scriptscriptstyle\mathrm{nox}}, 1}$ and $\hat{\beta}_{{\scriptscriptstyle\mathrm{nox}}, 0}$ are equal to $\hat{\beta}_{{\scriptscriptstyle\mathrm{nox}}}$ up to a small factor.
In practice, covariates are often added into the regression for the purpose of adjusting for confounders or improving statistical efficiency. We consider two versions of rank-OLS, one without interactions and one with interactions between the treatment and covariates:
and
where $\bar{X} = (1/n)\sum_{i=1}^{n}X_i$. The former estimator is widely used in applied work and the latter estimator is much less so. The latter estimator is equivalent to running separate regressions on the treated and control units with centered covariates $X_i - \bar{X}$ and contrasting the intercepts. Our motivation to consider $\hat{\beta}_{{\scriptscriptstyle\mathrm{int}}}$ is from the regression adjustment literature freedman2008regression, imbens2009recent, lin2013agnostic, li2020rerandomization, lei2021regression, negi2021revisiting.
The following result shows that the OLS estimand remains the same under the following mild regularity condition.
In this subsection, we relax the assumption of random assignments and assume the propensity score $\pi(z)$ varies with $z$. To make progress, we assume strong ignorability or selection on observables:
In addition, we assume strict overlap or positivity crump2009dealing, d2021overlap, lei2021distribution:
We state a generic result that shows $\hat{\beta}$ converges to a weighted average of conditional rank-ATEs. The proof is presented in Appendix (ref).
The condition (a) and (b) are similar to the ones in borusyak2024negative. In particular, the condition (b) is the analogue of the classical assumption of saturated specification angrist1995estimating. When $\tilde{\pi}(X)\in [0, 1]$, the effective estimand of $\hat{\beta}$ is a convex average of conditional rank-ATEs.
In this section, we consider a general treatment $W_i \in \mathcal{W} \subset\mathbb{R}$, where $\mathcal{W}$ can be a finite or infinite set. Define the rank of $W_i$ as \[R^{{\scriptscriptstyle W}}_i = \frac{1}{n}\sum_{j=1}^{n}I(W_j\le W_i).\] Let $F_W$ denote the marginal distribution of $W$, $Y(w)$ the potential outcome for dose $w$, and $F_{Y(w)}$ the marginal distribution of $Y(w)$. We make the following assumptions as analogues of Assumptions (ref) and (ref).
We consider the rank-OLS estimator with a transformed treatment:
where $h_n$ is any function that potentially depends on data (and hence is potentially random). This nests a broad class of regressions in applied work. For example, when $h_n$ is the identity mapping, $\hat{\beta}_{h_n}$ recovers $\hat{\beta}$ studied in Section (ref). Another widely-used class of regressions that is nested in (ref) is the rank-rank regressions:
In particular, $\hat{\gamma}_{g_n}$ is identical to $\hat{\beta}_{h_n}$ with $h_n = g_n\circ \hat{F}_{W}$ where $\hat{F}_W$ is the empirical CDF of $W_i$s. When $g_n$ is the identity mapping, (ref) is the standard rank-rank regression chetverikov2023inference. It also nests regressions that dichotomize $W_i$ by choosing $g_n(r) = I(r > 1/2)$ or those that coarsen $W_i$ into quartiles/quintiles/deciles by choosing a stepwise function.
It is known that, when $Y_i$ is continuous, $g_n(r) = r$, and no covariate is included, $\hat{\gamma}_{{\scriptscriptstyle\mathrm{nox}}}$ is equivalent to the Spearman's rank correlation chetverikov2023inference. The estimand can be equivalently expressed as $\mathrm{Cor}(F_W(W), F_Y(Y))$ where $\mathrm{Cor}$ denotes the Pearson correlation coefficient. However, the causal interpretation of this estimand is underexplored. We first derive an expression of the estimand in terms of potential outcomes when $W_i$ is randomly assigned:
To study the limit of $\hat{\beta}_{h_n}$, we impose the following assumption on limited dependence of $h_n$ on data. It is easy to check that the aforementioned examples all satisfy the assumption.
For rank-rank OLS estimators, by (ref), Assumption (ref) so long as \[\sum_{i=1}^{n}\left( g_n(W_i) - g(W_i)\right)^2 = O_\mathbb{P}(1),\] for some non-stochastic function $g$.
The proof is presented in Appendix (ref). Clearly, when $h$ is monotonely increasing, $b(W, \tilde{W})\ge 0$. In this case, Theorem (ref) implies that the effective estimand is a weighted sum of $\tau_{\mathrm{r}}\left( F_{Y(w)}, F_{Y(\tilde{w})}\right)$ for all dose pairs $(w, \tilde{w})$ with $w\ge \tilde{w}$ with non-negative weights.
To further illustrate the estimand, we consider the following examples as extensions of Example (ref) to (ref).
We further provide examples with different $h_n$.
To enable comparison of the OLS estimates across different treatment variables, we need to enforce the estimands to have the same scale. It suffices to ensure that $\beta^{*}_{h}$ is a convex average of pairwise effects $\{\tau_{\mathrm{r}}\left( F_{Y(w)}, F_{Y(\tilde{w})}\right): w > \tilde{w}\}$, in which case $\beta^{*}_{h}\in [-1/2, 1/2]$. To achieve this, we can rescale $h_n$ such that
For any given $h_n$, we can replace it by \[\frac{\mathbb{E}[(h_n(W) - h_n(\tilde{W}))I(W > \tilde{W})]}{\mathbb{E}[(h_n(W) - h_n(\tilde{W}))^2I(W > \tilde{W})]}\cdot h_n.\] Both the numerator and denominator can be estimated by U-statistics:
For example, when $h_n(w) = w$ as in the case of binary treatment, $\hat{\kappa} = 1$. When $h_n(w) = \hat{F}_W(w)$,
Thus, we need to replace $R^{{\scriptscriptstyle W}}_i / n$ by $2R^{{\scriptscriptstyle W}}_i / n$.
To ease exposition and highlight the main point, we limit our discussion to binary treatments and exclude the consideration of covariates. All technical proofs are presented in Appendix (ref).
Let $Z_i$ be a valid binary instrumental variable, $(W_i(1), W_i(0))$ be the potential treatment assignments, and $\{Y_i(w, z): w, z\in \{0, 1\}\}$ be the potential outcomes. We make the standard IV assumptions angrist1996identification.
Further, we let $G_i$ denote the type of the unit $i$ (i.e., always-taker, never-taker, and complier) with \[G_i = \left\{
\right..\] Throughout this secvtion we assume that $(Z_i, W_i(1), W_i(0), Y_i(1), Y_i(0), G_i)$ are i.i.d. and denote by $(Z, W(1), W(0), Y(1), Y(0), G)$ a generic draw. For each $g\in \{\texttt{a}, \texttt{n}, \texttt{c}\}$ and $w\in\{0, 1\}$, let $F_{Y(w)\mid g}$ denote the conditional distribution of $Y(w)$ given $G = g$.
Let $\hat{\beta}_{{\scriptscriptstyle\mathrm{2SLS}}}$, $\hat{\beta}_{{\scriptscriptstyle\mathrm{2SLS}}, 1}$, $\hat{\beta}_{{\scriptscriptstyle\mathrm{2SLS}}, 0}$ denote the 2SLS estimators applied to $R^{{\scriptscriptstyle Y}}_i / n$, $R^{{\scriptscriptstyle Y}}_{i,1}/n_1$, and $R^{{\scriptscriptstyle Y}}_{i,0}/n_0$, respectively. First, we derive the effective estimand of 2SLS applied to different types of ranks.
Theorem (ref) suggests that, unlike the rank-OLS estimator, the rank-2SLS estimator is no longer reference-robust. In addition, none of these estimands are as easily interpretable as the rank-ATE.
Since LATE is the mean difference of potential outcomes for compliers, it is natural to consider the rank-LATE, defined as the rank-ATE for compliers:
The following result shows that rank-LATE can be obtained by ranking the outcomes among the compliers.
Theorem (ref) implies that this new 2SLS estimator identifies the rank-LATE and enjoys reference robustness as the rank-OLS estimator under randomized assignments.
The next result shows that $F_{Y(1)\mid \texttt{c}}$ and $F_{Y(0)\mid \texttt{c}}$ can be identified under no additional assumptions.
For each $w,z\in \{0,1\}$, $F_{wz}$ can be consistently estimated by the empirical CDF of $Y_i$s among the units with $W_i = w, Z_i = z$. Classical IV theory imbens1994identification implies that \[\frac{n_{10}}{n_0}\stackrel{p}{\rightarrow} \pi_{\texttt{a}}, \quad \frac{n_{01}}{n_1}\stackrel{p}{\rightarrow} \pi_{\texttt{n}}, \quad \frac{n_{11}}{n_1} - \frac{n_{10}}{n_0}\stackrel{p}{\rightarrow} \pi_{\texttt{c}}\] where $n_{wz}$ is the number of units with $W_i = w, Z_i = z$ and $n_{z}$ is the number of units with $Z_i = z$.
We focus on the standard DiD setting with one pre-treatment period, indexed by $0$, and one post-treatment period, indexed by $1$. For each $t=0, 1$, let $(Y_{it}(1), Y_{it}(0))$ denote the potential outcomes of unit $i$ at time $t$ and $W_i$ denote the treatment status at time $1$. Further let $R^{{\scriptscriptstyle Y}}_{it}$ denote the rank of $Y_{it}$ among $\{Y_{jt}: j=1,\ldots,n\}$. Clearly, \[R^{{\scriptscriptstyle Y}}_{it} = n\hat{F}_{Y_t,n}(Y_{it}),\] where $\hat{F}_{Y_t,n}$ is the empirical CDF of $\{Y_{jt}: j=1,\ldots,n\}$.
Throughout the section we assume $(Y_{i0}(1), Y_{i0}(0), Y_{i1}(1), Y_{i1}(0), W_i)$ are i.i.d. and denote by $(Y_0(1), Y_0(0), Y_1(1), Y_1(0), W)$ a generic draw. Similar to the cross-sectional case, we make the following assumptions.
The standard rank-DiD estimator is defined as
Without rank transformation, the usual estimand for DiD is the average treatment effect on the treated (ATT) $\mathbb{E}[Y(1) - Y(0)\mid W = 1]$. A natural definition of rank-ATT is
The parallel trend assumption for original outcomes states that \[\mathbb{E}[Y_1(0) - Y_0(0)\mid W = 1] = \mathbb{E}[Y_1(0) - Y_0(0)\mid W = 0].\] This can be rewritten as \[\mathbb{E}[Y_1(0)\mid W=1] - \mathbb{E}[Y_1(0)\mid W = 0] = \mathbb{E}[Y_0(0)\mid W=1] - \mathbb{E}[Y_0(0)\mid W = 0].\] The left-hand side is the mean difference between $F_{Y_1(0)\mid W = 1}$ and $F_{Y_1(0)\mid W=0}$ and the right-hand side is the mean difference between $F_{Y_0(0)\mid W = 1}$ and $F_{Y_0(0)\mid W=0}$. Motivated by this expression, we define the rank parallel trend assumption as follows.
Note that Assumption (ref) is not equivalent to
We show that the identifying assumption for CiC implies Assumption (ref) but not (ref).
We prove that the effective estimand for the standard rank-DiD estimator is not the rank-ATT and could have a different sign.
To identify the rank-LATE, we need to rank outcomes based on another reference distribution. Let $\hat{F}_{Y_1(0)\mid W=1}$ be an estimate of the counterfactual distribution $F_{Y_1(0)\mid W=1}$. We define a modified rank-DiD estimator by replacing $R^{{\scriptscriptstyle Y}}_{i1} / n$ with $\hat{F}_{Y_1(0)\mid W=1}(Y_i)$:
Unfortunately, the rank parallel trend assumption does not imply identification of $F_{Y_1(0)\mid W=1}$. Nevertheless, it can be consistently estimated under the identifying assumption of CiC athey2006identification or distributional DiD roth2023parallel. Since the former implies the rank parallel trend condition, the rank-ATT can be identified under the same assumptions as CiC.
We consider the standard RDD setting where $X_i$ is a continuous one-dimensional running variable and $x^*$ is the common cutoff for all units. For the purpose of exposition, we focus on sharp RDDs where
The extensions to fuzzy RDDs is straightforward yet mathematically involved. Let $(Y_i(1), Y_i(0))$ denote the potential outcomes. Following imbens2008regression, we make the distributional continuity assumption.
As imbens2008regression noted, for level outcomes, the continuity assumption is stronger than required for identification but practically indistinguishable with the weakest possible assumption in most cases. In addition, when the outcomes are ranked, we need this stronger continuity assumption.
While there are many commonly-used RDD estimators for level outcomes imbens2008regression,lee2010regression,calonico2014robust,imbens2019optimized,eckles2020noise,cattaneo2022regression, we focus on the simplest kernel estimator
Further, let $\hat{\beta}_{{\scriptscriptstyle\mathrm{RDD}}, 1}$ and $\hat{\beta}_{{\scriptscriptstyle\mathrm{RDD}}, 0}$ denote the same estimator but with $R^{{\scriptscriptstyle Y}}_i/n$ replaced by $R^{{\scriptscriptstyle Y}}_{i,1}/n_1$ and $R^{{\scriptscriptstyle Y}}_{i,0}/n_0$, respectively. Here we make the following standard assumptions on the density of $X$ and the kernel.
The following result shows that the kernel RDD estimator does not satisfy the reference robustness and none of these estimators converge to the rank-ATE at cutoff $x^*$:
where $F_{Y(w)\mid x}$ denotes the conditional distribution of $Y(w)$ given $X = x$.
Since we only focus on identification, our results would continue to hold for local-polynomial estimators imbens2008regression,calonico2014robust using the standard asymptotic theory of local-polynomial regression estimators fan1992variable.
Intuitively, to identify rank-ATE at the cutoff, we need to rank based on units near the cutoff. This motivates the following kernel U-statistic:
Note that (ref) is equivalent to (ref) with $R^{{\scriptscriptstyle Y}}_i/n$ replaced by a weighted rank \[\frac{\sum_{j: X_j < x^*} K\left(\frac{X_j - x^*}{h_n}\right) I(Y_j \le Y_i)}{\sum_{j: X_j < x^*}K\left(\frac{X_j - x^*}{h_n}\right) }.\]
In this paper, we study the effective estimands and causal interpretations of popular econometric methods with ranked outcomes or treatments. We introduce the rank-ATE as a primitive causal parameter that enjoys many desirable properties and serves as the building block for other estimands studied in this paper. It also allows us to generalize the estimands for 2SLS, DiD, and RDD to the setting with ranks, though we show that they are not identified by directly applying these methods to ranked outcomes due to the nonlinearity of rank-ATE. We develop alternative identification strategies based on different reference distributions for ranking.
While the rank-ATE is a measure of the comparison between the distributions of treated and control potential outcomes, it cannot be written as an average of individual-level treatment effects, i.e., $\mathbb{E}[g(Y(1), Y(0))]$ chen2023logs. An analogue of rank-ATE is
It is not hard to find data generating distributions under which $\tau_{\mathrm{r}}(F_{Y(1)}, F_{Y(0)})$ and $\tau_{\mathrm{r}}^{*}(F_{Y(1)}, F_{Y(0)})$ have different signs. This is also referred to as the Hand's paradox in biostatistics fay2018causal.
Apparently, $\tau_{\mathrm{r}}^{*}(F_{Y(1)}, F_{Y(0)})$ is only partially identifiable without further assumptions. fan2010sharp derive the tight lower and upper bounds on $\tau^{*}$: \[\tau_{L}^{*}(F_{Y(1)}, F_{Y(0)}) = \max\left\{\sup_{y}(F_{Y(0)}(y) - F_{Y(1)}(y)), 0\right\} - \frac{1}{2}, \] and \[\tau_{R}^{*}(F_{Y(1)}, F_{Y(0)}) = \min\left\{\inf_{y}(F_{Y(0)}(y) - F_{Y(1)}(y)), 0\right\} + \frac{1}{2}.\] With covariates, the bounds can be sharpened lee2021partial.
However, to consistently estimate and quantify uncertainty of the sharper bounds, lee2021partial requires a correct model for the conditional distributions of potential outcomes given the covariates $X$, which is arguably too strong when $X$ is continuous, mixed-type, high-dimensional, or unstructured (e.g., texts). ji2023model overcomes this issue by exploiting Kantorovich duality in optimal transport theory. In particular, our method can wrap around any estimates of $\mathbb{P}(Y(1)\mid X)$ and $\mathbb{P}(Y(0)\mid X)$ and generate valid lower and upper bounds for $\tau^{*}$ even if the estimates are completely off. In the meantime, our bounds are efficient if the estimates of conditional distributions are consistent with the semiparametric rates (i.e., $O(n^{-1/4})$).