Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
89,447 characters · 24 sections · 74 citation commands
An Axiomatic Approach to Comparing Sensitivity Parameters
JEL classification: C18; C21; C51
Keywords: Linear Regression, Treatment Effects, Selection on Observables, Sensitivity Analysis, Unconfoundedness
\allowdisplaybreaks
\onehalfspacing
AngristPischke2015 argue that “careful reasoning about OVB [omitted variables bias] is an essential part of the 'metrics game.” Largely for this reason, researchers have eagerly adopted new tools that let them quantitatively assess the impact of omitted variables on their linear regression results. In particular, researchers now widely use the sensitivity analysis methods developed in AltonjiElderTaber2005 and Oster2019. These methods have been extremely influential, with about 4800 and 5600 Google Scholar citations as of January 2026, respectively. Every top journal in economics is now regularly publishing papers using these methods.\footnote{For example, we surveyed papers published in the top five economics journals in 2024 and 2025. We found 19 which discuss the results of Oster2019, 12 of which implement them numerically, including AndrabiBauDasKhwaja2025, AizerEtAl2024, FrakesGruber2025, HowellEtAl2025, and Rustagi2024, for example. For data from 2019--2021, see the survey in Appendix A of MastenPoirier2024, who note that this method is more commonly used than 2SLS.}
These sensitivity analyses are part of a broader econometric literature on robustness checks. After researchers perform a baseline analysis, using data to draw conclusions about a parameter of interest under a set of baseline assumptions, they typically use various methods to explore what could happen if those baseline assumptions do not hold. Which of these robustness checks should researchers use? There are infinitely many possibilities. In this paper, focusing on the context of assessing robustness to omitted variables, we provide a formal, axiomatic framework that applied researchers can use to narrow down the set of robustness checks that they should implement.
Specifically, a key innovation of AltonjiElderTaber2005 was the idea that the magnitude of OVB can be assessed by reasoning about the relative importance of observed and unobserved variables. This idea can be formalized by defining a specific sensitivity parameter which measures the impact of omitted variables on treatment and/or outcomes. Assumptions about this measure lead to bounds on OVB, thus providing researchers a way to quantitatively assess the impact of omitted variables. However, there are many different ways to formally define the “relative importance” of observed and unobserved variables, each leading to a different sensitivity parameter and thus a different robustness check. Moreover, assumptions about these different sensitivity parameters are usually non-nested and non-falsifiable. Thus the data alone cannot tell us which method or sensitivity parameter is the `correct' one to use. Indeed, the data alone cannot refute the stronger assumption that the omitted variable bias is zero. Consequently, researchers face a practical problem: Which of the many different sensitivity analyses should they implement?
The current, common answers are: (1) Use a method that is well known and widely used in their field\footnote{This choice can be clearly seen in citation patterns: For example, AltonjiElderTaber2005 and Oster2019 are primarily used in economics, CinelliHazlett2020 is primarily used in statistics and political science, and the marginal sensitivity model of Tan2006 is primarily used in statistics and medicine.} and (2) Use a method that is “interpretable,” a subjective criterion with no formal definition. In contrast, we provide the first formal, mathematical framework that can be used to select which sensitivity analysis to perform.
Since the data itself cannot be used to select among the various methods, we propose an approach that is inspired by classical frequentist theory. That classical theory starts from a frequentist thought experiment where repeated samples are drawn from a data generating process, which induces a frequentist distribution for any given estimator or test statistic. Estimators and tests are viewed as procedures which have properties like consistency and efficiency for estimators or size-validity and test consistency for hypothesis tests. Those properties are then used to justify the choice of a specific estimator or a specific hypothesis test. The early statistics literature defined many different formal properties that estimators and tests may have, which have since been used evaluate the desirability of newly proposed estimators and tests.
Analogously, for sensitivity analysis, inspired by the analysis in AltonjiElderTaber2005, we imagine repeatedly drawing sets of covariates from a universe of potentially observable covariates, with each draw partitioning that universe into observed and unobserved covariates. The sensitivity parameters measure the relative importance of the observed and unobserved covariates. Hence their exact value will differ across different draws of observed covariates. This repeated sampling of covariates induces a distribution of the sensitivity parameters, which we call their covariate sampling distribution. Following classical frequentist theory, we suggest using this distribution to compare and evaluate different sensitivity parameters. Specifically, we use this distribution to define two new properties for sensitivity parameters:
We argue that reasonable sensitivity parameters should satisfy both properties. At first glance, it may appear that any sensitivity parameter based on the equal selection benchmark of AltonjiElderTaber2005 would easily satisfy both properties. However, we show that a variety of parameters do not satisfy both properties. This includes the parameter studied in Oster2019, the most widely used parameter in the literature. This parameter is based on a particular residualization, which uses a version of the omitted variable that has been residualized to project out its correlation with the observed variables. Our results show that this transformation is not innocuous and changes the interpretation of the sensitivity parameter so that it no longer captures the notion of equal selection from AltonjiElderTaber2005. In the next two subsections we discuss the implications of our results for both empirical and theoretical research.
As described above, the purpose of our paper is to help researchers narrow down and formally justify the set of robustness checks they choose. Concretely, we focus on the question: How should researchers assess the impact of omitted variables in linear regression? An immediate, practical implication of our results is that applied researchers should carefully consider whether to continue using the very popular sensitivity analysis developed by Oster2019. This conclusion follows since that method relies on a sensitivity parameter which we prove to be inconsistent and non-monotonic, implying that researchers who use this parameter will generally draw the wrong conclusions about robustness. Instead, researchers can choose among various alternative methods, including CinelliHazlett2020 or our companion paper DMP2023v5, both of which use sensitivity parameters that we prove are consistent and monotonic in the present paper. Thus our paper provides a formal, theoretical justification that empirical researchers can use to explain why they use one of these robustness checks, but not another.
Finally, note that although our method can help researchers narrow down the set of robustness checks they use, our results do not deliver a single, unique “best” approach. This is not surprising, however, since it is unlikely that any theory could argue that researchers should perform one and only one robustness check. Moreover, we conjecture that further distinctions between different sensitivity parameters could be obtained by examining other distributional features of covariate sampling distributions besides the two properties we focus on; we leave these extensions to future work.
We develop our framework in the concrete setting of methods for assessing sensitivity to omitted variables in linear regression. Beyond this setting, our framework can be used to help researchers narrow down the set of robustness checks in any other setting where researchers are concerned about omitted variables. For example:
In both settings our framework can be used by theoretical researchers who must define and defend a sensitivity parameter used to measure those relative impacts; we leave the full development of these extensions to future work.
In section (ref), we review the identification problem caused by omitted variables, and set up the main challenge: How can the magnitude of the omitted variable bias be parameterized in a useful way? We survey eleven different measures from the literature in section (ref). We then develop our framework for choosing among these various measures in section (ref). As in AltonjiElderTaber2005, our framework asserts that observed covariates are selected probabilistically from a finite universe of potentially observable covariates. For any specific sensitivity parameter, a covariate selection mechanism induces a distribution over possible values of that sensitivity parameter, which we call the covariate sampling distribution of the sensitivity parameter. The key idea then is to compare this induced distribution across different sensitivity parameters.
We formally define equal selection of covariates as occurring when an equal number of covariates are observed and unobserved, and the specific split of these covariates is chosen uniformly at random. Under equal selection of covariates, from the ex ante perspective before the covariates have been selected, the covariates we observe are equally important for determining treatment as the covariates we do not observe. Put differently, equal selection of covariates implies equal impact of observed and unobserved variables on treatment, ex ante. Therefore, under equal selection, any sensitivity parameter that measures the relative importance of observed and unobserved variables should be approximately centered around the value 1. In this case, the value 1 represents the benchmark of equal selection against which other values of the sensitivity parameter can be compared.
While equal selection may not reflect the actual data-collection process in a given empirical study, we view it as a device that provides a tractable framework and also serves as a useful benchmark for analyzing sensitivity parameters. We perform these analyses in section (ref). Formally, we approximate the distribution of sensitivity parameters induced by covariate selection using an asymptotic approximation based on a data-generating process (dgp) with an increasing number of covariates. This approximation simplifies the analysis substantially, analogous to the simplification that occurs in standard frequentist analysis, whereby the exact finite sample distribution of a statistic is approximated by its asymptotic distribution. In our main result, we prove that the most commonly used approach in the literature is not necessarily centered at 1 under equal selection. Specifically, in Theorem (ref) we prove that Oster's Oster2019 delta parameter can converge to any real number, even under equal selection, unless an exogeneity assumption is imposed on control variables. Consequently, the value 1 is not the correct benchmark for equal selection, for Oster's delta. Intuitively, this bias away from 1 arises because this parameter involves an asymmetry in how the importance of observed and unobserved variables is measured. Empirical researchers widely use the value 1 as a cutoff for assessing robustness. This focus on the value 1 is motivated by its interpretation as the benchmark value of the parameter under equal selection. Thus, our result shows that when the controls are endogenous, researchers generally use the wrong benchmark for equal selection and therefore draw the wrong conclusions about robustness.
We also formally study the behavior of several other sensitivity parameters used in practice, including those proposed by AET2019, CinelliHazlett2020, and DMP2023v5. We show that all of these sensitivity parameters concentrate around 1 under equal selection. We also study whether these sensitivity parameters satisfy a monotonicity property which, roughly, states that the sensitivity parameter's values increase as the number of unobservable variables increases in our framework. We expect reasonable sensitivity parameters to satisfy this property as they purportedly measure the importance of unobservable covariates relative to observable covariates. We show the parameters of CinelliHazlett2020 and DMP2023v5 are monotonic, while the ones proposed by AltonjiElderTaber2005, AET2019, and Oster2019 are not.
These results are shown under high-level conditions on the distribution of covariates; in particular, the structure of the covariances across all observed and unobserved covariates. In section (ref) we show that these high-level conditions are implied by four different kinds of lower-level conditions on the distribution of covariates. In particular, we show that variance matrices with a moving-average, autoregressive, exchangeable, or factor structure all satisfy our high-level assumptions.
Finally, to complement the asymptotic results of section (ref), we study the exact, non-asymptotic distribution of sensitivity parameters in section (ref). We construct an empirically calibrated dgp by treating the data from the paper BFG2020 as if it were the true population. Hence we treat the twenty-two observed covariates in this dataset as the universe of all potentially observable covariates. We then compute and compare the exact distribution of various sensitivity parameters induced by random selection of a fixed number of these twenty-two covariates. Overall, our non-asymptotic findings agree with our asymptotic results.
AltonjiElderTaber2005 discuss the idea of sampling covariates. They argue that random covariate sampling can justify an assumption that their sensitivity parameter equals 1 (their Condition 1), which can then be used for identification. They do not use covariate sampling to compare and contrast different sensitivity parameters, however, which is our focus. The structure of our asymptotic analysis is similar to AET2019, who also use high level assumptions (e.g., their Assumption 2) accompanied by lower level sufficient conditions (e.g., their appendices A.1 and A.5). However, again the goals of the papers are quite different: AET2019 use assumptions about the covariates for identification. In contrast, we focus on the choice of sensitivity parameter, with the goal of selecting a sensitivity parameter that has desirable properties for a large class of covariate distributions. The selected sensitivity parameter can then be used in an identification analysis which does not explicitly impose any structure on the covariates, like Krauth2016, Oster2019, CinelliHazlett2020, or the one in our companion paper DMP2023v5. Another difference is that AET2019 assume that the inclusion indicators are iid Bernoulli random variables (their Assumption 4), whereas we use uniform selection (our (ref)), which allows us to control the ex post proportion of observed covariates.
The contribution of our companion paper DMP2023v5 is to propose new sensitivity parameters ($r_X$ and $r_Y$, see section (ref)) and derive new identification results and sensitivity analyses based on those parameters. In contrast, the present paper does not derive any identification results or propose any new sensitivity analyses. We instead focus on providing a general, axiomatic approach to comparing and selecting sensitivity parameters, as described in section (ref) and applied in section (ref). In particular, we also apply this approach to study the properties of the sensitivity parameters proposed in AltonjiElderTaber2005 and Oster2019. Several previous papers have also studied and critiqued those specific parameters and their uses, including DeLucaMagnusPeracchi2019, CinelliHazlett2020, Basu2022, and MastenPoirier2024. Our paper complements this literature by providing a new and distinct appraisal of those specific parameters. Furthermore, while this prior literature focused on specific parameters, this paper provides a general framework for studying a wide variety of sensitivity parameters, including the eleven parameters described in section (ref).
Finally, our framework uses a finite population design-based setup, similar to recent work like AbadieAtheyImbensWooldridge2020; see BorusyakHull2024 for a review and further citations. In our analysis, the population consists of covariates rather than units, and the design distribution is a set of indicators denoting which covariates are observed, rather than indicators denoting which units are treated.
For random vectors $A$ and $B$, let $\operatorname*{cov}(A,B)$ be the $\text{dim}(A) \times \text{dim}(B)$ matrix whose $(i,j)$th element is $\operatorname*{cov}(A_i, B_j)$. Define $A^{\perp B} \coloneqq A - \operatorname*{cov}(A,B)\operatorname*{var}(B)^{-1}B$. This is the sum of the residual from a linear projection of $A$ onto $(1,B)$ and the intercept in that projection. Many of our equations therefore do not include intercepts because they are absorbed into $A^{\perp B}$ by definition. Note also that $\operatorname*{cov}(A^{\perp B},B) = 0$ by definition. Let $R_{A \sim B \mathrel{\mathsmaller{\bullet}} C}^2$ denote the R-squared from a regression of $A^{\perp C}$ on $(1,B^{\perp C})$. This is sometimes called the partial R-squared. $\iota_K$ denotes a $K \times 1$ vector of ones and $\textbf{I}_K$ denotes a $K \times K$ identity matrix.
Let $Y$ be a scalar outcome variable, $X$ a scalar regressor of interest (often a treatment variable), and $W_1$ be a vector of observed covariates. Consider the OLS estimand of $Y$ on $(1,X,W_1)$. Let $(\beta_\text{med}, \gamma_{1,\text{med}})$ denote the coefficients on $(X,W_1)$. Then we can write \[ Y = \beta_\text{med} X + \gamma_{1,\text{med}}' W_1 + Y^{\perp X,W_1} \] where $Y^{\perp X,W_1}$ is defined to be the OLS residual plus the intercept term, and hence is uncorrelated with each component of $(X,W_1)$ by construction. Researchers often begin their analyses by computing an estimate of coefficients like these. The second step typically asks: How would the coefficient on $X$ change if we include additional unobserved covariates? Let $W_2$ denote the vector of these unobserved variables. Let $W \coloneqq (W_1, W_2)$, where $W_1$ has dimension $d_1$, $W_2$ has dimension $d_2$, and $W$ has dimension $K \coloneqq d_1 + d_2$. Consider the long OLS estimand $Y$ on $(1,X,W_1, W_2)$. Let $(\beta_\text{long}, \gamma_1, \gamma_2)$ denote the coefficients on $(X,W_1,W_2)$. Then we can write
where $Y^{\perp X,W}$ is defined to be the OLS residual plus the intercept term. To ensure that these OLS estimands are well defined, we maintain the following assumption throughout the paper.
Suppose our goal is to learn about the parameter $\beta_\text{long}$. Section 4 of our companion paper DMP2023v5 discusses causal models that lead to this specific OLS estimand as the parameter of interest, using either unconfoundedness, difference-in-differences, or instrumental variables as an identification strategy. Alternatively, it may be that we are simply interested in $\beta_\text{long}$ as a descriptive statistic. The specific motivation for interest in $\beta_\text{long}$ does not affect our technical analysis.
The identification problem is that $W_2$ is not observed, and therefore researchers cannot run the long regression. Hence the omitted variable bias, \[ \text{OVB} \coloneqq \beta_\text{med} - \beta_\text{long}, \] is completely unidentified without further assumptions. Rather than simply assuming that OVB is zero, there is now a large literature on sensitivity analysis that allows researchers to directly reason about and bound the magnitude of OVB. A first pass, naive approach would directly assume $| \text{OVB} | \leq M$ for some known sensitivity parameter $M \geq 0$, which immediately implies that $\beta_\text{long}$ is within $\pm M$ of $\beta_\text{med}$. The problem with this approach is that it is not clear how to select $M$. To avoid this problem, most methods do not reason about OVB directly, but rather parameterize OVB in terms of an “interpretable” sensitivity parameter, and then ask practitioners to make assumptions about this sensitivity parameter. These assumptions can then be translated into bounds on OVB, and hence bounds on $\beta_\text{long}$.
The key difference between the various methods in the literature is therefore how they define the sensitivity parameter which practitioners must reason about. In general, these parameters are completely unidentified from the data alone---recall that the data cannot even refute the claim that OVB is zero---and therefore the data cannot help practitioners select which sensitivity parameters to work with. In light of this challenge, we propose a new theoretical framework that can be used to compare and contrast the different sensitivity parameters. We develop this approach in section (ref). First, however, we review a variety of different sensitivity parameters from the literature in section (ref).
Many approaches to measuring the importance of omitted variables follow the important and influential work of AltonjiElderTaber2005, who proposed that researchers make assumptions on selection ratios, measures which compare the relative importance of observed and unobserved variables. This kind of sensitivity parameter allows empirical researchers to make statements like “to attribute the entire OLS estimate to selection effects, selection on unobservables would have to be at least three times greater than selection on observables” (NunnWantchekon2011, AER, page 3238). There are many different ways to formally define a selection ratio, however. In this section we survey eleven different measures from the literature, numbered P1--P11. Some of these are similar and hence we combine them into four groups.
Following AltonjiElderTaber2005, Oster2019 defines the following sensitivity parameter:
The numerator is a measure of selection on unobservables while the denominator is a measure of selection on observables. Oster2019 provides identification results for $\beta_\text{long}$ under three assumptions: (i) $\delta_\text{orig}$ is known, (ii) $R_{Y \sim X,W_1,W_2}^2$, the R-squared in the long regression of equation (ref), is known, and (iii) exogenous controls, $\operatorname*{cov}(W_1,W_2) = 0$. However, her identification results do not hold under knowledge of $\delta_\text{orig}$ if the controls are endogenous ($\operatorname*{cov}(W_1, W_2) \neq 0$). To allow for endogenous controls, she suggests replacing $\delta_\text{orig}$ with a different sensitivity parameter. Specifically, consider the linear projection of $\gamma_2'W_2$ onto $(1,W_1)$: \[ \gamma_2'W_2 = \phi' W_1 + (\gamma_2'W_2)^{\perp W_1}. \] Here $\operatorname*{cov}(W_1, (\gamma_2'W_2)^{\perp W_1}) = 0$ by construction and $\phi \coloneqq \operatorname*{var}(W_1)^{-1}\operatorname*{cov}(W_1,\gamma_2'W_2)$. Now define
Oster's Oster2019 results now hold so long as the assumptions are stated in terms of $\delta_\text{resid}$, even if exogenous controls fails. We further discuss the technical details behind this residualization in Appendix (ref).
AET2019 propose the following third variation on this type of parameter:
which uses the same numerator but a different denominator. They then provide identification results that involve assumptions on this parameter.
Our next parameter depends on the OLS estimand of $X$ on $(1,W_1,W_2)$. Let $(\pi_1,\pi_2)$ denote the coefficients on $(W_1,W_2)$. Then we can write
where $X^{\perp W}$ is defined to be the OLS residual plus the intercept term, and hence is uncorrelated with each component of $W$ by construction. Our companion paper DMP2023v5 measures selection on unobservables by the standard deviation of the index of unobservables, $\sqrt{\operatorname*{var}(\pi_2' W_2)}$. Likewise, that paper measures selection on observables by the standard deviation of the index of the observables, $\sqrt{\operatorname*{var}(\pi_1' W_1)}$. This leads to the following selection ratio:
That paper also considers a similar measure, but based on the outcome equation (ref),
The parameters in the next group are defined using R-squared's. CinelliHazlett2020 use the sensitivity parameter
They also consider the sensitivity parameter
The following variations can also be considered, although we are not aware of any identification results based on these:
\setcounter{peq}{10}
Finally, Krauth2016 derived identification results under assumptions on the following parameter:
where $W_2 = \phi W_1 + W_2^{\perp W_1}$, assuming $W_1$ and $W_2$ are scalars for simplicity here.
We now have eleven different ways of measuring selection on unobservables. And more measures are likely to be proposed in future research. Which should practitioners use to perform sensitivity analyses? All of the measures in section (ref) are functions of the unobserved variables and hence are not identified from the data. Therefore the data alone cannot tell researchers which parameter they should use to measure relative selection. To address this challenge, in this section we develop a formal, axiomatic approach for choosing between different measures of selection. This approach is based on comparing the properties of each selection ratio under a model of covariate selection that determines which covariates are observed.
We start by recalling the baseline model from section (ref), but now from the ex ante perspective where it is not yet known which covariates will actually be observed. Thus we let $W \in \ensuremath{\mathbb{R}}^K$ denote the random vector of all potentially observable covariates. $Y$ and $X$ are random scalars as before. Given the random vector $(Y,X,W)$, we can consider the linear projections
These equations are identical to equations (ref) and (ref), except that we have not specified which covariates are observed.
Inspired by the analysis in AltonjiElderTaber2005, we consider a model by which a subset of the components of $W$ will be observed. For each component $W_k$ of $W = (W_1,\ldots,W_K)$, let $S_k$ be a binary random variable denoting whether $W_k$ is observed or not. Let $S \coloneqq (S_1,\ldots,S_K)$ be a random vector with support $\{0,1\}^K$. The distribution of $S$ is called the design distribution. For any given realization $s$ of $S$, define
This notation emphasizes that the identity of the observed and unobserved covariates, along with their corresponding coefficients in equations (ref) and (ref), is determined by a random draw of $S$. For each sensitivity parameter in section (ref), we can use this notation to denote its value for each realization $s$. For example, the sensitivity parameter (ref) is \[ \delta_\text{resid}(s) \coloneqq \frac{\operatorname*{cov}(X, \gamma_2(s)'W_2(s)^{\perp W_1(s)})}{\operatorname*{var}( \gamma_2(s)'W_2(s)^{\perp W_1(s)} )} \hspace{-1mm} \Bigg/ \hspace{-1mm} \frac{\operatorname*{cov}(X, (\gamma_1(s) + \phi(s))' W_1(s))}{\operatorname*{var}( (\gamma_1(s) + \phi(s))' W_1(s))} \tag{\ref{eq:deltaResid}$^\prime$} \] where $\phi(s) \coloneqq \operatorname*{var}(W_1(s))^{-1}\operatorname*{cov}(W_1(s),\gamma_2(s)'W_2(s))$, and (ref) is \[ r_X(s) \coloneqq \frac{\sqrt{ \operatorname*{var}(\pi_2(s)' W_2(s))} }{ \sqrt{\operatorname*{var}(\pi_1(s)'W_1(s))} }. \tag{\ref{eq:rX}$^\prime$} \] This notational dependence on $s$ highlights the dependence of the sensitivity parameter values on the identity of the observed and unobserved covariates. Let $\theta(s)$ denote a generic sensitivity parameter as a function of the covariate selection realization, such as any of the parameters defined in section (ref).
We can now state the main idea of this paper:
This idea is directly analogous to the long-established practice in frequentist statistics of comparing the performance of estimators, hypothesis tests, etc., in terms of their sampling distributions. The main difference is the kind of frequentist thought experiment we are using. Here we consider repeated sampling of covariates, whereas traditional statistics considers repeated sampling of units. Despite this difference, this analogy is useful for framing our subsequent analysis, because many of the challenges that arise in implementing classical frequentist analysis arise here as well.
Specifically, the covariate sampling distributions of the sensitivity parameters depend on two things:
For the second aspect, our analysis will allow for lower-level assumptions on the joint covariance matrix of $(Y,X,W)$; we discuss this further in section (ref) below. For the first aspect, we in principle could carry out all of the analysis below under any posited design distribution of $S$, in the same way that properties of different estimators are studied under different sampling distributions of the observations. As we will discuss, the following design distribution provides a tractable and useful starting point.
Importantly, for most actual empirical studies, we do not view (ref) as an accurate description of the true process by which covariates are observed. Instead, we view it as a simplified and tractable model of covariate selection which is useful for comparing sensitivity parameters. If a sensitivity parameter has undesirable properties under (ref), it is unlikely to have better properties under more realistic and complicated models of covariate selection.
The main purpose of (ref) is to provide a benchmark for formally comparing sensitivity parameters. Indeed, a major motivation for using selection ratios to measure the importance of omitted variables is that they allow us to think about the benchmark case where the observed and unobserved variables are equally important, typically called “equal selection.” As we discuss in section (ref), “equal selection” is by far the most commonly used benchmark in empirical work, and is commonly considered to be the cutoff between a robust and non-robust result. It is therefore a particularly important case to study. One way to formalize this notion of “equal selection” is based on the hypothetical covariate selection distribution in (ref), as follows.
The motivation for this definition is similar to the analysis in AltonjiElderTaber2005: When we see exactly half of all covariates, and which covariates we see are chosen at random, then from the ex ante perspective before the covariates have been selected, the covariates we do observe will be equally important for determining treatment as the covariates we do not observe. So this definition explicitly formalizes the concept of “equal selection” in terms of a specific mechanism by which covariates are observed. Likewise, this definition formalizes the concept of more and less selection on unobservables.
The ex ante perspective here is directly analogous to the design-based analysis of randomized experiments (e.g., ImbensRubin2015 chapter 5). Randomized treatment assignment does not guarantee that potential outcomes will be perfectly balanced across treatment groups ex post. But it does guarantee ex ante balance, or balance across repeated randomizations. Similarly, equal selection does not guarantee that any given realization $s$ of $S$ will lead to a partition $W_1(s)$ and $W_2(s)$ of $W$ such that the observed variables are equally as important as the unobserved variables. But it does guarantee this ex ante, or across repeated selections of the covariates.
We can now define two properties of selection ratios that formalize the idea that they measure the relative importance of observed and unobserved variables.
Both properties are asymptotic, as the number of covariates gets large. Convergence in probability here refers to the covariate sampling distribution. We discuss these asymptotics in detail in section (ref) below. It is possible to define non-asymptotic versions of these properties, but just like the usual analysis of frequentist exact sampling distributions, it is difficult to prove exact results for large classes of data generating processes. Property 1 is a consistency requirement. We view this as a weak and natural requirement for any sensitivity parameter $\theta(\cdot)$ which aims to measure the relative importance of unobserved and observed covariates. Specifically, if $\theta(\cdot)$ is defined as the ratio of a measure of the importance of unobservables to a measure of the importance of observables, then Property 1 says that the absolute value of this ratio should be approximately centered around the value 1---nominally meaning that the unobservables and observables are equally important---when the covariates in fact satisfy equal selection. Property 2 is a stronger version of this requirement. This stronger version requires not only that the sensitivity parameter be approximately 1 under equal selection, but that it is approximately larger than 1 when there is more selection and it is approximately smaller than 1 when there is less selection. We also view this as a weak and natural requirement, since it says that the relative measure $\theta(\cdot)$ should say the unobservables are more important than the observables whenever there is more selection on unobservables. Likewise it says the relative measure $\theta(\cdot)$ should say unobservables are less important than the observables whenever there is less selection on unobservables.
Finally, as in classical frequentist statistics, there are other properties of the covariate sampling distributions we could consider. For example, we could study the asymptotic distribution of $\theta(S)$, rather than just its probability limit. However, as we show below, it turns out that even the weak requirements of Properties 1 and 2 above are often sufficient criteria to choose between the sensitivity parameters present in the literature. So we leave a careful study of other properties of this asymptotic distribution to future work.
In the previous section we described a general approach for comparing selection ratios: Compare their covariate sampling distributions. We then defined two properties that these sampling distributions could have, and argued that these are desirable properties. In this section we study whether these properties hold for several of the specific selection ratios defined in section (ref). For brevity we provide results for the first six parameters only, which include the most widely used approaches by empirical researchers.
As mentioned above, the covariate sampling distributions depend on features of the joint distribution of $(Y,X,W)$. In practice, these are unknown (because not all components of $W$ are observed). This is analogous to the fact that the exact sampling distribution of estimators also typically depends on the exact dgp, which is unknown. In this section we address this challenge by using asymptotics as a tool to approximate their exact distributions. Specifically, we study their probability limits as the number of covariates $K$ gets large. This approach is directly analogous to the literature on design-based inference on treatment effects in finite populations. That literature uses sequences of finite populations of growing size to approximate exact distributions for populations of fixed size. For example, see LiDing2017 or AbadieAtheyImbensWooldridge2020. In section (ref) we examine the quality of the asymptotic approximation by showing that several features of the asymptotic analysis appear in the exact, non-asymptotic covariate sampling distributions that arise when the population dgp is constructed using data from a published empirical application.
To analyze the probability limits of the sensitivity parameters, we make several high level assumptions on the dgp. In section (ref) we give various sets of lower-level sufficient conditions for these high level assumptions. These lower-level conditions suggest that our high level assumptions are compatible with a large range of data generating processes.
First, we consider sequences where the limiting relative proportion of unobserved to observed covariates is nontrivial.
This assumption requires both $d_1$ and $d_2$ to grow to infinity as $K$ grows. When $r = 1$, this corresponds to (asymptotic) equal selection. This is a weaker version of equal selection than that given in definition (ref), which would require $d_1 = d_2$ at every point along the sequence.
Second, we impose some regularity assumptions on the data generating process. Here we use the notation $W^K$ synonymously with $W$ to emphasize that the dimension of the covariates depends on $K$. Likewise for their coefficients, $\pi^K$ and $\gamma^K$. In this assumption, the subscript $S$ in $\operatorname*{var}_S$ refers to operators taken with respect to the design distribution of $S$. This is to distinguish it from operators taken with respect to the distribution of the random vector $(Y,X,W)$, which appear without subscripts.
Assumption (ref).1 says that the variance of $\pi'W$, the projection $X$ on $(1,W)$, is bounded from above and away from zero as the number of covariates $K$ varies. By properties of projections $\operatorname*{var}(\pi'W) \leq \operatorname*{var}(X)$. Thus an unbounded variance $\operatorname*{var}(\pi'W)$ would be incompatible with a finite variance for $X$. This assumption requires that either the $K$ entries in the coefficient vector $\pi$ shrink with $K$ or that the $K^2$ elements of $\operatorname*{var}(W)$ shrink with $K$.
Assumption (ref).2 is a tail condition that we use to apply a law of large numbers to replace sample averages with their expectations. We then use Assumption (ref).3 to ensure convergence of a scaled version of those expectations. To understand that assumption, note that the variance of any sum of random variables equals the sum of the variances plus other terms that depend on covariances. For a given $K$ and distribution of $(Y,X,W)$, the ratio of these two components is fixed at some number. Assumption (ref).3 allows this ratio to vary with $K$ so long as it eventually converges. In this respect it is similar to Assumption (ref). When the covariates are all mutually uncorrelated---which implies that the observed variables are always exogenous---Assumption (ref).3 holds with $c_\pi =1$. When the covariates are correlated, however, generally $c_\pi \neq 1$.
We will also use the same regularity conditions, but replacing the selection equation coefficients with the outcome equation coefficients.
The interpretation of (ref) is similar to (ref). Our final assumption restricts the impact of a single covariate on (a) treatment and (b) the portion of outcomes accounted for by covariates.
For example, $\sup_{i =1,\ldots,K} | \operatorname*{cov}(X, \gamma_i^K W_i^K) | = o(1 / \sqrt{K})$ is a sufficient condition for the first part of (ref).
We can now derive the probability limits of the sensitivity parameters under consideration. We start with $r_X$ and $r_Y$, because they are particularly simple.
The limiting value of this sensitivity parameter depends on two factors: the ratio of unobserved to observed covariates $r$, and the nonnegative constant $c_\pi$ which depends on the correlation in the covariates. Importantly, for any value of $c_\pi \geq 0$, this ratio is strictly increasing in $r$, and it equals 1 when $r = 1$. Hence, under assumptions (ref)--(ref), $r_X(S)$ satisfies Properties 1 and 2 from section (ref). In this sense, the sensitivity parameter $r_X$ correctly measures the importance of unobservables relative to observables.
If we replace $r_X$ with $r_Y$ and $c_\pi$ with $c_\gamma$ in Theorem (ref), the result continues to hold. This follows since these two parameters have similar definitions. See Appendix (ref) for details.
Next consider \[ \delta_\text{orig}(s) \coloneqq \frac{\operatorname*{cov}(X,\gamma_2(s)'W_2(s))}{\operatorname*{var}(\gamma_{2}(s)'W_2(s))} \hspace{-1mm} \Bigg/ \frac{\operatorname*{cov}(X,\gamma_1(s)'W_1(s))}{\operatorname*{var}(\gamma_1(s)'W_1(s))}. \]
Like our other high level assumptions, the additional assumption on $\operatorname*{cov}(X, \gamma'W)$ in this theorem holds under a variety of lower level conditions, like MA, AR, or factor covariates; see section (ref). Similar to $r_X$, the limiting value of the sensitivity parameter $\delta_\text{orig}(S)$ depends on the ratio of unobserved to observed covariates $r$ and the nonnegative constant $c_\gamma$ which depends on the correlation in the covariates. And like $r_X$, for any $c_\gamma \geq 0$, this limiting value equals 1 when $r=1$. Hence $\delta_\text{orig}(S)$ is consistent. However, it does not satisfy the monotonicity Property 2. Specifically, when $c_\gamma \in [0,1)$, such as when the covariates are exchangeable ($c_\gamma = 0$), this limiting value is monotonic, but in the wrong direction---when a higher proportion of variables are not observed ($r > 1$) the limiting value is smaller than 1, nominally suggesting that the unobserved variables are less important than the observed variables. Conversely, when a higher proportion of variables are observed ($r < 1$), the limiting value is larger than 1, nominally suggesting that the unobserved variables are more important than the observed values. Another unusual property is that when the covariates are uncorrelated, $c_\gamma = 1$, which implies that $\delta_\text{orig}(S) \xrightarrow{p} 1$ regardless of the value of $r$; that is, regardless of whether most covariates are observed, or whether most covariates are unobserved.
Next we consider the residualized version of this sensitivity parameter, $\delta_\text{resid}(S)$. Residualization makes both the behavior and analysis of this parameter much more complicated than the previous parameters. The following theorem considers the case where the covariates are exchangeable. Note that (ref) is formally defined on page (ref).
Theorem (ref) provides an asymptotic representation for $\delta_\text{resid}(S)$. This representation has two main consequences.
Thus, for any value of $r$ and any number $C$, we can find a sequence of dgps such that $\delta_\text{resid}$ converges to $C$ as $K\rightarrow \infty$. In particular, this result holds even if $r = 1$. Corollary (ref) therefore shows that, under the same assumptions by which we showed that $r_X(S)$ satisfies Properties 1 and 2, $\delta_\text{resid}(S)$ does not satisfy either of those properties. It is possible that $\delta_\text{resid}(S)$ may satisfy these properties under stronger assumptions on the dgp, and this is one possible direction for future work. In this case, empirical researchers will have to argue that they believe their dgp satisfies these narrower restrictions in order to ensure that $\delta_\text{resid}$ satisfies the properties in their application.
To build intuition for Corollary (ref), compare equation (ref) for $\delta_\text{orig}$ with equation (ref) for $\delta_\text{resid}$. The equation for $\delta_\text{orig}$ treats $W_1$ and $W_2$ symmetrically. Thus, under equal selection, the numerator and denominator of $\delta_\text{orig}(S)$ have the same covariate sampling distribution. In contrast, $\delta_\text{resid}$ treats $W_1$ and $W_2$ asymmetrically: $W_2$ is always projected onto $W_1$ because of residualization. This residualization breaks the symmetry between the numerator and denominator of $\delta_\text{orig}$. Consequently, when the controls are endogenous, the numerator and denominator do not have the same covariate sampling distribution, even under equal selection; equation (ref) in Theorem (ref) illustrates this asymmetry formally.
The following corollary gives an even stronger negative result.
This result shows that, like $\delta_\text{orig}(S)$, $\delta_\text{resid}(S)$ can exhibit a reverse monotonicity property---it is large when $r$ is small and small when $r$ is large. This is the opposite of Property 2. In section (ref) we show that this reverse monotonicity is not just a theoretical phenomenon but can arise in real empirical datasets.
Next consider the third variation on this type of sensitivity parameter, \[ \delta_\text{ACET}(s) \coloneqq \frac{\operatorname*{cov}(X,\gamma_2(s)'W_2(s)^{\perp \gamma_1(s)'W_1(s)})}{\operatorname*{var}(\gamma_2(s)'W_2(s)^{\perp \gamma_1(s)'W_1(s)})} \hspace{-1mm} \Bigg/ \hspace{-1mm} \frac{\operatorname*{cov}(X,\gamma_1(s)'W_1(s)^{\perp \gamma_2(s)'W_2(s)})}{\operatorname*{var}(\gamma_1(s)'W_1(s)^{\perp \gamma_2(s)'W_2(s)})}. \]
Like $\delta_\text{orig}(S)$, $\delta_\text{ACET}(S)$ treats the observed and unobserved variables symmetrically, and hence it generally satisfies Property 1. Unlike $\delta_\text{orig}(S)$ or $\delta_\text{resid}(S)$, however, it does not exhibit a reverse monotonicity property. Instead, it is asymptotically a constant function in $r$. Theorem (ref) is similar to Theorem 1 and Corollary 1 of AET2019, which show that their parameter converges to 1 under certain assumptions.
Finally, consider \[ k_X(s) \coloneqq \frac{R^2_{X \sim W} - R^2_{X \sim W_1(s)}}{R^2_{X \sim W_1(s)}}. \] Like $\delta_\text{resid}(S)$, this sensitivity parameter is somewhat complicated to analyze. Here we consider the case where the covariates are exchangeable with shrinking variance (formally stated as (ref) in section (ref)).
Theorem (ref) shows that, when the covariates are exchangeable with small covariances, $k_X(S)$ satisfies both Properties 1 and 2. Under these same assumptions, both $r_X(S)$ and $r_Y(S)$ also satisfy Properties 1 and 2. This shows that these two properties alone are not sufficient to fully distinguish between all possible selection ratios. This is analogous to how many---but not all---estimators are consistent, and hence we have to look beyond their probability limits to further distinguish between them; for example, by studying their asymptotic distributions. The same situation arises here, and we conjecture that examining the asymptotic distribution of selection ratios that satisfy Properties 1 and 2 will raise further differences between them. We leave that to future work.
Note that if the covariates are uncorrelated, $k_X(s) = r_X(s)^2$ and hence $k_X(S) \xrightarrow{p} r$ as $K \rightarrow \infty$ in this case. Another open question concerns the behavior of $k_X(S)$ under more general conditions on the dgp than we give in Theorem (ref). In section (ref) we show that $k_X(S)$ generally performs well in an empirical dataset that does not satisfy the exchangeability assumption. This suggests that it likely performs well in wider classes of dgps, although we leave the full theoretical analysis to future work.
Our asymptotic analysis in section (ref) used various high level conditions on the dgp, assumptions (ref), (ref), and (ref). In this section we give simple lower level conditions on the covariance structure of $\operatorname*{var}(W^K)$ for this set of assumptions. Specifically, we show that if the components of $W^K$ (a) satisfy a moving average (MA) process, (b) satisfy an autoregressive (AR) process, (c) satisfy a factor model, or (d) are exchangeable with shrinking covariances, then assumptions (ref), (ref), and (ref) hold.
We begin by considering covariates that have moving average type dependence.
(ref).1 says that covariates are correlated only if they are at most one index apart. The magnitude restriction on $\rho$ ensures that the covariates' variance matrix is positive semi-definite for all $K$. (ref).2 says that none of the covariates is too important, in the sense that their maximal coefficient value is bounded. This bound also shrinks with $K$ to ensure that $\operatorname*{var}(\sum_{i=1}^K \pi_i^K W_i^K)$ does not diverge. To understand (ref).3, note that (ref).1 implies
So (ref).3 says that the $\rho$ parameter and the coefficient sequence are such that the relative contribution of these two terms of the variance converges. This is a more primitive version of (ref).3. Like that high level assumption, (ref).3 says that we only consider sequences where this relative contribution is stable. Finally, (ref).4 simply says that covariates are not degenerate in the treatment effect equation.
Most of the selection ratios we consider in section (ref) are invariant to the scale of $\pi$ or $\gamma$. This suggests that (ref).2 can always be satisfied by simply scaling the coefficients down appropriately. However, a naive rescaling will violate (ref).4. So the two assumptions are in fact restrictive.
This result can be extended to the $\text{MA}(q)$ for $q \geq 1$ case at the expense of a longer proof.
Consider the following assumption.
When $\rho = 0$ in (ref).1, we adopt the convention that $0^0 = 1$. (ref).1 says that the magnitude of the covariance between any two covariates decays as their indices become farther apart. The interpretation of the rest of (ref) is very similar to (ref), which we discussed above. The main difference is (ref).3, which arises since (ref).1 yields the following structure on the variance of the sum: \[ \operatorname*{var} \left( \sum_{i=1}^K \pi_i^K W_i^K \right) = \sum_{i=1}^K (\pi_i^K)^2 + \sum_{i,j:i \neq j} \pi_i^K \pi_j^K \rho^{|i-j|}, \] reflecting the different covariance structure for the AR dgp compared to the MA dgp.
This result can be extended to the $\text{AR}(p)$ for $p \geq 1$ case at the expense of a longer proof.
Consider the following assumption.
(ref).1 imposes a factor structure on the covariates' variance matrix. The special case of exchangeable covariates with covariance $\rho \in (0,1)$ obtains by letting $R=1$, $\Lambda^K = \iota_K \sqrt{\rho}$ and $\sigma^2_E = 1-\rho$. (The exchangeable case with covariance $\rho = 0$ is included as a special case of either the earlier MA(1) or AR(1) assumptions.) (ref).2 is similar to the bounded coefficients assumptions from the MA and AR cases, except that the bound is now of order $K^{-1}$ instead of $K^{-1/2}$. In the exchangeable case, for example, $\operatorname*{var}(\sum_{i=1}^K \pi_i^K W_i^K)$ is of order $K^2 \cdot O(\sup_{i =1,\ldots,K} | \pi_i^K |^2)$. Hence we need the largest $\pi_i^K$ values to shrink at least at the $K^{-1}$ rate to keep the variance finite. (ref).3 requires the factor loadings to be uniformly bounded.
For one of our results in section (ref) we use the following low level conditions, which impose that the covariates are exchangeable with a shrinking covariance. We state these assumptions separately from (ref)---which includes exchangeable covariates with non-shrinking covariances as a special case---because the shrinking covariance requires a different scaling of the coefficients $\pi$ and $\gamma$.
In section (ref) we used asymptotics to approximate and compare the covariate sampling distributions of various selection ratios. That analysis raises several questions: (1) How accurate are the asymptotic approximations? (2) How restrictive are the regularity conditions used to obtain these approximations? (3) Is the non-convergence of $\delta_\text{resid}(S)$ to 1 under equal selection that we showed in Theorem (ref) typical in any sense, or is it a rare knife-edge case that can be safely ignored? We next address these concerns by comparing the selection ratios' covariate sampling distributions in an exact, non-asymptotic setting based on an empirically calibrated dgp using data from BFG2020. We find similar patterns as in the asymptotic analysis: For example, the distribution of $r_X(S)$ is tightly centered around 1 under equal selection, and satisfies a natural monotonicity property. In contrast, the distribution of $| \delta_\text{resid}(S) |$ is highly variable under equal selection, and follows a reverse monotonicity property: It is larger when most covariates are observed and it is smaller when most covariates are not observed. This is the exact opposite pattern from its nominal interpretation, where large values of $| \delta_\text{resid}(S) |$ are supposed to represent cases where the unobserved variables are substantially more important than observed variables. Overall, our findings show that the asymptotic results reflect properties of the sensitivity parameters that arise in real data sets with a finite number of covariates.
Since our goal is simply to create a realistic joint distribution of $(Y,X,W)$ from which we can sample covariates, we omit a description of the dataset and refer interested readers to BFG2020 for details. Here we provide the minimum detail needed for replication. We let $Y$ denote Republican vote share, $X$ total frontier experience. For the covariates $W$, we pick the 10 covariates that the authors use in their baseline specifications (their Table 3) as well as 12 of the 13 additional variables that the authors consider in their appendix. We omit one variable (contemporary population density) so that the total number of covariates $K$ is even; similar results obtain if we include it but then equal selection is not exactly satisfied. This makes for a total of 22 covariates in the vector $W$. As in the authors' sensitivity analysis, we also treat state fixed effects as variables that are not used for calibration, and hence we project them out from all other variables. We thus let $(Y,X,W)$ have the empirical distribution of all these residualized variables in the data. Using this joint distribution of $(Y,X,W)$, $(\beta_\text{long},\gamma)$ are defined as the estimated coefficients on $(X,W)$ from OLS of $Y$ on $(1,X,W)$. Likewise, $\pi$ is the vector of estimated coefficients on $W$ from OLS of $X$ on $(1,W)$. For this exercise, we therefore treat these estimates as the population values of these coefficients.
Having specified the distribution of $(Y,X,W)$, we can now compute the exact covariate sampling distributions of any sensitivity parameter $\theta(S)$ by fixing a number $d_1$ of covariates to observe from the $K=22$ total covariates, computing $\theta(s)$ for all possible values of $s$ with $\sum_{k=1}^K s_k = d_1$, and then putting equal weight on all of these values.
We compute these covariate sampling distributions for the six sensitivity parameters which we analyzed asymptotically in section (ref). For brevity, we focus on $r_X(S)$ and $| \delta_\text{resid}(S) |$ here; Appendix (ref) shows the results for the other four parameters. Figure (ref) plots histograms of the covariate sampling distributions for three choices of $d_1$, from left to right: More covariates observed ($d_1 = 19$), Equal selection ($d_1 = 11$), and More covariates unobserved ($d_1 = 3$). First consider the top row, which shows the distributions of $r_X(S)$. From the middle plot, we see an empirical analog of Property 1: The distribution of $r_X(S)$ is tightly centered at 1 under equal selection. From the left and right plots we see an empirical analog of Property 2: When most covariates are observed, $r_X(S)$ is mostly below 1, indicating that the unobserved covariates are not as important as the observed covariates. Likewise, when most covariates are not observed, $r_X(S)$ is mostly above 1, indicating that the unobserved covariates are more important than the observed covariates.
Next consider the bottom row, which shows the distributions of $| \delta_\text{resid}(S) |$. From the middle plot, which shows the equal selection case, we see that the distribution is very spread out, and does not appear to have any discernible concentration near one. From the left and right plots, we see that $| \delta_\text{resid}(S) |$ also does not satisfy any kind of empirical analog of the monotonicity Property 2. In fact, the distribution shifts closer to zero as we increase the number of unobserved covariates.
This reverse monotonicity property can also be seen by considering various summary statistics for these distributions. Table (ref) shows the proportion of realizations of the sensitivity parameters that are below the equal selection benchmark of 1. For $r_X(S)$ this happens 99.2% of the time when most covariates are observed, exactly 50% of the time under equal selection, and only 0.8% of the time when most covariates are not observed. In contrast, $| \delta_\text{resid}(S) |$ is smaller than 1 only 28.7% of the time when most covariates are observed, and this increases up to 49.5% of the time when most covariates are unobserved. This is the opposite pattern we would expect from the nominal interpretation of $\delta_\text{resid}(S)$, whereby values smaller than 1 are supposed to indicate that the unobserved variables are less important than the observed variables. Table (ref) also shows the results for the other four sensitivity parameters we consider. Again we see that these finite sample properties mirror their asymptotic properties.
Similar findings can be seen in Table (ref), which shows additional summary statistics for two of these distributions (Appendix Table (ref) shows results for the other four distributions). When most covariates are observed, the maximum value of $r_X(S)$ is 1.107, with a median value of 0.525. Thus the $r_X(S)$ parameter almost always correctly reports that the unobserved covariates are less important than the observed covariates. Likewise, when most covariates are not observed, the minimum value is 0.904, with a median value of 1.906. Again, the $r_X(S)$ parameter almost always correctly reports that the unobserved covariates are more important than the observed covariates. In contrast, $| \delta_\text{resid}(S) |$ is mostly large when most covariates are observed, with a median of 1.774, a 25th percentile of 0.879, and a standard deviation of 3.06. When most covariates are not observed, this distribution shifts down, with the 25th, 50th, and 75th percentiles all decreasing---the opposite direction we would expect from the nominal interpretation of $| \delta_\text{resid}(S) |$. Under equal selection, $| \delta_\text{resid}(S) |$ is somewhat biased upwards, with a median of 1.118, but it is also very spread out, with a very large upper tail and a standard deviation of about 12,300. In contrast, $r_X(S)$ is tightly distributed exactly around 1 under equal selection.
Since the original work of AltonjiElderTaber2005, empirical researchers now regularly discuss the robustness of their results to omitted variables in relative terms, comparing the magnitudes of selection on observables with unobservables. In particular, they use the value 1 as an important benchmark for their sensitivity parameter, nominally interpreted as meaning “equal selection” (e.g., this is true for all of the empirical papers published in top 5 journals from 2019--2021 in the survey in Appendix A of MastenPoirier2024). From our results in section (ref), we see that the value 1 is usually not the correct benchmark of equal selection for the most popularly used sensitivity parameter, $\delta_\text{resid}$. This implies that researchers who wish to compare values of their sensitivity parameter to the benchmark of equal selection will generally draw the wrong conclusions about robustness if they use 1 as the benchmark.
One possible response to this result is to continue to use $\delta_\text{resid}$, but to change the benchmark value. The problem with this approach is that the correct benchmark of equal selection generally depends on the unknown dgp. Equation (ref) gives a simple example of how this limit depends on various features of the unknown dgp when the covariates are exchangeable. Hence the correct benchmark is currently unknown. Moreover, even if the correct benchmark were known, we observed that $\delta_\text{resid}$ can exhibit a counterintuitive reverse monotonicity property. In contrast, we have shown that alternative sensitivity parameters can achieve a limiting value that does not depend on the dgp under equal selection (e.g., Theorem (ref)), and thus provides a feasible benchmark. This is the property we called consistency. In particular, the sensitivity parameters in both CinelliHazlett2020 and our companion paper DMP2023v5 are consistent, and also satisfy the monotonicity in selection property we introduced. Based on this criterion, one immediate practical implication of our results is that researchers should consider using sensitivity analyses beyond those involving $\delta_\text{resid}$. Hence we recommend that applied researchers use either of those methods, or any other methods that satisfy those two properties.
We showed that consistency alone is sufficient to rule out some sensitivity parameters. We also showed that some parameters are consistent, but do not satisfy monotonicity in selection, such as $\delta_\text{orig}$, so that this second property is also useful for distinguishing between sensitivity parameters. Nonetheless, there are multiple parameters that satisfy both requirements. We conjecture that further refinements can be obtained by examining the distributional features of covariate sampling distributions, rather than just their probability limits, although it is unlikely that future work will ever recommend a single unique robustness check. Finally, as discussed in section (ref), our framework can be straightforwardly extended to help researchers narrow down the set of robustness checks in other settings as well, including assessing the parallel trends assumption in difference-in-differences analyses or the exogeneity assumption in instrumental variable models.
\singlespacing