Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
26,425 characters · 10 sections · 20 citation commands
The Role of Measured Covariates in Assessing Sensitivity to Unmeasured Confounding
Keywords: Sensitivity Analysis, Omitted Variable Bias, Multicollinearity, Proximal Learning.
Confounding, especially unmeasured confounding, remains a central concern in observational studies. Few examples illustrate this more clearly than the long-standing debate over the relationship between smoking and lung cancer. The seminal paper by cornfield1959smoking played a pivotal role in this discussion by examining how strong an unmeasured confounder would need to be to overturn a causal interpretation. Their analysis showed that any such factor would have to be implausibly powerful given the scientific and social context of the time. This conclusion was supported by a careful consideration of alternative explanations, including age, sex, marital status, diagnostic practices, urban residence, and environmental exposure, thereby providing compelling robustness for the causal link between smoking and lung cancer.
The success of sensitivity analysis is intricately tied to the quality of the measured covariates in an observational study hammond1964smoking. Sensitivity analysis can strengthen the credibility of a well-designed study, but it cannot substitute for thoughtful covariate selection. The central role of measured confounders in the design of an observational study cannot be overstated and is pivotal in assessing the rigor of such an analysis. In fact, under reasonable conditions, measured covariates can even attenuate bias arising from unmeasured confounders zhang2025general. In this regard, it is imperative to consider the effect of measured covariates on the sensitivity of causal conclusions.
In a correctly specified standard linear regression of an outcome $Y$ on an exposure $A$ and a measured confounder $X$, strong multicollinearity between $A$ and $X$ is traditionally viewed as a problem of statistical power farrar1967multicollinearity. When variables are highly collinear, standard errors increase, and estimates become less precise, making it harder to detect treatment effects.
A skeptic may argue that even under multicollinearity, effect estimates from a correctly specified linear model remain valid, and that observing a significant effect despite multicollinearity therefore strengthens causal evidence. However, this conclusion relies critically on the assumption of no unmeasured confounding. When such an assumption is questionable, the implications of strong multicollinearity between the exposure and measured confounders on the sensitivity of a causal claim remain unclear.
For instance, consider two studies as shown in Figure (ref), with an exposure $A$, a measured covariate $X$, and an outcome $Y$. Suppose $X$ has the same distribution in both studies and serves as a stable proxy for an unmeasured confounder $U$, as is common in observational studies. Figure (ref) shows that the two studies report similar effect size estimates, standard errors, and sensitivity analyses (using the sensitivity model of cinelli2020making). However, exposure is more strongly predicted by $X$ in Study 1 than in Study 2 (see Appendix B for the data-generating mechanisms). Analogous situations could arise if one instead used sensitivity models such as those in rosenbaum2010design or tan2006distributional. In this setting, does one study warrant greater confidence in its causal claims than the other?
Observational studies in which the exposure is heavily confounded with other covariates have often been shown to yield misleading causal conclusions. Notable examples include early findings on the protective effects of hormone replacement therapy on coronary artery disease grodstein2001postmenopausal and the adverse effects of caffeine consumption during pregnancy on birth weight vlajinac1997effect, both of which were later overturned by experimental evidence. In each case, the exposure was strongly associated with measured differences in lifestyle, rendering the studies highly susceptible to unmeasured confounding. These examples serve as cautionary tales, underscoring the need for heightened skepticism when drawing causal conclusions from non-experimental data in settings where exposure is strongly confounded rutter2007identifying.
In this article, we investigate this question further, studying how the relationship between measured covariates and exposure shapes the interpretation of sensitivity analyses in observational studies, particularly when these covariates proxy for latent confounders. Revisiting smoking as a motivating example, we show how evolving socioeconomic patterns in exposure uptake alter the sensitivity of causal claims over time. Together, these considerations offer a framework for a more nuanced interpretation of sensitivity analyses in the presence of strong exposure–covariate associations.
We begin by revisiting the standard linear regression setup where the explanatory variable is measured with error.
It is well known that an ordinary least squares (OLS) regression coefficient $\beta_{Y\sim X}$ of $Y$ on $X$ in Model (ref) underestimates the true $\beta$, as given by Equation (ref).
This phenomenon is often referred to as attenuation bias in econometrics wooldridge2010econometric, and it foreshadows the problem we address when there is strong multicollinearity between $A$ and $X$ in a regression, and when $X$ is itself associated with an unmeasured variable $U$. In such settings, the bias in estimating the slope of $A$ on $Y$ is affected by the association between $U$ and $X$, with direct implications for the sensitivity of the estimated coefficient to unmeasured confounding.
Consider the following structural equation
where a linear regression of $Y$ on $A$ and $X$ identifies the structural parameter $\beta$. In this setting, the absence of unmeasured confounding implies that strong multicollinearity between the exposure $A$ and the measured covariates $X$ does not introduce bias, but instead affects only the precision of the estimated treatment effect by inflating its standard error. If the outcome is conditionally mean-independent of $X$ given $A$ (given by $\theta_X = 0$), including $X$ is unnecessary for unbiasedness and may reduce efficiency wooldridge2016should. In practice, however, the assumption of no unobserved confounding is difficult to verify, and bias from omitted variables may persist despite adjustment using imperfect proxies.
A well-known illustration with substantial multicollinearity is the Equality of Educational Opportunity (Coleman) Report, which found that after controlling for students’ family background and proxies for socioeconomic status, differences in measured school resources like funding, teacher qualifications, and available facilities, had limited explanatory power for academic achievement, compared to family background and socioeconomic status coleman1968equality.
These conclusions have since been challenged by modern experimental and \sloppy{quasi-experimental} evidence showing that specific school inputs can have meaningful causal effects on achievement, including class-size reductions and teacher quality angrist1999using, chetty2011does. One possible reconciliation of this paradox is that socioeconomic status is a latent construct that is only imperfectly captured by observed proxies such as parental education, family income, and home resources. When these proxies are highly correlated with school resources, strong multicollinearity combined with proxy measurement error can increase the sensitivity of causal estimates to residual unmeasured confounding, potentially exacerbating omitted-variable bias in observational analyses such as those in the Coleman Report.
The issue of multicollinearity in measured covariates has been particularly salient in the evolution of smoking behavior over the last few decades. Socioeconomic gradients have long been associated with heterogeneity in smoking exposure, and the cigarette epidemic is commonly understood to have evolved through four stages lopez1994descriptive, mackenbach2006health. mackenbach2006health describes smoking as following a diffusion pattern across socioeconomic groups: early adoption occurs primarily among higher–socioeconomic-status men, but as smoking spreads more broadly, it becomes increasingly evenly distributed across the population, weakening socioeconomic gradients. In later stages, smoking prevalence declines first among more advantaged groups while remaining concentrated among less affluent populations, eventually becoming predominantly a habit of lower socioeconomic groups. These patterns have been empirically documented in subsequent studies pampel2005diffusion, vedoy2014tracing, quirmbach2016gender, di2019smoking, garrett2019socioeconomic.
To aid this discussion, we analyze data from the National Health and Nutrition Examination Survey (NHANES), a large-scale program designed to assess the health and nutritional status of adults and children in the United States nhanes. To examine changes in smoking exposure over time, we focus on two survey periods for which NHANES data are available: 1971–1974 (NHANES I) and 2015–2016. We fit logistic propensity score models following rosenbaum1983central for current smoking status in each period, adjusting for age at interview, race, highest grade completed, and poverty index, and stratifying by sex. Using NHANES I data, the model achieves a C-statistic of 0.61 for males and 0.65 for females; in contrast, the corresponding models using 2015–2016 NHANES data yield a C-statistic of approximately 0.70 for both sexes. This increase in predictive performance indicates that smoking status has become more strongly determined by observed demographic and socioeconomic covariates over time, suggesting that multicollinearity between smoking exposure and confounders is more pronounced today than it was fifty years ago.
Consistent with this observation, patterns in the logistic regression coefficients in Table (ref) indicate that smoking is more tightly associated with socioeconomic and demographic characteristics in the contemporary data than in the early 1970s. Together, these findings motivate the \sloppy{real-data} counterpart to the conundrum posed in Figure (ref): are effect estimates of smoking on lung cancer today more sensitive to potential unmeasured confounding than they were in earlier periods?
Consider an observed outcome $Y$, exposure of interest $A$, measured covariate $X$ and unmeasured covariate $U$, governed by the following structural equation \footnote{A structural equation specifies how a variable is generated as a function of other variables and an error term, encoding assumed causal relationships in the data-generating process goldberger1991course.}.
Furthermore, suppose the measured covariate $X$ is a proxy for $U$, given by
In addition, we also impose the following restriction.
In simple words, Assumption (ref) is satisfied if $A$ and $X$ do not share any common cause except $U$. Note that, if a linear structural equation also holds for $A$, then Assumption (ref) also implies that $X$ is not a direct cause of $A$, and the only association between $A$ and $X$ occurs through $U$.
Proposition (ref), the proof of which can be found in Appendix A, decomposes the bias into three distinct components: (i) $\gamma$, the confounding strength of $U$ on $Y$, (ii) $\text{Var}(\varepsilon_X)$, which encapsulates the quality of $X$ as a proxy of $U$, and (iii) the observable strength of collinearity between $A$ and $X$, quantified by $\beta_{A\sim X}/(\text{Var}(A)(1-R^2_{A\sim X}))$. In particular, as the association between $X$ and $A$ strengthens, reflected in a larger effect size $\beta_{A\sim X}$, while the residual variance of $A$ (i.e., $\text{Var}(A)(1-R^2_{A\sim X})$) remains stable, the discrepancy between the estimated coefficient $\beta_{Y\sim A,X}$ and the true causal parameter $\beta$ increases. This amplification heightens the sensitivity of causal conclusions to residual unmeasured confounding.
Notably, Proposition (ref) imposes no structural assumptions on the relationship between $A$ and $U$ beyond Assumption (ref), and is agnostic to the functional form governing this association. The bias amplification it describes depends only on observable OLS regression coefficients of $A\sim X$, regardless of whether such a linear relationship holds. Furthermore, Proposition (ref) does not require $\theta_X$ to be null, and therefore its implications remain valid even when $X$ functions as a negative control outcome.\footnote{A negative control outcome is a variable that is not causally affected by the exposure of interest but may be associated with the underlying sources of bias.}
Proposition (ref) offers a method to assess whether effect estimates of smoking on lung cancer from later data are more sensitive to unmeasured confounding. Let $Y$ denote lung-cancer incidence, $A$ smoking prevalence, $U$ the true unmeasured socioeconomic status and $X$ a measured proxy of $U$, such as the poverty index. We expect the structural equations governing $Y$ to be relatively stable across time periods, so that Equation (ref) holds in both periods of our analysis. We make an analogous assumption for Equation (ref), arguing that the relationship between socioeconomic status and measured covariates such as race and poverty index has remained relatively stable over time.
Note that the above conclusion is valid under Assumption (ref), which excludes additional common causes between $A$ and $X$ beyond $U$. While violations of this assumption can complicate the relationship between $\beta_{Y\sim A,X}$ and $\beta$ (see Appendix A), the result remains most informative for covariates that plausibly act as proxies for latent socioeconomic status, such as the poverty index in the smoking application, and that are unlikely to directly affect smoking behavior. With additional measured covariates, analogous conclusions hold conditional on those covariates.
Table (ref) summarizes the linear regression coefficients for smoking uptake on measured covariates, including age at interview, race, highest education grade earned, and poverty index. The results indicate that the association between poverty index ($X$) and smoking uptake ($A$) has strengthened substantially over time, while the residual variance of smoking uptake has remained relatively stable. Under the assumption that poverty index does not directly affect smoking uptake except through latent socioeconomic status, we examine confidence intervals for the ratio of the poverty index coefficient to the residual variance across sexes and time periods as a diagnostic for differences in sensitivity to unmeasured confounding.
Conservative confidence intervals for the ratio $\beta_{A\sim X}/(\text{Var}(A)(1-R^2_{A\sim X}))$, controlling for all other measured covariates, are reported in Table (ref). The lack of overlap between these intervals across time periods, consistent across sexes, illustrates how the sensitivity of smoking–lung cancer effect estimates to unmeasured confounding has heightened over time. More pertinently, this exercise highlights the need for heightened caution in assessing the sensitivity of causal effects in settings where proxies for latent sources of bias become increasingly strongly associated with the exposure of interest.
Sensitivity analysis serves as an important tool for assessing the credibility of causal claims, but it cannot substitute for careful consideration of measured covariates. While high association between treatment uptake and observed covariates has traditionally been viewed as a problem of statistical power, there has been limited attention on its implications for validity and sensitivity. In this article, we argue that when latent confounders are difficult to adjust for directly and are instead controlled for using proxies, stronger associations between the exposure and measured proxies imply effect estimates that are more prone to exacerbated omitted-variable bias.
We show that the ratio of the measured covariate coefficient in the exposure model to the residual variance of the exposure provides a natural, observable metric for evaluating this added sensitivity in linear regression based analyses. Applying this framework to smoking, we demonstrate that smoking exposure has become increasingly intertwined with socioeconomic factors over time, and the increased multicollinearity makes causal assertions from more recent data less robust to unmeasured confounding. We hope this work serves as a cautionary note when assessing the sensitivity of causal conclusions to potential unmeasured confounding, particularly when strong multicollinearity exists between exposure and covariates intended to proxy latent confounders.
{The authors are grateful to the participants of the Wharton Causal Data Science Lab for their insightful discussions and thoughtful feedback.}