Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
73,392 characters · 16 sections · 58 citation commands
Distributionally Robust Instrumental Variables Estimation
\def\spacingset#1{ {#1}} \spacingset{1}
\if11 \fi
\if01 \fi
{\it Keywords:} Causal Inference; Distributionally Robust Optimization; Square Root Ridge; Invalid Instruments; Distribution Shift
\spacingset{1.8}
Instrumental variables (IV) estimation, also known as IV regression, is a fundamental method in econometrics and statistics to infer causal relationships in observational data with unobserved confounding. It leverages access to additional variables (instruments) that affect the outcome exogenously and exclusively through the endogenous regressor to yield consistent causal estimates, even when the standard ordinary least squares (OLS) estimator is biased by unobserved confounding imbens1994identification,angrist1996identification,imbens2015causal. Over the years, IV estimation has become an indispensable tool for causal inference in empirical works in economics card1994minimum, as well as in the study of genetic and epidemiological data davey2003mendelian.
Despite the widespread use of IV in empirical and applied works, it has important limitations and challenges, such as invalid instruments sargan1958estimation,murray2006avoiding, weak instruments staiger1997instrumental, non-compliance imbens1994identification, and heteroskedasticity, especially in settings with weak instruments or highly leveraged datasets andrews2019weak,young2022consistency. These issues could significantly impact the validity and quality of estimation and inference using instrumental variables jiang2017have. Many works have since been devoted to assessing and addressing these issues, such as statistical tests hansen1982large,stock2002testing, sensitivity analysis rosenbaum1983assessing,bonhomme2022minimizing, and additional assumptions or structures on the data generating process kolesar2015identification,kang2016instrumental,guo2018confidence.
Recently, an emerging line of works have highlighted interesting connections between causality and the concepts of invariance and robustness peters2016causal,meinshausen2018causality,rothenhausler2021anchor,buhlmann2020invariance,jakobsen2022distributional,fan2024environment. Their guiding philosophy is that causal properties can be viewed as robustness against changes across heterogeneous environments, represented by a set $\mathcal{P}$ of data distributions. The robustness of an estimator against $\mathcal{P}$ is often represented in a distributionally robust optimization (DRO) framework via the min-max problem
where $\ell(W;\beta)$ is a loss function of data $W$ and parameter $\beta$ of interest.
In many estimation and regression settings, one assumes that the true data distribution satisfies certain conditions, e.g., conditional independence or moment equations. Such conditions guarantee that standard statistical procedures based on the empirical distribution $\mathbb{P}_0$ of data, such as M-estimation and generalized method of moments (GMM), enable valid estimation and inference. In practice, however, it is often reasonable to expect that the distribution $\mathbb{P}_0$ of the observed data might deviate from that generated by the {ideal} model that satisfies such conditions, e.g., due to measurement errors or model mis-specifications. DRO addresses such distributional uncertainties by explicitly incorporating possible deviations into an ambiguity set $\mathcal{P}=\mathcal{P}(\mathbb{P}_0,\rho)$ of distributions that are “close” to $\mathbb{P}_0$. The parameter $\rho$ quantifies the degree of uncertainty, e.g., as the radius of a ball centered around $\mathbb{P}_0$ and defined by some divergence measure between probability distributions. By minimizing the worst-case loss over $\mathcal{P}(\mathbb{P}_0,\rho)$ in the min-max optimization problem (ref), the DRO approach achieves robustness against deviations captured by $\mathcal{P}(\mathbb{P}_0,\rho)$.
DRO provides a useful perspective for understanding the robustness properties of statistical methods. For example, it is well-known that members of the family of $k$-class estimators anderson1949estimation,nagar1959bias,theil1961economic are more robust than the standard IV estimator against weak instruments andrews2007inference. Recent works by rothenhausler2021anchor and jakobsen2022distributional show that $k$-class estimators in fact have a DRO representation of the form (ref), where $\ell$ is the square loss, $W=(X,Y)$, and $X,Y$ are endogenous and outcome variables generated from structural equation models parameterized by the natural parameter of $k$-class estimators. See Appendix (ref) for details. The general robust optimization problem (ref) can trace its roots in the classical robust statistics literature huber1964robust,huber2011robust as well as classic works on robustness in economics hansen2008robustness. Drawing inspirations from them, recent works in econometrics have also explored the use of robust optimization to account for (local) deviations from model assumptions kitamura2013robustness,armstrong2021sensitivity,chen2021robust,bonhomme2022minimizing,adjaho2022externally,fan2023quantifying. These works, together with works on invariance and robustness, highlight the emerging interactions between econometrics, statistics, and robust optimization.
Despite new developments connecting causality and robustness, many questions and opportunities remain. An important challenge in DRO is the choice of the ambiguity set $\mathcal{P}(\mathbb{P}_0,\rho)$ to adequately capture distributional uncertainties. This choice is highly dependent on the structure of the particular problem of interest. While some existing DRO approaches use ambiguity sets $\mathcal{P}(\mathbb{P}_0,\rho)$ based on {marginal} or joint distributions of data, such $\mathcal{P}(\mathbb{P}_0,\rho)$ may not effectively capture the structure of IV estimation models. In addition, as the min-max problem (ref) minimizes the loss function under the worst-case distribution in $\mathcal{P}(\mathbb{P}_0,\rho)$, a common concern is that the resulting estimator is too conservative when $\mathcal{P}(\mathbb{P}_0,\rho)$ is too large. In particular, although DRO estimators enjoy better empirical performance in finite samples, their asymptotic validity typically requires the ambiguity set to vanish to a singleton, i.e., $\rho \rightarrow 0$ blanchet2019robust, blanchet2022confidence. However, in the context of IV estimation, distributional uncertainties about untestable model assumptions could persist in large samples, necessitating the need for an ambiguity set that does not vanish to a singleton. It is therefore important to ask whether and how one can construct an estimator in the IV estimation setting that can sufficiently capture the distributional uncertainties about model assumptions, and at the same time remains asymptotically valid with a non-vanishing robustness parameter.
In this paper, we propose to view common challenges to IV estimation through the lens of DRO, whereby uncertainties about model assumptions, such as the exclusion restriction and homoskedasticity, are captured by a suitably chosen ambiguity set in (ref). Based on this perspective, we propose DRIVE, a general DRO approach to IV estimation. Instead of constructing the ambiguity set based on marginal or joint distributions as in existing works, we construct $\mathcal{P}(\mathbb{P}_0,\rho)$ from distributions conditional on the instrumental variables. More precisely, we construct $\mathbb{P}_0$ as the empirical distribution of outcome and endogenous variables $Y,X$ projected onto the space spanned by instrumental variables. When the ambiguity set of DRIVE is based on the 2-Wasserstein metric, we show that the resulting estimator minimizes a square root version of ridge regularized two stage least squares (TSLS) objective, where the radius $\rho$ of the ambiguity set becomes the regularization parameter. This regularized regression formulation relies on the general duality of Wasserstein DRO problems gao2016distributionally,blanchet2019robust,kuhn2019wasserstein.
We next next reveal a surprising statistical property of the square root ridge by showing that Wasserstein DRIVE is consistent as long as the regularization parameter $\rho$ is bounded above by an estimable constant, which depends on the first stage coefficient of the IV model and can be interpreted as a measure of instrument quality. To our knowledge, this is the first consistency result for regularized regression estimators where the regularization parameter does not vanish as the sample size $n \rightarrow \infty$. One implication of our results is that Wasserstein DRIVE, being a regularized regression estimator, enjoys better finite sample properties, but does not introduce bias asymptotically even for non-vanishing $\rho$, unlike standard regularized regression estimators such as the ridge and LASSO.
We further characterize the asymptotic distribution of Wasserstein DRIVE and propose data-driven procedures to select the regularization parameter. We demonstrate with numerical experiments that Wasserstein DRIVE improves over the finite sample performance of IV and $k$-class estimators, thanks to its ridge type regularization, while at the same time retaining asymptotic validity whenever instruments are valid. In particular, Wasserstein DRIVE achieves significant improvements in mean squared errors (MSEs) over IV and OLS when instruments are moderately invalid. These findings suggest that Wasserstein DRIVE can be an attractive option in practice when we are concerned about model assumptions.
The rest of the paper is organized as follows. In (ref), we discuss the standard IV estimation framework and common challenges. In (ref), we propose the Wasserstein DRIVE framework and provide the duality theory. In (ref), we develop asymptotic results for the Wasserstein DRIVE, including consistency under a non-vanishing robustness/regularization parameter. (ref) conducts numerical studies that compare Wasserstein DRIVE with other estimators including IV, OLS, and $k$-class estimators. Background materials, proofs, and additional results are included in the appendices in the supplementary material.
Notation. Throughout the paper, $\|v \|_p$ denotes the $p$-norm of a vector $v$, while $\|v \| := \|v \|_2$ denotes the Euclidean norm. $\text{Tr}(M)$ denotes the trace of a matrix $M$. $\lambda_{k}(M)$ represents the $k$-th largest eigenvalue of a symmetric matrix $M$. Boldfaced variables, such as $\mathbf{X}$, represents a matrix whose $i$-th row is the variable $X_i$.
In this section, we first provide a brief review of the standard IV estimation framework. We then motivate the DRO approach to IV estimation by viewing common challenges from the perspective of distributional uncertainties. In (ref), we propose the Wasserstein distributionally robust instrumental variables estimation (DRIVE) framework.
Consider the following standard linear instrumental variables regression model with $X\in\mathbb{R}^{p},Z\in\mathbb{R}^{d}$ where $d\geq p$, and $\beta_{0}\in\mathbb{R}^{p},\gamma\in\mathbb{R}^{d\times p}$:
In (ref), $X$ are the endogenous variables, $Z$ are the instrumental variables, and $Y$ is the outcome variable. The error terms $\epsilon$ and $\xi$ capture the unobserved (or residual) components of $Y$ and $X$, respectively. We are interested in estimating the causal effects $\beta_0$ of the endogenous variables $X$ on the outcome variable $Y$ given independent and identically distributed (i.i.d.) samples $\{X_i, Y_i, Z_i\}_{i=1}^{n}$. However, $X$ and $Y$ are confounded through some unobserved confounders $U$ that are correlated with both $Y$ and $X$, represented graphically in the directed acyclic graph (DAG) below:
Mathematically, the unobserved confounding can be described by the moment condition \[\mathbb{E}\left[X\epsilon\right]\neq \mathbf{0}.\] As a result of the unobserved confounding, the standard ordinary least squares (OLS) regression estimator of $\beta_0$ that regresses $Y$ on $X$ is biased. To address this problem, the IV estimation approach leverages access to the instrumental variables $Z$, also often called instruments, which satisfy the moment conditions
Under these conditions, a popular IV estimator is the two stage least squares (TSLS, sometimes also stylized as 2SLS) estimator theil1953repeated. With $\Pi_{Z}:=\mathbf{Z}(\mathbf{Z}^{T}\mathbf{Z})^{-1}\mathbf{Z}^{T}$ and $\mathbf{X,Y,Z}$ matrix representations of $\{X_i, Y_i, Z_i\}_{i=1}^{n}$ whose $i$-th rows correspond to $X_i, Y_i, Z_i$, respectively, the TSLS estimator $\hat \beta^{\text{IV}}:=(\mathbf{X}^T\Pi_{Z}\mathbf{X})^{-1}\mathbf{X}^T\Pi_{Z}\mathbf{Y}$ minimizes the objective
In contrast, the standard OLS estimator $\hat \beta^{\text{OLS}}$ solves the problem
When the moment conditions (ref) and (ref) hold, the TSLS estimator is a consistent estimator of the causal effect $\beta_0$ under standard assumptions wooldridge2020introductory, and valid inference can be performed by constructing variance estimators based on the asymptotic distribution of $\hat \beta^{\text{IV}}$ imbens2015causal. Although not the most common presentation of TSLS, the optimization formulation in (ref) provides intuition on how IV estimation works: when the instruments $Z$ are uncorrelated with the unobserved confounders $U$ affecting $X$ and $Y$, the projection operator $\Pi_{Z}$ applied to $\mathbf{X}$ “removes” the confounding from $X$, so that $\Pi_{Z}\mathbf{X}$ becomes (asymptotically) uncorrelated with $\epsilon$. Regressing $\mathbf{Y}$ on $\Pi_{Z}\mathbf{X}$ then yields a consistent estimator of $\beta_0$.
The validity of estimation and inference based on $\hat \beta^{\text{IV}}$ relies critically on the moment conditions (ref) and (ref). Condition (ref) is often called the relevance condition or rank condition, and requires $\mathbb{E}\left[ZX^T\right]$ to have full rank (recall $p\leq d$). In the special case of one-dimensional instrumental and endogenous variables, i.e., $d=p=1$, it simply reduces to $\mathbb{E}\left[ZX\right]\neq 0$. Intuitively, the relevant condition ensures that the instruments $Z$ can explain sufficient variations in the endogenous variables $X$. In this case, the instruments are said to be relevant and strong. When $\mathbb{E}\left[ZX^T\right]$ is close to being rank deficient, i.e., the smallest eigenvalue $\lambda_p(\mathbb{E}\left[ZX^T\right])\approx 0$, IV estimation suffers from the so-called weak instrument problem, which results in many issues in estimation and inference, such as small sample bias and non-normal statistics stock2002survey. Some $k$-class estimators, such as limited information maximum likelihood (LIML) anderson1949estimation, are partially motivated to address these problems. Condition (ref) is often referred to as the exclusion restriction or instrument exogeneity imbens2015causal, and instruments that satisfy this condition are called valid instruments. When an instrument $Z$ is correlated with the unobserved confounder that confounds $X,Y$, or when $Z$ affects the outcome $Y$ through an unobserved variable other than the endogenous variable $X$, the instrument becomes invalid, resulting in biased estimation and invalid inference of $\beta_0$ murray2006avoiding. These issues can often be exacerbated when the instruments are weak, when there is heteroskedasticity andrews2019weak, or the data is highly leveraged young2022consistency.
Although many works have been devoted to addressing the problems of weak and invalid instruments, there are fundamental limits on the extent to which one can test for these issues. Given the popularity of IV estimation in practice, it is therefore desirable to have estimation and inference procedures that are robust to the presence of such issues. Our work is precisely motivated by these considerations. Compared to many existing robust approaches to IV estimation, we take a more agnostic approach via distributionally robust optimization. More precisely, we argue that many common challenges in IV estimation can be viewed as uncertainties about the data distribution, i.e., deviations from the ideal model that satisfies IV assumptions, which can be explicitly taken into account by choosing an appropriate ambiguity set in a DRO formulation of the standard IV estimation. To demonstrate this perspective more concretely, we now examine some common problems in IV estimation and show that they can be viewed as distributional shifts under a suitable metric, and therefore amenable to a DRO approach.
Consider now the following one-dimensional version of the IV model in (ref)
where we assume that $X,Y$ are confounded by an unobserved confounder $U$ through \[\epsilon=Z\eta+U, \quad \xi=U.\] Note that in addition to $U$, there is also potentially a direct effect $\eta$ from the instrument $Z$ to the outcome variable $Y$. We focus on the resulting model for our subsequent discussions:
The standard IV assumptions can be succinctly summarized for (ref). The relevance condition (ref) requires that $\gamma\neq0$, while the exclusion restriction (ref) requires that $Z$ is uncorrelated with $U$ and that in addition $\eta=0$. Assume that $U,Z$ are i.i.d. standard normal. $X,Y$ are then determined by $\eqref{eq:simple-IV}$. We are interested in the shifts in data distribution, appropriately defined and measured, when the exogeneity and relevance conditions are violated.
Next, we consider the distributional shift resulting from heteroskedastic errors, which are known to yield the TSLS estimator inefficient and the standard variance estimator invalid baum2003instrumental. Some $k$-class estimators, such as the LIML and the Fuller estimators, also become inconsistent under heteroskedasticity hausman2012instrumental.
The preceding discussions demonstrate that distributional uncertainties resulting from violations of common model assumptions in IV estimation are well captured by the 2-Wasserstein distance on the distributions of the conditional variables $(\tilde{X},\tilde{Y})$. We therefore propose to construct an ambiguity set in (ref) using a Wasserstein ball around the empirical distribution on $(\tilde{X},\tilde{Y})$. We provide details of this framework in the next section.
In this section, we propose a distributionally robust IV estimation framework. We propose to use Wasserstein ambiguity sets to account for distributional uncertainties in IV estimation. We develop the dual formulation of Wasserstein DRIVE as regularized regression, and discuss its connections and distinctions to other regularized regression estimators.
Motivated by the intuition that common challenges to IV estimation in practice, such as violations of model assumptions, can be viewed as distributional uncertainties on the conditional distributions of $(\tilde{X},\tilde{Y})=(X\mid Z, Y\mid Z)$, we propose the Distributionally Robust IV Estimation (DRIVE) framework, which solves the following DRO problem given a dataset $\{(X_i,Y_i,Z_i)\}_{i=1}^n$ and robustness parameter $\rho$:
where $\tilde {\mathbb{P}}_n(\mathcal{X}\times \mathcal{Y})$ is the empirical distribution on $(X,Y)$ induced by the projected samples \[\{\tilde X_i,\tilde Y_i\}_{i=1}^{n}\equiv \{(\Pi_{\mathbf{Z}}\mathbf{X})_i,(\Pi_{\mathbf{Z}}\mathbf{Y})_i\}_{i=1}^n.\] Here $\mathbf{X}\in \mathbb{R}^{n\times p},\mathbf{Y}\in \mathbb{R}^{n},\mathbf{Z}\in \mathbb{R}^{n\times d}$ are the matrix representations of observations, and $\Pi_{\mathbf{Z}} = \mathbf{Z}(\mathbf{Z}^T\mathbf{Z})^{-1}\mathbf{Z}^T$ is the projection matrix onto the column space of $\mathbf{Z}$. $D(\cdot,\cdot)$ is a metric or divergence measure on the space of probability distributions on $\mathcal{X}\times \mathcal{Y}$. Therefore, in our DRIVE framework, we first regress both the outcome $Y$ and covariate $X$ on the instrument $Z$ to form the $n$ predicted samples $(\Pi_{\mathbf{Z}}\mathbf{X},\Pi_{\mathbf{Z}}\mathbf{Y}$). Then an ambiguity set is constructed using $D$ around the empirical distribution $\tilde {\mathbb{P}}_n$. This choice of the reference distribution $\mathbb{P}_0$ is a key distinction of our work from previous works that leverage DRO in statistical models. In the standard regression/classification setting, the reference distribution is often chosen as the empirical distribution $\hat{\mathbb{P}}_n$ on $\{X_i, Y_i\}_{i=1}^{n}$ blanchet2019robust. In the IV estimation setting where we have additional access to instruments $Z$, we have the choice of constructing ambiguity sets around the empirical distribution on the marginal quantities $\{(X_i,Y_i,Z_i)\}_{i=1}^n$, which is the approach taken in bertsimas2022distributionally. In contrast, we choose to use the empirical distribution on the conditional quantities $\{(\tilde X_i,\tilde Y_i)\}_{i=1}^n$. This choice is motivated by the intuition that violations of IV assumptions can be captured by conditional distributional shifts, as illustrated by examples in the previous section.
The choice of the divergence measure $D(\cdot,\cdot)$ is also important, as it characterizes the potential distributional uncertainties that DRIVE is robust to. In this paper, we propose to use the 2-Wasserstein distance $W_2(\mu,\nu)$ between two probability distributions $\mu,\nu$ mohajerin2018data,gao2016distributionally. One advantage of the Wasserstein distance is the tractability of its associated DRO problems blanchet2019robust, which can often be formulated as regularized regression problems with unique solutions. See also Appendix (ref). In (ref), we provided several examples that demonstrate the 2-Wasserstein distance is able to capture common distributional uncertainties in the IV estimation setting. Alternative distance measures of probability distributions, such as the class of $\phi$-divergences ben2013robust, can also be used instead of the Wasserstein distance. For example, kitamura2013robustness use the Hellinger distance to model local perturbations in robust estimation under moment restrictions, although not in the IV estimation setting. In this paper, we focus on the Wasserstein DRIVE framework based on $D=W_2$, and leave studies of DRIVE with other choices of $D$ to future works.
We next begin our formal study of Wasserstein DRIVE. In (ref), we will show that the Wasserstein DRIVE objective is dual to a convex regularized regression problem. As a result, the solution to the optimization problem (ref) is well-defined, and we denote this estimator by $\hat{\beta}_{\mathrm{DRIVE}}$. In (ref), we show $\hat{\beta}_{\mathrm{DRIVE}}$ is consistent with potentially non-vanishing choices of the robustness parameter and derive its asymptotic distribution.
It is well-known in the optimization literature that min-max optimization problems such as (ref) often have equivalent formulations as regularized regression problems. This correspondence between regularization and robustness already manifests itself in the ridge regression, which is equivalent to an $\ell_{2}$-robust OLS regression bertsimas2018characterization. Importantly, the regularized regression formulations are often more tractable in terms of solving the resulting optimization problem, and also facilitate the study of the statistical properties of the estimators. We first show that the Wasserstein DRIVE objective can also be written as a regularized regression problem similar to, but distinct from, the standard TSLS objective with ridge regularization. Proofs can be found in Appendix (ref).
Note that the robustness parameter $\rho$ of the DRO formulation (ref) now has the dual interpretation as the regularization parameter in (ref). This convex regularized regression formulation implies that the min-max problem (ref) associated with Wassertein DRIVE has a unique solution, thanks to the strict convexity of the regularization term $\sqrt{\rho(\|\beta\|^{2}+1)}$, and is easy to compute despite not having a closed form solution. In particular, (ref) can be reformulated as a standard second order conic program (SOCP) el1997robust, which can be solved efficiently with off-the-shelf convex optimization routines, such as CVX. More importantly, we leverage this formulation of Wasserstein DRIVE as a regularized regression problem to study its novel statistical properties in (ref).
The equivalence between Wasserstein DRO problems and regularized regression problem is a familiar general result in recent works. For example, blanchet2019robust and gao2016distributionally derive similar duality results for distributionally robust regression with $q$-Wasserstein distances for $q>1$. Compared to previous works, our work is distinct in the following aspects. First, we apply Wasserstein DRO to the IV estimation setting instead of standard regression settings, such as OLS or logistic regression. {Although from an optimization point of view there is no substantial difference, the IV setting motivates a new asymptotic regime that uncovers interesting statistical properties of the resulting estimators.} Second, the regularization term in (ref) is distinct from those in previous works, which often use $\|\beta\|_p$ with $p\geq 1$. This seemingly innocuous difference turns out to be crucial for our novel results on the Wasserstein DRIVE. Lastly, compared to the proof in blanchet2019robust, our proof of (ref) is based on a different argument using the Sherman-Morrison formula instead of H\"{o}lder's inequality, which provides an independent proof of the important duality result for Wasserstein distributionally robust optimization.
The regularized regression formulation of the Wasserstein DRIVE problem in (ref) resembles the standard ridge regularized hoerl1970ridge TSLS regression:
We therefore refer to (ref) as the square root ridge regularized TSLS. However, there are three major distinctions between (ref) and (ref) that are essential in guaranteeing the statistical properties of Wasserstein DRIVE not enjoyed by the standard ridge regularized TSLS. First, the presence of square root operations on both the risk term and the penalty term; second, the presence of a constant in the regularization term; third, an additional projection on the outcomes in $\Pi_{\mathbf{Z}}\mathbf{Y}$. We further elaborate on these features in (ref).
In the standard regression setting without instrumental variables, the square root ridge
also resembles the {“square root LASSO”} of belloni2011square:
In particular, both can be written as dual problems of Wasserstein DRO problems blanchet2019robust. However, the square root LASSO is motivated by high-dimensional regression settings where the dimension of $X$ is potentially larger than the sample size $n$, but $\beta$ is very sparse. In contrast, our study of the square root ridge is motivated by its robustness properties in the IV estimation setting, where the dimension of the endogenous variable is small (often one-dimensional). In other words, variable selection is not the main focus of this paper. A variant of the square root ridge estimator in (ref) was also considered in the standard regression setting by owen2007robust, who instead uses the penalty term $\|\beta\|_2$.
As is well-known in the regularized regression literature fu2000asymptotics, when the regularization parameter decays to 0 at a rate $O_p(1/\sqrt{n})$, the ridge estimator is consistent. A similar result also holds for the square root ridge (ref) in the standard regression setting as $\rho \rightarrow 0$. However, in the IV estimation setting, our distributional uncertainties about model assumptions, such as the validity of instruments, could persist even in large samples. Recall that $\rho$ is also the robustness parameter in the DRO formulation (ref). Therefore, the usual requirement that $\rho \rightarrow 0$ as $n \rightarrow \infty$ cannot adequately capture distributional uncertainties in the IV estimation setting. In the next section, we study the asymptotic properties of Wasserstein DRIVE when $\rho$ does not necessarily vanish. In particular, we establish the consistency of Wasserstein DRIVE leveraging the three distinct features of (ref) that are absent in the standard ridge regularized TSLS regression (ref). This asymptotic result is in stark contrast to the conventional wisdom on regularized regression that regularized regression achieves lower variance at the cost of non-zero bias.
In this section, we leverage distinct geometric features of the square root ridge regression to study the asymptotic properties of the Wasserstein DRIVE. In (ref), we show that the Wasserstein DRIVE estimator is consistent for any $\rho \in [0,\overline \rho]$, where $\overline{\rho}$ depends on the first stage coefficient $\gamma$. This property is a consequence of the consistency of the square root ridge estimator in settings where the objective value at the true parameter vanishes, such as the GMM estimation setting. It ensures that Wasserstein DRIVE can achieve better finite sample performance thanks to its ridge type regularization, while at the same time retaining asymptotic validity when instruments are valid. In (ref), we characterize the asymptotic distribution of Wasserstein DRIVE, and discuss several special settings particularly relevant in practice, such as the just-identified setting with one-dimensional instrumental and endogenous variables.
Recall the linear IV regression model in (ref)
where $X\in\mathbb{R}^{p}, Z\in\mathbb{R}^{d}$, and $\beta_{0}\in\mathbb{R}^{p},\gamma\in\mathbb{R}^{d\times p}$ with $d\geq p$ to ensure identification. In this section, we make the standard assumptions that the instruments satisfy the relevance and exogeneity conditions in (ref) and (ref), $\epsilon,\xi$ are homoskedastic, the instruments $Z$ are not perfectly collinear, and that $\mathbb{E}\|Z\|^{2k}<\infty, \mathbb{E}\|\xi\|^{2k}<\infty, \mathbb{E}|\epsilon|^{2k}<\infty$ for some $k>2$. The results can be extended in a straightforward manner when we relax these assumptions, e.g., only requiring that exogeneity holds asymptotically. Given i.i.d. samples from the linear IV model, recall the regularized regression formulation of the Wasserstein DRIVE objective
where $\Pi_{\mathbf{Z}}=\mathbf{Z}(\mathbf{Z}^{T}\mathbf{Z})^{-1}\mathbf{Z}^{T}\in \mathbb{R}^{n\times n}$, and $\Pi_{\mathbf{Z}}\mathbf{Y}\in\mathbb{R}^{n}$ and $\Pi_{\mathbf{Z}}\mathbf{X}\in \mathbb{R}^{n\times p}$ are $\mathbf{Y}\in\mathbb{R}^{n},\mathbf{X}\in\mathbb{R}^{n\times p}$ projected onto the instrument space spanned by $\mathbf{Z}\in\mathbb{R}^{n\times d}$.
Therefore, Wasserstein DRIVE is consistent as long as the limit of the regularization parameter is bounded above by $\overline{\rho }=\lambda_{p}(\gamma^T \Sigma_Z \gamma)$. In the case when $\Sigma_Z = \sigma^2_Z I_d$ for $\sigma_Z^2>0$, the upper bound $\overline{\rho }$ is proportional to the square of the smallest singular value of the first stage coefficient $\gamma$, which is positive under the relevance condition (ref). Recall that $\rho_n$ is the radius of the Wasserstein ball in the min-max formulation of DRIVE in (ref). (ref) therefore guarantees that even when the robustness parameter $\rho_n \equiv \rho \neq 0$, which implies the solution to the min-max problem is different from the TSLS estimator ($\rho =0$), the resulting estimator always has the same limit as long as $\rho \leq \lambda_{p}(\gamma^T \Sigma_Z \gamma)$.
The significance of (ref) is twofold. First, it provides a meaningful interpretation of the robustness parameter $\rho$ in Wasserstein DRIVE in terms of problem parameters, more precisely the variance covariance matrix $\Sigma_Z$ of $Z$ and the first stage coefficient $\gamma$ in the IV regression model. The maximum amount of robustness that can be afforded by Wasserstein DRIVE without sacrificing consistency is $\lambda_{p}(\gamma^T \Sigma_Z \gamma)$, which directly depends on the strength and variance of the instrument. This relation can be described more precisely when $\Sigma_Z = \sigma^2_Z I_d$, in which case the bound is proportional to $\sigma^2_Z$ and $\lambda_p(\gamma^T\gamma)$. Both quantities improve the quality of the instruments: $\sigma^2_Z$ improves the proportion of variance of $X$ and $Y$ explained by the instrument vs. noise, while a $\gamma$ far from rank deficiency avoids the weak instrument problem. Therefore, the robustness of Wasserstein DRIVE depends intrinsically on the strength of the instruments. The quantity $\gamma^T \Sigma_Z \gamma$ is not unfamiliar in the IV setting, as it is proportional to the inverse of the asymptotic variance of the standard TSLS estimator when errors are homoskedastic. This observation suggests an intrinsic connection between robustness and efficiency in the IV setting. See the discussions after (ref) for more on this point.
More importantly, (ref) is the first consistency result for regularized regression estimators where the regularization parameter does not vanish with sample size. Although regularized regression such as ridge and LASSO is often associated with better finite sample performance at the cost of introducing some bias, our work demonstrates that, in the IV estimation setting, we can get the best of both worlds. On one hand, the ridge type regularization in Wasserstein DRIVE improves upon the finite sample properties of the standard IV estimators, which aligns with conventional wisdom on regularized regression. On the other hand, with a bounded level of the regularization parameter $\rho$, Wasserstein DRIVE can still achieve consistency. This is in stark contrast to existing asymptotic results on regularized regression fu2000asymptotics. Therefore, in the context of IV estimation, with Wasserstein DRIVE we can achieve consistency and a certain amount of robustness at the same time, by leveraging additional information in the form of valid instruments. The maximum degree of robustness that can be achieved also has a natural interpretation in terms of the strength and variance of the instruments.
(ref) also suggests the following procedure to construct a feasible and valid robustness/regularization parameter $\hat \rho$ given data $\{X_i, Y_i, Z_i\}_{i=1}^{n}$. Let $\hat{\gamma}=(\mathbf{Z}^{T}\mathbf{Z})^{-1}\mathbf{Z}^{T}\mathbf{X}$ be the OLS regression estimator of the first stage coefficient $\gamma$ and $\hat \Sigma_Z$ an estimator of $\Sigma_Z$, such as the heteroskedasticity-robust estimator white1980heteroskedasticity. We can use $\hat \rho \leq \lambda_p(\hat{\gamma}^T \hat \Sigma_Z\hat{\gamma})$ to construct the Wasserstein DRIVE objective, i.e., any value bounded above by the smallest eigenvalue of $\hat{\gamma}^T \hat \Sigma_Z\hat{\gamma}$. Under the assumptions in (ref), $\lambda_p(\hat{\gamma}^T\hat \Sigma_Z\hat{\gamma})\rightarrow \lambda_p({\gamma}^T\Sigma_Z{\gamma})$, which guarantees that the Wasserstein DRIVE estimator with parameter $\hat \rho$ is consistent. In Appendix (ref), we discuss the construction of feasible regularization parameters in more detail. We demonstrate the validity and superior finite sample performance of DRIVE based on these proposals in simulation studies in (ref).
One may wonder why Wasserstein DRIVE can achieve consistency with a non-zero regularization $\rho$. Here we briefly discuss the phenomenon that the limiting objective (ref)
has a unique minimizer at $\beta_0$ for bounded $\rho>0$. The first term $\sqrt{(\beta-\beta_{0})^{T}\gamma^T \Sigma_Z \gamma(\beta-\beta_{0})}$ achieves its minimum value of 0 at $\beta=\beta_0$. When $\rho$ is small, the effect of adding the regularization term $\sqrt{\rho(\|\beta\|^{2}+1)}$ does not overwhelm the first term, especially when its curvature at $\beta_0$ is large. As a result, we may expect the minimizer to not deviate much from $\beta_0$. While this intuition is reasonable qualitatively, it does not fully explain the fact that the minimizer does not change for small $\rho$. In the standard regression setting, the same intuition can be applied to the standard ridge regularization, but we know shrinkage occurs as soon as $\rho>0$. The key distinction of (ref) turns out to be the square root operations we apply to the loss and regularization terms, which endows the objective with a geometric interpretation, and ensures that the minimizer does not deviate from $\beta_0$ unless $\rho$ is above some positive threshold. We call this phenomenon the “delayed shrinkage” of the square root ridge, as shrinkage does not happen until the regularization is large enough. We illustrate it with a simple example in (ref), where the minimizer of the limiting square root objective does not change for a bounded range of $\rho$.
Lastly, we comment on the importance of projection operations in Wasserstein DRIVE. A crucial feature of the Wasserstein DRIVE objective is that both the outcome and the covariates are regressed on the instrument to compute their predicted values. In other words, the objective (ref) uses $\Pi_{\mathbf{Z}}\mathbf{Y}-\Pi_{\mathbf{Z}}\mathbf{X}\beta$ instead of $\mathbf{Y}-\Pi_{\mathbf{Z}}\mathbf{X}\beta$. For standard IV estimation ($\rho=0)$, there is no substantial difference between the two objectives, since their minimizers are exactly the same, due to the idempotent property $\Pi^2_\mathbf{Z}=\Pi_\mathbf{Z}$. In fact, in applications of TSLS, the outcome variable is often not regressed on the instrument. However, Wasserstein DRIVE is consistent for positive $\rho$ only if the outcome $Y$ is also projected onto the instrument space. In other words, the following problem does not yield a consistent estimator when $\rho > 0$:
The reason behind this phenomenon is that ${\frac{1}{n}\|\Pi_{\mathbf{Z}}\mathbf{Y}-\Pi_{\mathbf{Z}}\mathbf{X}\beta\|^{2}}$ is a GMM objective
which when $n \rightarrow \infty$ achieves a minimal value of 0 at $\beta_0$, while the limit of ${\frac{1}{n}\|\mathbf{Y}-\Pi_{\mathbf{Z}}\mathbf{X}\beta\|^{2}}$ does not vanish even at $\beta_0$. In the former case, the geometric properties of the square root ridge ensure that the minimizer of the regularized objective is $\beta_0$.
Having established the consistency of Wasserstein DRIVE with bounded $\rho$, we now turn to the characterization of its asymptotic distribution. In general, the asymptotic distribution of Wasserstein DRIVE is different from that of the standard IV estimator. However, we will also examine several special cases relevant in practice where they coincide.
Recall that the maximal robustness parameter $\rho$ of Wasserstein DRIVE while still being consistent is equal to the smallest eigenvalue of $\gamma^{T}\Sigma_Z\gamma$, which is proportional to the inverse of the asymptotic variance of the TSLS estimator. Therefore, as the efficiency of TSLS increases, so does the robustness of the associated Wasserstein DRIVE estimator. The “price” to pay for robustness when $\rho>0$ is an interesting question. It is clear from (ref) that the curvature of the population objective decreases as $\rho$ increases. Since the objective is not continuous at $\beta_0$, however, a generalized notion of curvature is needed to precisely characterize this behavior. Note that the asymptotic distribution of the TSLS estimator minimizes the objective \[(\mathcal{Z}+\Sigma_Z\gamma\delta)^{T}\Sigma_Z^{-1}(\mathcal{Z}+\Sigma_Z\gamma\delta),\] (ref) implies that in general the asymptotic distributions of Wasserstein DRIVE and TSLS are different when $\rho>0$. However, there are still several cases relevant in practice where their asymptotic distributions do coincide, which we discuss next.
In particular, the just-identified case with $d=p=1$ covers many empirical applications of IV estimation, since in practice we are often interested in the causal effect of a single endogenous variable, for which we have a single instrument. The case when $\beta_{0}\equiv \mathbf{0}$ is also very relevant, since an important question in practice is whether the causal effect of a variable is zero. Our theory suggests that the asymptotic distribution of Wasserstein DRIVE should be the same as that of the TSLS when the causal effect is zero and $d>1$, even for $\rho>0$. Based on this observation, intuitively, we should expect that Wasserstein DRIVE and TSLS estimators to be “close” to each other. If the estimators or their asymptotic variance estimators differ significantly, then $\beta_0$ may not be identically 0. We can design statistical tests by leveraging this intuition. For example, we can construct test statistics using the TSLS estimator and the DRIVE estimator with $\rho>0$, such as the difference $\hat \beta^{\text{DRIVE}}-\hat \beta^{\text{TSLS}}$. Then we can use bootstrap-based tests, such as a bootstrapped permutation test, to assess the null hypothesis that $\hat \beta^{\text{DRIVE}}-\hat \beta^{\text{TSLS}}=0$. If we fail to reject the null hypothesis, then there is evidence that the true causal effect $\beta_{0}=\mathbf{0}$.
(ref) can be seemingly pessimistic because it demonstrates that the asymptotic distribution of Wasserstein DRIVE could be the same as that of the TSLS in special cases. However, recall that Wasserstein DRIVE is formulated to minimize the worst-case risk over a set of distributions that are designed to capture deviations from model assumptions. Therefore, there is not actually any a priori reason that it should coincide with the TSLS when $\rho>0$. In this sense, the fact that the Wasserstein DRIVE is consistent with $\rho>0$ and may even coincide with TSLS is rather surprising. In the latter case, the worst-case distribution for Wasserstein DRIVE in the large sample limit must coincide with that of the standard population distribution, which may be worth further investigation.
The asymptotic results we develop in this section provide the basis on which one can perform estimation and inference with the Wasserstein DRIVE estimator. In the next section, we study the finite sample properties of DRIVE in simulation studies and demonstrate that it is superior in terms of estimation error and out of sample prediction compared to other popular estimators.
In this section, we study the empirical performance of Wasserstein DRIVE. Our results deliver three main messages. First, we demonstrate with simulations that Wasserstein DRIVE, with non-zero robustness parameter $\rho$ based on (ref), has comparable performance as the standard IV estimator whenever instruments are valid. Second, when instruments become invalid, Wasserstein DRIVE outperforms other methods in terms of RMSE. Third, on the education dataset of card1993using, Wasserstein DRIVE also has superior performance at prediction for a heterogeneous target population.
We use the data generating process
where $U,\epsilon_Z \sim \mathcal{N}(0,\sigma^2)$ and we allow a direct effect $\eta$ from the instruments $Z$ to the outcome $Y$. Moreover, the instruments $Z$ can also be correlated with the unobserved confounder $U$ ($\beta_{UZ}\neq 0)$. We fix the true parameters and generate independent datasets from the model, varying the degree of instrument invalidity. In (ref), we report the MSE of estimators averaged over 500 repeated experiments. We control the degree of instrument invalidity by varying $\eta$, the direct effect of instruments on the outcome, and $\beta_{UZ}$, the correlation between unobserved confounder and instruments. Results in (ref) are based on data where $\|\gamma\| \gg 0$ is large. We see that when instruments are strong, Wasserstein DRIVE performs as well as TSLS when instruments are valid, but performs significantly better than OLS, TSLS, anchor, and TSLS ridge when instruments become invalid. This suggests that DRIVE could be preferable in practice when we are concerned about instrument validity.
We further investigate the empirical performance of Wasserstein DRIVE when instruments are potentially invalid or weak. We present box plots (omitting outliers) of MSEs in (ref). The Wasserstein DRIVE estimator with regularization parameter $\rho$ based on bootstrapped quantiles of the score function consistently outperforms OLS, TSLS, anchor ($k$-class), and TSLS with ridge regularization. Moreover, the selected penalties increase as the direct effect of $Z$ on $Y$ or the correlation between the unobserved confounder $U$ and the instrument $Z$ increases, i.e., as the model assumption of valid instruments becomes increasingly invalid. See (ref) in Appendix (ref) for more details. This property is highly desirable, because based on the DRO formulation of DRIVE, $\rho$ represents the amount of robustness against distributional shifts associated with the estimator, which should increase as the instruments become more invalid (larger distributional shift). Box plots of estimation errors in (ref) also verify that even when instruments are valid, the finite sample performance of Wasserstein DRIVE is still better compared to the standard IV estimator, suggesting that there is no additional cost in terms of MSE when applying Wasserstein DRIVE, even when instruments are valid.
We now turn our attention to a different task that has received more attention in recent years, especially in the context of policy learning and estimating causal effects across heterogeneous populations dehejia2021local,adjaho2022externally,menzel2023transfer. We study the prediction performance of estimators when they are estimated on a dataset (training set) that has a potentially different distribution from the dataset for which they are used to make predictions (test set). We demonstrate that whenever the distributions between training and test datasets have significant differences, the prediction error of Wasserstein DRIVE is significantly smaller than that of OLS, IV, and anchor ($k$-class) estimators.
We conduct our numerical study using the classic dataset on the return of education to wage compiled by David Card card1993using. Here, the causal inference problem is estimating the effect of additional school years on the increase in wage later in life. The dataset contains demographic information about interviewed subjects. Importantly, each sample comes from one of nine regions in the United States, which differ in the average number of years of schooling and other characteristics, i.e., there are covariate shifts in data collected from different regions. Our strategy is to divide the dataset into a training set and a test set based on the relative ranks of their average years of schools, which is the endogenous variable. We expect that if there are distributional shifts between different regions, then predicting wages using education and other information using conventional models trained on the training data may not yield a good performance on the test data.
Since each sample is labeled as working in 1976 in one of nine regions in the U.S., we split the samples based on these labels, using number of years of education as the splitting variable. For example, we can construct the training set by including samples from the top 6 regions with the highest average years of schooling, and the test set to consist of samples coming from the bottom 3 regions with the lowest average years of schooling. In this case, we would expect the training and test sets to have come from different distributions. Indeed, the average years of schooling differs by more than 1 year, and is statistically significant.
In splitting the samples based on the distribution of the endogenous variable, we are also motivated by the long-standing debates revolving around the use of instrumental variables in classic economic studies card1994minimum. A leading concern is the validity of instruments. In the case of the study on educational returns, the validity of estimation and inference require that the instruments (proximity to college and quarter of birth) are not correlated with unobserved characteristics that may also affect their earnings. The following quote from card1999causal illustrates this concern:
When this assumption is violated, the estimates based on a particular subpopulation becomes unreliable for the wider population, and we evaluate the performance based on how well they generalize to other groups of the population with potential distributional or covariate shifts. In Table (ref), we compare the test set MSE of OLS, IV, Wasserstein DRIVE, anchor regression, ridge, and ridge regularized IV estimators. We see that Wasserstein DRIVE consistently outperforms other estimators commonly used in practice.
In this paper, we propose a distributionally robust instrumental variables estimation framework. Our approach is motivated by two main considerations in practice. The first is the concern about model mis-specification in IV estimation, most notably the validity of instruments. Second, going beyond estimating the causal effect for the endogenous variable, practitioners may also be interested in making good predictions with the help of instruments when there is heterogeneity between training and test datasets, e.g., generalizing from findings using samples from a particular population/geographical group to other groups. We argue that both challenges can be naturally unified as problems of distributional shifts, and then addressed using frameworks from distributionally robust optimization.
We provide a dual representation of our Wasserstein DRIVE framework as a regularized TSLS problem, and reveal a distinct property of the resulting estimator: it is consistent with non-vanishing penalty parameter. We further characterize the asymptotic distribution of the Wasserstein DRIVE, and establish a few special cases when it coincides with that of the standard TSLS estimator. Numerical studies suggest that Wasserstein DRIVE has superior finite sample performance in two regards. First, it has lower estimation error when instruments are potentially invalid, but performs as well as the TSLS when instruments are valid. Second, it outperforms existing methods at the task of predicting outcomes under distributional shifts between training and test data. These findings provide support for the appeal of our DRO approach to IV estimation, and suggest that Wasserstein DRIVE could be preferable in practice to standard IV methods. Finally, there are many future research directions of interest, such as further results on inference and testing, as well as connections to sensitivity analysis. Extensions to nonlinear models would also be useful in practice.
We are indebted to Han Hong, Guido Imbens, and Yinyu Ye for invaluable advice and guidance throughout this project, and to Agostino Capponi, Timothy Cogley, Rajeev Dehejia, Yanqin Fan, Alfred Galichon, Rui Gao, Wenzhi Gao, Vishal Kamat, Samir Khan, Frederic Koehler, Michal Koles{\'a}r, Simon Sokbae Lee, Greg Lewis, Elena Manresa, Konrad Menzel, Axel Peytavin, Debraj Ray, Martin Rotemberg, Vasilis Syrgkanis, Johan Ugander, and Ruoxuan Xiong for helpful discussions and suggestions. This work was supported in part by a Stanford Interdisciplinary Graduate Fellowship (SIGF).