Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
26,762 characters · 5 sections · 13 citation commands
Bias Reduction in Instrumental Variable Estimation through First-Stage Shrinkage
\begingroup \footnote{ Jann Spiess, Department of Economics, Harvard University, [email removed]}. I thank Gary Chamberlain, Maximilian Kasy, and Jim Stock for helpful comments. } \addtocounter{footnote}{-1} \endgroup
The standard two-stage least-squares (2SLS) estimator is known to be biased towards the OLS estimator when instruments are many or weak. In a linear instrumental variables model with one endogenous regressor, at least four instruments, and Normal noise, I propose an estimator that combines James--Stein shrinkage in a first stage with a second-stage control-function approach. Unlike other IV estimators based on James--Stein shrinkage, my estimator reduces bias uniformly relative to 2SLS. Unlike LIML, it is invariant with respect to the structural form and translation of the target parameter.
I consider the first stage of a two-stage least-squares estimator as a high-dimensional prediction problem, to which I apply rotation-invariant shrinkage akin to James:1992jm. Regressing the outcome on the resulting predicted values of the endogenous regressor directly would shrink the 2SLS estimator towards zero, which could increase or decrease bias depending on the true value of the target parameter. Conversely, shrinking the 2SLS estimator towards the OLS estimator can reduce risk hansen2017stein, but increases bias towards OLS. Instead, my proposed estimator uses the first-stage residuals as controls in the second-stage regression of the outcome on the endogenous regressor. If no shrinkage is applied, the 2SLS estimator is obtained as a special case, while a variant of James:1992jm shrinkage that never fully shrinks to zero uniformly reduces bias.
The proposed estimator is invariant to a group of transformations that include translation in the target parameter. While the limited-information maximum likelihood estimator (LIML) can be motivated rigorously as an invariant Bayes solution to a decision problem Chamberlain:2007uz, these transformations rotate the (appropriately re-parametrized) target parameter and invariance applies to a loss function that has a non-standard form in the original parametrization. In particular, unlike LIML, the invariance of my estimator applies to squared-error loss.
The two-stage linear model is set up in Section (ref). Section (ref) proposes the estimator and establishes bias improvement relative to 2SLS. Section (ref) develops invariance properties of the proposed estimator.
I consider estimation of the structural parameter $\beta \in \mathbb{R}$ in the standard two-stage linear regression model
from $n$ iid observations $(Y_i,X_i,Z_i,W_i)$, where $X_i \in \mathbb{R}$ is the regressor of interest (assumed univariate), $W_i \in \mathbb{R}^k$ control variables, $Z_i \in \mathbb{R}^\ell$ instrumental variables, and $(U_i,V_i)' \in \mathbb{R}^2$ is homoscedastic (wrt $Z_i$), Normal noise. $\alpha$ is an intercept, \footnote{We could alternatively include a constant regressor in $X_i$ and subsume $\alpha$ in $\beta$. I choose to treat $\alpha$ separately since I will focus on the loss in estimating $\beta$, ignoring the performance in recovering the intercept $\alpha$.} and $\gamma$ and $\pi$ are nuisance parameters. This model could be motivated by a latent variable present in both outcome and first-stage equation under appropriate exclusion restrictions as in Chamberlain:2007uz.\footnote{In this section, the intercepts $\alpha, \alpha_X$ could be subsumed in the control coefficients $\gamma,\gamma_X$ without loss, but I maintain this notation to keep it consistent.}
Throughout this document, I write upper-case letters for random variables (such as $Y_i$) and lower-case letters for fixed values (such as when I condition on $X_i = x_i$). When I suppress indices, I refer to the associated vector or matrix of observations, e.g. $Y \in \mathbb{R}^n$ is the vector of outcome variables $Y_i$ and $X \in \mathbb{R}^{n \times m}$ is the matrix with rows $X'_i$.
For the noise I use the notation
for some $\rho \in (-1,1)$. The reduced form is
with
Note that there is a one-two-one mapping between reduced-form and structural-form parameters provided that the proportionality restriction $\pi_Y = \pi \beta$ holds. I develop a natural many-means form directly from the structural model, which is thus without loss, but not without consequence. Throughout, our interest will be in estimating $\beta$ for many instruments (large $\ell$).
We have $ U_i | X_i=x_i,Z_i=z_i,W_i=w_i \sim \mathcal{N}\left(\frac{\rho \sigma}{\tau} v_i,(1-\rho^2) \sigma^2\right) $ where $v_i = x_i - \alpha_X - z_i'\pi - w_i'\gamma_X$. Given $w \in \mathbb{R}^{n \times k}$ and $z \in \mathbb{R}^{n \times \ell}$, where I assume that $(\mathbf{1},w,z)$ has full rank $1 + k + \ell \leq n - 1$, let $q = (q_{\mathbf{1}},q_w,q_z,q_r) \in R^{n \times n}$ orthonormal where $q_{\mathbf{1}} \in \mathbb{R}^n, q_w \in \mathbb{R}^{n \times k}, q_z \in \mathbb{R}^{n \times \ell}$ such that $\mathbf{1}$ is in the linear subspace of $\mathbb{R}^n$ spanned by $q_{\mathbf{1}} \in \mathbb{R}^{n}$ (that is, $q_{\mathbf{1}} \in \{\mathbf{1}/n,-\mathbf{1}/n\}$), the columns of $(\mathbf{1},w)$ are in the space spanned by the columns of $(q_{\mathbf{1}},q_w)$, and the columns of $(\mathbf{1},w,z)$ are in the space spanned by the columns of $(q_{\mathbf{1}},q_w,q_z)$. (As above, such a basis exists, for example, by an iterated singular value decomposition.) Then,
where $n^* = n - 1 - k$. Writing $X^*_z, X^*_r,Y^*_z, Y^*_r$ for the respective subvectors,
$\mu = q'_z z \pi$, and $s= n - 1 - k - \ell$, we arrive at the canonical structural form
where I have suppressed conditioning on $Z{=}z,W{=}w$ (and omit it from here on).
Given an estimator $\hat{\mu} = \hat{\mu}(X^*)$ of $\mu$, a feasible implied estimator for $\beta$ in (ref) is the coefficient on $X^*$ in a linear regression of $Y^*$ on $X^*$ and the control function $X^* - (\hat{\mu}',\mathbf{0}'_s)'$. (The two-stage least-squares estimator $\hat{\beta}^{\textnormal{2SLS}} = \frac{(Y^*_z)'X^*_z}{(X^*_z)'X^*_z}$ is obtained from the first-stage OLS solution $\hat{\mu}^{\textnormal{OLS}} = X^*_z$. It is biased towards the OLS estimator $\hat{\beta}^{\textnormal{OLS}} = \frac{(Y^*)'X^*}{(X^*)'X^*}$.)
For high-dimensional $\mu$, a natural estimator for $\mu$ is a shrinkage estimator of the form $ \hat{\mu}(X^*) = c(X^*) X^*_z $ with scalar $c(X^*)$. The conditional bias of the implied control-function estimator $\hat{\beta}$ takes a particularly simple form for this class of estimators:
Shrinkage in the James:1992jm estimator (for unknown $\tau^2$) takes the form $c(x^*) = 1 - p \frac{\|x^*_r \|^2}{\| x^*_z \|^2}$. This shrinkage pattern (and its positive-part variant) is unappealing here, as it can cross zero, around which point the estimator diverges. A natural variant that mitigates this problem is
which behaves as $1 - p \frac{\|x^*_r \|^2}{\| x^*_z \|^2}$ for small $p \frac{\|x^*_r \|^2}{\| x^*_z \|^2}$, but never quite reaches zero.
The requirement $\ell \geq 4$ is an artifact of this specific shrinkage pattern and dominance should extend to $\ell = 3$ for an appropriate modification.
The estimator $\hat{\beta}$ developed in the previous section has invariance properties in a decision problem, where in spirit and notation I follow the treatment of LIML in Chamberlain:2007uz.
First I fix the sample and action spaces, as well as a class of loss functions, for the decision problem of estimating $\beta$. Starting with (ref), I write $\mathcal{Z} = (\mathbb{R}^{\ell + s})^2$ for the sample space from which $(X^*,Y^*)$ is drawn according to $\mathop{}\!\textnormal{P}_\theta$, where I parametrize $\theta = (\beta,\mu,\rho,\sigma,\tau) \in \Theta = \mathbb{R} \times \mathbb{R}^\ell \times \mathbb{R}^3_{\geq 0}$. The action space is $\mathcal{A}= \mathbb{R}$, from which an estimate of $\beta$ is chosen. I assume that the loss function $L: \Theta \times \mathcal{A} \rightarrow \mathbb{R}$ can be written as $L(\theta,a) = \ell(a - \beta)$ for some sufficiently well-behaved $\ell: \mathbb{R} \rightarrow \mathbb{R}$ (such as squared-error loss $L(\theta,a) = (a - \theta)^2$). The estimator $\hat{\beta}: \mathcal{Z} \rightarrow \mathcal{A}$ from the previous section is a feasible decision rule in this decision problem.
For an element $g = (g_\beta,g_z,g_r)$ in the (product) group $G = \mathbb{R} \times O(\ell) \times O(s)$, where $\mathbb{R}$ denotes the group of real numbers with addition (neutral element $0$) and $O(\ell)$ the group of ortho-normal matrices in $\mathbb{R}^{\ell \times \ell}$ with matrix multiplication (neutral element $\mathbb{I}_\ell$), consider the following set of transformations (which are actions of $G$ on $\mathcal{Z}, \Theta, \mathcal{A}$):
These transformations are tied together by leaving model and loss invariant. Indeed, the following result is immediate from (ref):
A decision rule $d: \mathcal{Z} \rightarrow \mathcal{A}$ is invariant if, for all $(g,(x^*,y^*)) \in G \times \mathcal{Z}$, $d(m_\mathcal{Z}(g,(x^*,y^*))) = m_\mathcal{A}(g,d((x^*,y^*)))$. The estimator $\hat{\beta}$ above is included in a class of invariant decision rules:
An application of James--Stein shrinkage to instrumental variables in a canonical structural transformation consistently reduces bias. The specific estimator is invariant to a group of transformations of the structural form that involves translation of the target parameter.
In a companion paper JSC, I show how analogous shrinkage in at least three control variables provides consistent loss improvement over the least-squares estimator without introducing bias, provided that treatment is assigned randomly. Together, these results suggests different roles of overfitting in instrumental variable and control coefficients, respectively: while overfitting to instrumental variables in the first stage of a two-stage least-squares procedure induces bias, overfitting to control variables induces variance.