Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
55,610 characters · 17 sections · 12 citation commands
Canonical correlation regression with noisy data
\doparttoc \faketableofcontents
\def\spacingset#1{ {#1}} \spacingset{1}
\makeatletter \apptocmd{ {2pt} {2pt} {2pt} {2pt} } \apptocmd{ {2pt} {2pt} {2pt} {2pt} } \apptocmd{ {2pt} {2pt} {2pt} {2pt} } \makeatother
{\it Keywords:} Factor model, instrumental variable, measurement error.
\spacingset{1.9}
Linear instrumental variable regression is among the most widely used methods in economics, where it is increasingly implemented in data rich environments: settings where the economist has access to many noisy covariates and many noisy instruments. In macroeconomics and finance, covariates and instruments are often constructed from many noisy measurements of a few underlying factors. In applied microeconomics, covariates and instruments constructed from texts, surveys, and transactions are high dimensional and contaminated by measurement error.
In data rich environments, a common practice is to use some version of principal component analysis to compress the information in the covariates and instruments before running two stage least squares regression. However, there is little practical guidance on how to choose among the competing spectral procedures. Moreover, there is little theoretical guidance on whether the recovered low-dimensional signal is strong enough---and, more subtly, well-aligned enough---to recover the parameter of interest despite the endogeneity and noise. Our research question is how to optimally estimate the linear instrumental variable regression parameter in noisy, data rich environments where the signal of the covariates differs from the signal of the instruments.
Our key assumption is that the covariates and instruments are each repetitive in nature: they each are low rank, with a few factors providing the signal component to many noisy measurements. Crucially, we analyze not only regimes in which the covariate factors and instrument factors are well aligned, but also regimes in which they are poorly aligned. We interpret the extreme case in which the factors are perfectly aligned as a repeated measurement model. We interpret the other extreme case, in which the factors are perfectly misaligned, as an instrumental variable model with irrelevant instruments. We interpret the case of poorly aligned factors as an instrumental variable model with weak instruments. This paper provides a unified analysis across this continuum of regimes for a family of regularized estimators.
Our first contribution is to derive nonasymptotic upper bounds on the estimation error for a family of estimators that we refer to as canonical correlation regression. Different estimators in the family have different “first stages” and the same “second stage.” The first stage may be principal component analysis applied separately to the covariates and instruments; canonical correlation analysis applied jointly to the covariates and instruments; or any interpolation thereof. The second stage is regression upon the learned signal from the first stage. By providing a unified analysis for this family of estimators in the presence of noise, we characterize which version an economist should use, depending upon the regime. Our upper bound analysis culminates in a phase diagram, summarizing when to use which estimator.
Our second contribution is to derive a nonasymptotic lower bound on the estimation error in the presence of endogeneity and noise. Like the upper bound, the lower bound is valid across a continuum of regimes. By combining our upper bound and lower bound, we conclude that canonical correlation regression is sometimes optimal in noisy, data rich environments. A particular variation of the estimator is optimal in particular regimes. Intuitively, those regimes are ones in which the instrument is strong enough.
We contribute to a mature literature on ordinary least squares and two stage least squares with noisy factor models, by considering the realistic scenario in which the covariate factors and instrument factors are not perfectly aligned. The regression model is an instrumental variable regression model in which the covariates and instruments coincide. In such case, the estimator simplifies to principal component analysis of the covariates, followed by least squares. This estimator has been extensively studied in the context of noisy data under various names: factor augmented regression, diffusion indices, or principal component regression StockWatson2002,BaiNg2006,agarwalpcr,agarwal2024causalinferencecorrupteddata. To the best of our knowledge, previous analyses of two stage least squares with noisy data all place a strong assumption that we seek to relax: that the factors of the covariates and instruments are the same BaiNg2010,BaiWang2016, and therefore that the instruments are essentially repeated measurements of the covariates. By contrast, in this work, we allow the factors of the covariates and the factors of the instruments to differ. We characterize the entire continuum from perfect alignment to perfect misalignment, which we view as the realistic continuum of settings faced by empirical economists.
We contribute to the mature literatures on high dimensional and nonparametric instrumental variable regression, with possibly weak instruments, by considering the realistic scenario in which the covariates and instruments are noisy. The literature on many weak instruments is vast and growing Andrews2016CLC,Andrews2018TwoStep,MikushevaSun2022,LimWangZhangValidAR,LimWangZhangCLCWeak, and some papers explicitly consider regularization CarrascoTchuente2016,DingGuoShiWang2025, though all of these analyses appear to require clean data. For sparse and clean data, other works consider regularized approaches, under first stage restricted eigenvalue conditions belloni_chernozhukov_hansen,CARRASCO2012383,gautier2021highdimensionalinstrumentalvariablesregression. Finally, for nonparametric relationships in clean data, several works consider spectral regularization carrasco2007linear,darolles2011nonparametric,chen2018optimal,RenSunDaiSpectralIV,MeunierEtAlDemystify,van2025nonparametric. The linear versions of these estimators often belong to the family of estimators we call canonical correlation regression. Our analysis is complementary to these various works because we consider measurement error.
Finally, we contribute to the literature on canonical correlation analysis by emphasizing its central role in economic data analysis. Canonical correlation analysis is a classical tool hotelling1936relations; see e.g., bykhovskaya2025canonicalcorrelationanalysisreview,bykhovskaya2025highdimensionalcanonicalcorrelationanalysis for recent reviews. We argue that it is often implicit in the analysis of instrumental variables, and is, in fact, a crucially important lens through which to view many estimators and many regimes. The asymptotic literature on canonical correlation analysis BenaychGeorgesNadakuditi2012,OnatskiMoreiraHallin2013,OnatskiMoreiraHallin2014,Dobriban2017, BaoHuPanZhou2019 clarifies the meaning of our finite-sample bounds. In particular, it helps us to interpret the weak-alignment settings in which sample canonical correlations are unreliable for instrumental variable estimation.
Section (ref) introduces the continuum of models. Section (ref) describes the family of estimators. Section (ref) derives the upper bounds on estimation error, and summarizes them with phase diagrams. Section (ref) derives the minimax lower bound. Section (ref) illustrates our theory with simulations. Section (ref) concludes.
We formalize a data rich environment with many noisy covariates and instruments. We assume factor structure but allow a continuum of models in which the factors of covariates and instruments may be well aligned or poorly aligned.
As running convention, for a matrix $A$, we denote its transpose by $A^{\top}$ and its pseudo-inverse by $A^{\dagger}$. We denote its operator norm by $\|A\|_2$, and its $r$th singular value by $\sigma_r(A)$. Finally, we denote the projection onto the column space of $A$ by $\mathrm{proj}_A$. For a vector $v$, we denote its $\ell_2$ norm by $\|v\|_2$. For scalars $a$ and $b$, we write $a\lesssim b$ when $a\leq C b$ for some universal constant $C$. Expectations are with respect to the randomness of the structural disturbance $\varepsilon^*$ introduced below: $\mathbb{E}(\cdot \mid \cdot)=\int (\cdot )\mathrm{d}\mathbb{P}(\varepsilon^* \mid \cdot)$.
Consider the linear instrumental variable regression model with a scalar outcome, high dimensional covariates, and high dimensional instruments. In matrix notation, we express these quantities as $Y\in\mathbb{R}^n$, $X\in \mathbb{R}^{n\times p}$, and $W\in\mathbb{R}^{n\times q}$, respectively. We posit that the outcome $Y$ is generated by the linear model $$ Y =X\beta_0+ \varepsilon^*, \quad \mathbb{E}(\varepsilon^* \mid W) = 0, $$ where $\varepsilon^*\in\mathbb{R}^n$ is a structural disturbance. In this causal model, $\mathbb E(\varepsilon^* \mid X)\neq 0$ yet $\mathbb{E}(\varepsilon^* \mid W) = 0$; the covariates $X$ are endogenous, so we leverage the instruments $W$. The causal parameter of interest is the coefficient vector $\beta^*\in\mathbb R^{p}$, which we aim to estimate well. We place a weak regularity condition on this coefficient.
The crux of our problem is that the economist does not directly observe $(Y,X,W)$; instead, the economist observes $(Y,Z_X,Z_W)$, where $Z_X\in\mathbb{R}^{n\times p}$ are noisy measurements of the covariates, and $Z_W\in \mathbb{R}^{n\times q}$ are noisy measurements of the instruments. We define the perturbations as $H_X=Z_X-X$ and $H_W=Z_W-W$. Under these definitions, $$ Z_X=X+H_X,\quad Z_W=W+H_W, $$ and we have not yet placed any independence assumptions on the noise. These noise terms can capture measurement errors, discretizations, or privacy mechanisms applied to the true covariates and instruments.
What independence conditions are required on the structural disturbance $\varepsilon^*$ and the noise terms $H_X$ and $H_W$? For the structural disturbance, we have already assumed validity of the instrument via $\mathbb{E}(\varepsilon^* \mid W) = 0$. Throughout the paper, we also maintain that the structural disturbance is independent across observations and has bounded variance.
For the noises, we remain largely agnostic: the noise terms on different covariates, on different instruments, and on covariates and instruments, may be weakly correlated and heteroscedastic. Our general, non-asymptotic results only use the independence and randomness of the structural disturbance; in this sense, they are conditional upon the signal and noise $(X,W,H_X,H_W)$.
An economist may place additional structure on the problem to make our general results concrete. For example, the operator norms of the noise $H_X$ and $H_W$ appear in our upper bounds. If the economist assumes the noise is sub-Gaussian and independent across observations, then these operator norms are small with high probability, appealing to the randomness of the noise. When deriving the lower bounds, we consider Gaussian noise for simplicity.
Our strongest assumption is that the covariates and instruments each have factor structure. Algebraically, we assume they are low rank. Geometrically, we assume that the covariates and the instruments each have a low dimensional approximating subspace, though they are high dimensional. Moreover, the dimension of the covariate subspace is less than or equal to the dimension of the instrument subspace, generalizing the standard “first stage” rank condition in instrumental variable analysis.
Since the covariates and instruments are low rank, i.e., well approximated by low dimensional subspaces, we now introduce notation to describe those subspaces. Denote the singular value decompositions of the covariates and instruments by \[ X = U_* \Sigma_* V_*^\top,\quad W=\widetilde{U}_* \widetilde{\Sigma}_* \widetilde{V}_*^\top. \] The diagonal matrices $\Sigma_*\in \mathbb{R}^{k\times k}$ and $\widetilde{\Sigma}_* \in \mathbb{R}^{\ell \times \ell}$ contain the singular values. They are pre and post multiplied by matrices that contain left and right singular vectors, respectively. The spans of these singular vectors define the low dimensional approximating subspaces.
To focus on the subspaces, it is sometimes helpful to discard the scaling by the singular values. For any rank-$r$ matrix $A=U_A\Sigma_A V_A^\top$, define its “whitened” factorization $\underline A = U_A V_A^\top$. In particular, $$\underline X = U_* V_*^\top,\quad \underline W=\widetilde{U}_* \widetilde V_*^\top.$$
In our results, the continuum of models is indexed by the alignment between the covariate factors and the instrument factors. Geometrically, this is the alignment between the covariate subspace and the instrument subspace. Algebraically, it is captured by an object we call the subspace overlap matrix \[ \widetilde{U}_*^\top U_* \in \mathbb R^{\ell\times k}. \] The singular values of $\widetilde{U}_*^\top U_*$ lie in $[0,1]$ and equal the cosines of the principal angles between $\mathrm{span}(U_*)$ and $\mathrm{span}(\widetilde{U}_*)$. Equivalently, they are the canonical correlations between $\underline X$ and $\underline W$.
The rank of $\widetilde{U}_*^\top U_*$ is the number of covariate factors that can be learned from the instrument factors. By assuming $\widetilde{U}_*^\top U_*$ is full rank, we ensure that all of the directions of the covariate factors can be detected by the directions of the instrument factors. If $\widetilde{U}_*^\top U_*$ were rank deficient, then some covariate factors would be lost; the instruments would be irrelevant to those covariate factors. Importantly, we allow the smallest singular value of $\widetilde{U}_*^\top U_*$ to be close to zero, meaning that some covariate factors may only have weak instruments. In this way, we consider a continuum of instrumental variable models.
So far, we have placed assumptions on the signal: the covariates and instruments have factor structure, and the covariate factors can be at least weakly detected by the instrument factors. Now, we place an assumption on the noise: it is dominated by the signal in a particular sense.
Assumption (ref) essentially means that the signal is concentrated while the noise is diffuse. By a concentrated signal, we mean that the smallest nonzero singular values of $X$ and $W$ are large enough, similar to strong factor, pervasiveness, and incoherence assumptions bai2016econometric. By diffuse noise, we mean that the largest singular values of $H_X$ and $H_W$ are small enough, which is the case when their tails are well behaved. For example, Assumption (ref) is satisfied when $X$ and $W$ have balanced nonzero singular values while $H_X$ and $H_W$ are sub-Gaussian, weakly dependent across columns, and independent across rows.
We describe the family of two stage least squares estimators with spectral regularization. Then we articulate our research question in more detail.
A consequence of Assumption (ref) is that spectral regularization will be able to separate signal from noise. In this paper, we study a family of spectrally regularized two stage least squares estimators, which differ in how they regularize in the first stage. We will prove that different forms of first stage regularization lead to different performance, depending on the alignment between covariate factors and instrument factors as well as the strength of those factors.
To state the family of estimators, we extend the singular value decomposition notation from the true covariates and instruments to their noisy measurements. For the noisy analogues, we drop the subscript $*$. For example, we write $$ Z_X = U\Sigma V^\top, \quad Z_W = \widetilde{U} \widetilde{\Sigma} \widetilde{V}^\top. $$ Let $U_k\Sigma_k V_k^\top$ denote the rank-$k$ truncated singular value decomposition of $Z_X$, and $\widetilde{U}_\ell \widetilde{\Sigma}_\ell \widetilde{V}_\ell^\top$ the rank-$\ell$ truncated decomposition of $Z_W$. When convenient, we abbreviate $$ \widehat{X}= U_k\Sigma_k V_k^\top, \quad \widehat{W} = \widetilde{U}_\ell \widetilde{\Sigma}_\ell \widetilde{V}_\ell^\top. $$ Finally, recall the whitening notation $\underline{\widehat{X}}= U_k V_k^\top$ and $\underline{\widehat{W}} = \widetilde{U}_\ell \widetilde{V}_\ell^\top.$
The spirit of two stage least squares is to project the covariates onto the instruments in the first stage, then to project the outcomes onto these projections in the second stage. Let $\widecheck{P}\in \mathbb{R}^{n\times p}$ be a high level representation of the projected covariates from the first stage. Then the abstract formulation of the second stage is $$ \widehat{\beta} = \arg\min_{\beta}\|Y-\widecheck{P}\beta\|_2^2. $$ Compactly, the solution can be written as $ \widehat{\beta} = \widecheck{P}^{\dagger} Y$.
As a warm up, we write the oracle 2SLS estimator that has access to clean data $(Y,X,W)$. The oracle first stage would be the projection of $X$ onto $W$.
Intuitively, the first factor $\widetilde{U}_*$ is the signal of $W$. The last factor $V_*^\top$ is the signal of $X$. In the middle, we have the subspace overlap matrix $\widetilde{U}_*^\top U_*$, i.e. the true canonical correlations. It is pre-multiplied by the left weight $I$ and post multiplied by the right weight $\Sigma_*$.
In practice, the economist actually observes $(Y,Z_X,Z_W)$. Some sensible estimators spectrally regularize $Z_X$ and $Z_W$ in the first stage, either separately or jointly, before conducting least squares in the second stage.
This example is clearly an empirical analogue of the previous example. Instead of the true canonical correlations $\widetilde{U}_*^\top U_*$, it contains the estimated canonical correlations $\widetilde{U}_\ell^\top U_k$ as the middle factor. These are multiplied by the same left weight $I$ and the analogous right weight $\Sigma_k$.
Alternative first stage procedures amount to different left and right weights on the estimated canonical correlations.
The left weight remains $I$ but now the right weight becomes $I$ as well.
In this final example, the left weight becomes $\widetilde{\Sigma}_\ell^{-1}$ while the right weight becomes $I$.
The initial factor is $\widetilde{U}_\ell$, i.e. the left singular vectors of the noisy instruments, while the final factor is $V_k$, i.e. the right singular vectors of the noisy covariates. The middle factor $\widetilde{U}_\ell^\top U_k$ is the empirical canonical correlations. The matrices $A_L \in \mathbb{R}^{\ell \times \ell}$ and $A_R \in \mathbb{R}^{k \times k}$ are diagonal weighting matrices that vary across different estimators within the family. By differently weighting the canonical correlations, they trade off signal usage against conditioning. As a consequence, different choices can dominate in different signal and noise regimes.
A growing empirical literature uses some variation of principal component analysis to denoise and compress high dimensional covariates and instruments. These procedures are attractive because they are computationally simple and robust to a wide range of measurement errors, but they leave a fundamental question unanswered: how to optimally estimate the linear instrumental variable regression parameter in noisy, data rich environments. This question has several components.
First, many different estimators that combine spectral regularization and two stage least squares are possible. It is not clear how to compare them on a common footing, and how to select one as a user. Practical guidance is missing.
Second, the statistical limits of such procedures are incompletely understood. How does the magnitude of the measurement error, together with the strength and alignment of instruments, jointly determine the best achievable accuracy? Theoretical guidance is missing.
Third, results from high-dimensional canonical correlation analysis suggest that sufficiently weak correlations can be spectrally indistinguishable from noise. When is it the case that some components of $\beta^*$ may be impossible to learn?
In what follows, we develop theory for instrumental variable estimation with noisy, high dimensional data. Our analysis unifies across alignment regimes and estimators. We establish upper and lower bounds that clarify which directions of $\beta^*$ are estimable, and at what rates.
As our first contribution, we analyze the estimation error for the family of estimators we call canonical correlation regression. We consider noisy, data rich environments. Our nonasymptotic upper bound decomposes the mean square error into interpretable terms. The main bias–variance trade-off is governed by the choice of weights $(A_L,A_R)$.
Our bound yields a “phase diagram” in terms of singular values and condition numbers of $(X,W)$. In some regimes, PCA-style denoising is preferred, in others whitening is preferred, and in a wide intermediate region CCA and closely related choices of $(A_L,A_R)$ dominate.
A few interpretable quantities will recur in our upper and lower bounds: noise to signal ratios, and condition numbers. The noise to signal ratios are $$ \texttt{NSR}_X = \frac{\|H_X\|_2^2}{\sigma_k(X)^2}, \quad \texttt{NSR}_W = \frac{\|H_W\|_2^2}{\sigma_{\ell}(W)^2}. $$ Intuitively, they measure the diffusion of the noise relative to the concentration of the signal, as measured by the largest singular value of the former and the smallest nonzero singular value of the latter. More concentrated signal and more diffuse noise, reflected by smaller noise to signal ratios, will lead to better rates.
The condition numbers are $\kappa(A)=\sigma_{\max}(A)/\sigma_{\min}(A)$ and, when needed, the generalized condition number $\kappa(A,A')=\sigma_{\max}(A)/ \sigma_{\min}(A')$. We will see condition numbers for the weights $(A_L,A_R)$ as well as for the canonical correlations $(\widetilde{U}_\ell^\top U_k)$. Intuitively, the condition numbers quantify how imbalanced the spectra of these objects are. Better conditioning, reflected by smaller condition numbers, will lead to better rates.
Finally, to lighten notation, we abbreviate a key object in the definition of the canonical correlation regression: the weighted canonical correlations are $\Delta=A_L(\widetilde{U}_\ell^\top U_k)A_R \in\mathbb{R}^{\ell\times k}$, so the second stage regressors can be expressed as $\widecheck{P}= \widetilde{U}_\ell \Delta V_k^\top.$ Denote $r=\operatorname{rank}(\Delta) \leq \min\{\ell,k\}$. Different $\Delta$ are given by different choices of $(A_L,A_R)$. Clearly, $\Delta$ is an estimator of the corresponding oracle quantity in Example (ref): $\Delta_*=I( \widetilde{U}_*^\top U_*)\Sigma_*$. Considering alternative weights $(A_L,A_R)$ introduces bias, but possibly reduces variance.
As an initial step, we decompose the mean square error $\|\widehat{\beta}-\beta^*\|_2^2$ into three terms. The decomposition reflects the weighted canonical correlations $\Delta$. Specifically, we decompose the space of covariates $\mathbb{R}^p$ space via three operators: $$ \Pi_{\text {row }}=V_k \mathrm{proj}_{\Delta} V_k^{\top}, \quad \Pi_{\text {null }}=V_k\left(I_k-\mathrm{proj}_{\Delta}\right) V_k^{\top}, \quad \Pi_{\perp}=I-V_k V_k^{\top} . $$
These operators decompose the covariates based on their estimated signal. The projection onto the estimated signal is $V_k V_k^{\top}$. Clearly the initial two operators add up to the projection $V_k V_k^{\top}$, while the third operator is the rejection $I-V_k V_k^{\top}$. We call the third operator $\Pi_{\perp}$
The initial two operators further decompose the estimated covariate signal $V_k V_k^{\top}$ in a way that reflects the estimator, i.e. in a way that reflects the choice of weighted canonical correlations $\Delta$. The first operator $\Pi_{\text {row }}$ subsets to the directions of the estimated covariate signal that are preserved by the first stage of canonical correlation regression. Equivalently, it subsets to the row space of $\widecheck P$. The second operator $\Pi_{\text {null }}$ is the remainder.
Our decomposition of mean square error, formally derived in Lemma (ref), is $$ \|\hat\beta-\beta^*\|_2^2 = \|\Pi_{\text{row}}(\hat\beta-\beta^*)\|_2^2 + \left\|\left(\Pi^*_{\text{null}}- \Pi_{\text{null}}\right)\beta^*\right\|_2^2 + \left\|\left(\Pi^*_{\perp}-\Pi_{\perp}\right)\beta^*\right\|_2^2, $$ where $\Pi^*_{\text{null}} = V_*(I-\mathrm{proj}_{\Delta_*})V_*^\top$ and $\Pi^*_\perp = I-V_*V_*^\top$.
We bound the first term in Theorem (ref) below. We bound the second and third terms in Theorem (ref) below. The first term generally dominates the second and third.
These three terms have an intuitive interpretation. The first term is the error of estimating the coefficient, restricted to the directions that canonical correlation regression can actually learn, namely the row space of $\widecheck P$. The second term is error from extracting the causally relevant signal from the covariates using the estimated signal from the instruments. The third term is the error from estimating the signal of the noisy covariates.
To begin, we bound the first term in the mean square error.
We interpret the upper bound as having three components: a weighting bias, a conditioning bias, and a variance.
The weighting bias is the first term. Some objects that feature prominently are $\left\|\left(I-A_L\right)\right\|_2^2$, $\left\|A_L\right\|_2^2$, and $\left\|\Sigma_*-A_R\right\|_2^2$, which are biases from weights that deviate from $A_L=I$ and $A_R=\Sigma_*$, which appear in the oracle estimator's $\Delta_*$. This term is larger when the noise to signal ratios are worse, and when the canonical correlations are closer to zero. We may view the denominator as the effective strength of the first stage of canonical correlation regression.
The conditioning bias is the second term. Some objects that feature prominently are the condition numbers of the weights, $\kappa(A_L)$ and $\kappa(A_R)$, and of the canonical correlations $\kappa(\widetilde{U}_\ell^\top U_k)$. This term is larger when these objects are poorly conditioned, which magnifies the covariate noise.
The variance is the third term. Similar to the weighting bias, the denominator is the effective strength of the first stage of canonical correlation regression. Now, the numerator is the effective rank of canonical correlation regression: $\mathrm{rank}(\Delta)$. This term is large when the effective first stage is weak, or when the complexity of the estimator is high.
To complete our analysis of mean square error, we now bound its remaining components.
Corollary (ref) shows how the bound in Theorem (ref) is generally dominated by bound in Theorem (ref).
To conclude our discussion of the upper bound on estimation error, we provide practical guidance for economists choosing among the estimators within the family we call canonical correlation regression. We simplify the bounds in Theorems (ref) and (ref) by focusing on leading regimes. Under these simplifications, we summarize which estimators perform best via phase diagrams.
To begin, we articulate a condition under which the bound in Theorem (ref) dominates the bound in Theorem (ref). Specifically, under this condition, Corollary (ref) confirms that the weighting bias in Theorem (ref) dominates the quantities in Theorem (ref).
Having established that Theorem (ref) dominates Theorem (ref), we now establish which terms dominate within Theorem (ref). Under the following condition, the weighting bias dominates the conditioning bias. The condition essentially means that the covariates have a low enough noise to signal ratio.
Finally, to simplify the diagrams, we assume the weighting bias is never too extreme.
Under these simplifications, there are two remaining scenarios: either the weighting bias in Theorem (ref) dominates, or the variance in Theorem (ref) dominates.
These scenarios can be visualized in Figure (ref). The horizontal axis quantifies how poorly the covariate factors and the instrument factors are aligned, via the cross conditioning $\kappa(X,W)$. A larger value on this axis means less signal alignment and less signal strength. The vertical axis quantifies how poorly behaved the noise to signal ratios are, via $\texttt{NSR}=\texttt{NSR}_X+\texttt{NSR}_W$. A larger value on this axis means more noise relative to the signal. Intuitively, bias dominates if the signals are poorly aligned and weak relative to the noise.
Finally, we present our practical contribution. Within the bias-dominated region, we characterize which estimator performs best. This depends on $\sigma_{\min}(X)$ and $\kappa(W)$. Within the variance-dominated region, we characterize which estimator performs best. This depends on $\sigma_{\min}(X)$ and $\sigma_{\max}(W)$.
Having established an upper bound on the estimation error of canonical correlation regression, we now turn to a lower bound on the estimation error for any possible estimator that uses noisy covariates and instruments $(Z_X,Z_W)$. We show that canonical correlation regression is essentially rate optimal among possible estimators, in certain regimes. Specifically, in regimes with balanced spectra, the variance term of our upper bound (Theorem (ref)) matches the minimax lower bound (Theorem (ref)) up to constants.
Our analysis also sheds light on the weak alignment regime, where canonical correlations are small due to poor alignment of the covariate factors and instrument factors, and therefore estimation breaks down. We interpret this negative result in light of classic work on weak instruments and recent work on high dimensional canonical correlation analysis.
Throughout this section, we work conditionally on $(X,W)$. We view the structural disturbance and measurement error $(\varepsilon^*,H_X,H_W)$ as Gaussian with known variances $(\sigma^2_{\varepsilon},\sigma^2_X,\sigma^2_W)$. Even with this additional structure, the lower bound will match our upper bound. The lower bound is obtained by a Fano-type packing argument over $(X,W,\beta)$, where the different hypotheses are separated in the parameter $\beta$ but induce probability distributions on $(Y,Z_X,Z_W)$ that are difficult to distinguish.
The lower bound fundamentally depends on the three quantities: the largest canonical correlation via $\sigma_{\max}(\widetilde{U}_*^\top U_*)^2$, the weakest instrument factor via $\sigma_{\min}(W)^2$, and the dimension of the oracle weighting via $r_*=\operatorname{rank}(\Delta_*)$. The fraction $\frac{\sigma_{\max}(\widetilde{U}_*^\top U_*)^2}{\sigma_{\min}(W)^2}$ may be viewed as the effective strength of the first stage. Theorem (ref) shows that even with the best possible estimator, the squared error must be at least $\sigma_\varepsilon^2 r_*$ divided by this fraction. The simplified lower bound closely resembles the variance term in Theorem (ref).
Recall that the oracle estimator (Example (ref)) contained the true canonical correlations $\widetilde{U}_*^\top U_*$, and their weighted versions $\Delta_*$. We denote the number of these true canonical correlations by $r_*$. These oracle quantities appear in our mimimax lower bound (Theorem (ref)).
Meanwhile, the feasible estimators (Example (ref), (ref), (ref)) must be computed from empirical canonical correlations $\widetilde{U}_\ell^\top U_k$, weighted as $\Delta$. We denote the number of the (weighted) empirical canonical correlations by $r$. The hyperparameters $(\ell,k)$, together with positive definite weights $(A_L,A_R)$, imply $r$. These empirical quantities appear in our upper bound of canonical correlation regression (Theorems (ref) and (ref)).
Our upper bounds and lower bounds match when the empirical quantities match the oracle quantities. In practice, this amounts to “oracle tuning” of the hyperparameters $(\ell,k)$: they are correctly chosen, perhaps based on scree plots or some data driven procedure. As long as they are well chosen, and the factors are aligned well enough, the empirical canonical correlations will estimate the true canonical correlations well with high probability, implying that the upper and lower rates match.
By preserving the generality on non-oracle tuning in our upper bound analysis (Theorems (ref) and (ref)), our upper rates for canonical correlation regression continue to describe its behavior under general forms of tuning.
In this work, we view weak instruments as the setting where the covariate factors and instrument factors are weakly aligned, as measured by the canonical correlations $\widetilde{U}_*^\top U_*$. With weak enough canonical correlations, our lower bound explodes, and it becomes impossible to learn the instrumental variable regression coefficient $\beta^*$. This negative result echoes classic results on weak instruments StaigerStock1997. For the remainder of this section, we relate this negative result to the more modern asymptotic theory on high dimensional canonical correlation regression BaikBenArousPeche2005,BenaychGeorgesNadakuditi2012,BaoHuPanZhou2019.
In a high-dimensional canonical correlation setting with $p,q,n\to\infty$ at proportions, the random-matrix literature shows that there is a critical canonical correlation $\rho_c$ such that: (i) if a population canonical correlation $\rho_j\leq \rho_c$ then the associated spike is spectrally indistinguishable from noise and the sample canonical directions are asymptotically orthogonal to the truth; (ii) if $\rho_j>\rho_c$ then the sample canonical direction has a nontrivial alignment with the population one, but its variability increases as $\rho_j\downarrow\rho_c$.
Our nonasymptotic, minimax lower bound is conditional upon the realized signal; in this sense, $(X,W)$ are fixed. In proportional asymptotics where $(X,W)$ are random, existing theory describes their spectral behavior and, in particular, separates weak-alignment directions (below $\rho_c$) from strong-alignment directions (above $\rho_c$). Reading our lower bound through this lens suggests a refined understanding of the fundamental difficulty: there is an irreducible error associated with components of $\beta^*$ that lie in weakly-aligned canonical directions---which cannot be stably recovered when the corresponding effective canonical correlations do not separate from noise---and a variance inflation phenomenon for the estimable directions as the effective canonical correlations approach the critical level from above.
We corroborate our theoretical results via simulations. The simulations document impressive finite-sample behavior of the canonical correlation regression estimators defined in Section (ref).
We compare four estimators. As a benchmark, we implement 2SLS with the noisy data, without any spectral regularization. Then, we implement three variations of canonical correlation regression: PCA-2SLS (Example (ref)), Whiten-2SLS (Example (ref)), and CCA-2SLS (Example (ref)). These spectrally regularized estimators use the same truncation hyperparameters $(k,\ell)$ for comparability.
We implement several synthetic low-rank data generating processes, each with many noisy covariates and instruments, and each with correlated measurement error in covariates and instruments. Figures (ref) and (ref) consider high and moderate dimensional regimes, respectively. Within each figure, the subfigures consider regimes with very weak, weak, or strong alignment between the instrument factors and covariate factors. Within each subfigure, we consider different sample sizes and report mean square error. See Appendix (ref) for details on the data generating processes.
For each configuration, we run $250$ independent replications with fixed base seed and independent replication seeds, and report $ \mathrm{MSE}(\hat\beta)=\frac{1}{p}\|\hat\beta-\beta^*\|_2^2, $ together with pointwise $2.5\%$--$97.5\%$ quantile bands across replications.
In these data generating processes, the dominant contribution to error is the variance term, rather than bias from weighting or conditioning. This is exactly the regime in which our phase-diagram comparison is informative: performance differences are driven primarily by how each weighting shapes the spectrum of $\Delta$.
In this variance-dominated region, the phase diagram predicts that CCA-2SLS should outperform PCA-2SLS because the CCA weighting effectively stabilizes the contribution of poorly conditioned instrument directions. Equivalently, it avoids inflating the variance term by weighting towards the instruments.
The empirical curves match our theoretical prediction: CCA-2SLS yields the smallest MSE, followed by Whiten-2SLS and then PCA-2SLS. Because we simulate matched power-law spectra for $X$ and $W$ (a “balanced-spectrum” design), our minimax lower bound matches the variance term up to constants in this regime. Therefore, improvements beyond the CCA-2SLS variance level are information-theoretically limited.
This paper provides a unified analysis of low rank instrumental variable estimators with noisy, high dimensional data. Our main contributions are sharp upper and lower bounds on estimation error when covariate factors and instrument factors differ from each other. Theoretically, our results clarify how measurement error as well as alignment between the covariate factors and instrument factors govern the achievable accuracy. Practically, our results distinguish the regimes when an economist should choose an estimator based on PCA, whitening, or CCA. We conclude that canonical correlation regression is a natural and sometimes optimal method for exploiting low-rank structure in the presence of measurement error.
\spacingset{1} \spacingset{1.5}