Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
58,257 characters · 5 sections · 0 citation commands
The Ridge Path Estimator for Linear Instrumental Variables
\pagenumbering{arabic}
\thispagestyle{plain}
Author Information:
\setcounter{page}{1}
This paper presents the asymptotic distribution for a ridge regression estimator for the linear instrumental variable (IV) model. The ridge estimator requires a regularization tuning parameter and can achieve lower MSE than two-stage least squares. This estimator differs from previously studied ridge regression estimators in three important dimensions. First, a nonzero prior. The estimators are allowed to be shrunk towards a economically meaningful prior. This is particularly important when the estimates are structural parameters with subject matter meaning. Second, the regularization tuning parameter is selected empirically using the observed data. Instead of stating asymptotic rates the tuning parameter needs to satisfy we consider a empirically selected tuning parameter and report the resulting asymptotic distribution.
Third, the traditional GMM framework is used to characterize the asymptotic distribution of this ridge estimator. Both adding a regularization penalty term and splitting the observed data into a training and test samples, takes the estimator out of the traditional GMM framework. New moment conditions are presented that fit into the traditional GMM framework and include the first order conditions for the ridge estimator.
Currently, it is becoming fashionable for empirical work to use tuning parameters selected with a holdout or test sample. However, there is a limited theoretical work on the asymptotic properties of the resulting estimators.
The tuning parameters for ridge, Lasso and Bridge estimators are typically required to satisfy asymptotic rates of convergence to allow asymptotic results (see, \citeA{huang2008}, \citeA{Caner_2009}, and \citeA{doi:10.1080/07474938.2015.1092806}). This leaves uncertainty because there are typically an infinite number of values that satisfy the restrictions. In finite samples, different values for the regularization tuning parameter result in different estimates for the parameters of interest. To avoid this indeterminacy, the observed sample is used to optimally select the value of the tuning parameter.
The ridge path estimator is the “best” parameter estimate over a one-dimensional path in the parameter space between the global minimum and a prior. The global minimum is associated with low bias and high variance whereas the prior is associated with higher bias and zero variance. The trade-off between bias and variance is exploited to find the estimate with lower Mean Squared Error (MSE). The data is split into training and test samples. The linear IV objective function using the training sample determines the one-dimensional path and the estimate is the parameter value associated with the point on the path which minimizes the linear IV objective function using the test sample. The ridge path estimator is compared to traditional 2SLS for simulated models. We find that for low precision models with small samples, the new ridge estimator is always superior to the 2SLS estimator. However, if the model has high precision and the sample size is large, the ridge path estimator is competitive.
Precision problems in linear IV estimation can occur with several models. The past 20 years has shown a large growth in our understanding of the possible types of identification and asymptotic distributions that can occur with linear IV models (see \citeA{10.2307/23116599} for a summary): e.g. strong instruments, nearly-strong instruments, nearly-weak instrument and weak instruments. For this taxonomy, this paper and estimator is in the strong instruments setting. A related but different model is when the number of instruments grow with the sample size (see \citeA{10.2307/2692218}). In this paper we restrict attention to fixed number of instruments. The models considered in this paper are closest to the situation considered in \citeA{SANDERSON2016212}. However, unlike \citeA{SANDERSON2016212} we have small parameters on the instruments instead of having some of the parameters drifting to zero. In addition we focus on providing estimates for a given sample instead of testing for weak instruments. The models we study are explicitly strongly identified, however in a finite sample the precision can be low.
The ridge path estimator belongs to a family of estimators which utilize regularization. \citeA{bickel2006regularization} provides an overview of the properties of various regularization procedures in statistics. They loosely define regularization as “the class of methods needed to modify maximum likelihood to give reasonable answers in unstable situations.” These estimates tend to have significantly lower variance which usually comes at the price of higher bias, i.e. the “bias-variance trade-off”. Nonparametric density estimation, ridge penalty estimation, LASSO penalty estimation, elastic net and spectral cutoff are all examples of regularization. For a review of methods see \citeA{hastie2009unsupervised}.
Within the structural econometrics literature, regularization concepts have recently been used by a few authors, however the intersection is still largely open. Notable contributions are the set of papers by Carrasco et al. [\citeA{carrasco2000generalization}, \citeA{carrasco2007linear}, \citeA{carrasco2012regularization}, \citeA{doi:10.1080/07474938.2015.1092806}], \citeA{caner2010adaptive} and \citeA{liao2013adaptive}. The first set of papers extend the $m$ moment conditions to a continuum of moment conditions. The authors use ridge regularization to find the inverse of the optimal weighting operator (instead of optimal weighting matrix in traditional GMM). \citeA{caner2010adaptive} attach a linear penalty term like in the LASSO framework and argues that this helps by forcing parameters not significant down to zero. Finally, \citeA{liao2013adaptive} augments the $m$ moment conditions with another $k$ moment conditions where the second set of augmented moment conditions is constructed from the subset of the original $m$ moment conditions which may be misspecified. The new set of $m+k$ moment conditions and a LASSO-type penalty permit simultaneous estimation and moment selection. \citeA{doi:10.1080/07474938.2015.1092804} present a comparative analysis of different moment selection techniques via simulation studies.
The ridge path estimator extends the literature in three important dimensions. First, a meaningful prior is incorporated into the estimator. When the prior is ignored, or equivalently set to zero, the model penalizes variability about the origin. However, in structural economic models a more appropriate penalty will be variability about some economically meaningful prior values. The parameters have meaning in the economic environment implying that prior knowledge and expertise can be incorporated by shrinking towards a prior.
Second, the data are explicitly used to select the tuning parameter. This is in agreement with the advice to use the data in the model selection and/or tuning parameter selection. Following \citeA{10.1257/jep.31.2.3} and \citeA{10.1111/ectj.12097} we accept this sample split to determine the optimal model as a powerful tool to be embraced. A key feature of this new estimator is splitting the sample into a training and test samples. Lemma 1 gives the consistency and root-$n$ convergence of the empirically estimated tuning parameter.
Third, empirically selecting the tuning parameter impacts the asymptotic distribution of the parameter estimates. As stressed in \citeA{10.2307/3533623}, the final asymptotic distribution will depend on empirically selected tuning parameters. We address this directly by characterizing the joint asymptotic distribution that include both the parameters of interest and the tuning parameter. The resulting asymptotic distribution is nonstandard because the population parameter value is at the boundary of the parameter space. We show how the ridge path estimator can be represented as a GMM estimator and are able to apply results in \citeA{andrews2002generalized}. To our knowledge, this approach and result have not been previously presented.
Section (ref) presents the linear IV framework, describes the precision problem and the ridge path estimator. Section (ref) characterizes the asymptotic distribution of the ridge path estimator in the traditional GMM framework. Small sample properties are analyzed via simulations in Section (ref). Section (ref) concludes.
This section introduces the linear IV model notation. Ridge regression is presented as an approach to improve the MSE. The regularization tuning parameter is empirically determined by splitting the data into training and test samples. Conditional on the tuning parameter the ridge estimate for the training samples creates a path from the prior to the IV estimator. The IV objective function for the test sample is then evaluated along this path to empirically determine the optimal tuning parameter and parameters of interest. The asymptotic distribution of these estimates will be investigated using the GMM framework. The first order conditions that characterize the estimates do not immediately fit into the GMM framework. However, an alternative system of equations is presented which include the estimates.
Consider the linear instrumental variables model where $Y$ is $n \times 1$, $X$ is $n \times k $ and $Z$ is $n \times m $ with $m \geq k$
The IV estimator
where $P_Z$ is the projection matrix for $Z$, has the asymptotic distribution $$ \sqrt{n} \left( \hat{\beta}_{IV} - \beta_0 \right) \sim_a N \left( 0, \sigma^2_\varepsilon \left( \Gamma_0' R_{z} \Gamma_{0} \right)^{-1} \right). $$ The covariance can be consistently estimated with
where $\hat{\varepsilon} = Y - X\hat{\beta}_{IV}$. Let $S_0 = E[z_{i} x_{i}' ] = R_z \Gamma_0 $.
For a finite sample let\footnote{This term is both the second derivative of the objective function ((ref)) and the matrix being inverted in the last term of the covariance ((ref)).} $\frac{X' P_Z X}{n}$ have the spectral decomposition $ C \Lambda C'$, where $\Lambda$ is a positive definite diagonal $k \times k$ matrix, and $C$ is orthonormal, $C' C = I_k$. A precision problem occurs when some of the eigenvectors explain very little variation, as represented by the magnitude of the corresponding eigenvalues. This occurs when the objective function is relatively flat along these dimensions and the resulting covariance estimates are large because as equation ((ref)) shows, the variance of $\hat{\beta}_{IV}$ is proportional to $ \left( \frac{X' P_Z X}{n} \right)^{-1} = \left( C \Lambda C'\right)^{-1} = C \Lambda^{-1} C'.$ The flat objective function, or equivalently large estimated variances, leads to a relatively large MSE. The ridge path estimator addresses this problem by shrinking the estimated parameter toward a prior. The IV estimate still has low bias (it is consistent) and has the asymptotically minimum variance. However, accepting a little higher bias can have a dramatic reduction in the variance and thus provide a point estimate with lower MSE.
The ridge objective function augments the usual IV objective function ((ref)) with a quadratic penalty centered at a prior value, $\beta^p$, weighted by a regularization tuning parameter $\alpha$
The objective function's second derivative is $ \left(\frac{X' P_Z X}{n} + \alpha I_k \right) = C(\Lambda + \alpha I_k )C'. $ The regularization parameter injects stability since $ \left( \frac{X' P_Z X}{n} + \alpha I_k\right)^{-1} = C \left(\Lambda + \alpha I_k \right) ^{-1} C' $ has eigenvalues $ 1/(\lambda_i + \alpha) $ for $i=1, \ldots, k$ which are decreasing in $\alpha$. This results in smaller variance but higher bias.
Denote the ridge solution given $\alpha$ as
Equation ((ref)) shows how the tuning parameter, $\alpha$ creates a smooth curve in the parameter space between the low bias-high variance IV estimate, $\hat{\beta}_{IV}$, (when $\alpha = 0$) to the high bias-no variance prior, $\beta^p$, (when $\alpha \rightarrow \infty$). The ridge estimator should be evaluated using equation ((ref)) because the IV estimator is poorly defined for the situations considered in this paper.
Different values of $\alpha$ result in different values of $\beta$. The optimal value of $\alpha$ is determined empirically as follows. The data are split into training and test samples. The training sample is the first $[\tau n]$ observations, denoted, $Y_{\tau n}$, $X_{\tau n}$, and $Z_{\tau n}$, and are used to calculate a path between the IV estimate and the prior as in equation ((ref)). The estimate using the training sample, conditional on $\alpha,$ is
where $P_{Z_{\tau n}} $ is the projection matrix onto $Z_{\tau n}$ and $[\cdot]$ is the greatest integer function. The first order conditions for an internal solution are
or alternatively
The closed form solution is
As $\alpha$ goes from 0 towards infinity, this gives a path from the IV estimator, $\hat{\beta}_{IV, \tau n}$ (at $\alpha = 0$), to the prior, $\beta^p$ (the limit as $\alpha \rightarrow \infty$). Following this path, the optimal $\alpha$ is selected to minimize the IV least squares objective function ((ref)) over the remaining $(n - [\tau n])$ observations, the test sample, denoted $Y_{n(1-\tau)}$, $X_{n(1-\tau)}$ and $Z_{n(1-\tau)}$. The optimal value for the tuning parameter is defined by $ \hat{\alpha} = \operatorname*{arg\,min}_{\alpha \in [0, \infty)} Q_{n(1- \tau)}(\alpha)$ where
where $P_{Z_{n(1-\tau)}}$ is the projection matrix onto $Z_{n(1-\tau)}.$ The first order condition for an internal solution is
or alternatively
The ridge path regression estimate is $\hat{\beta}_{ \hat{\alpha}} \equiv \hat{\beta}_{IV, \tau n}( \hat{\alpha})$.
The first order conditions that characterize the ridge path estimator, equations ((ref)) and ((ref)), are $k+1$ equations in the $k+1$ parameters and have the structure of sample averages being set to zero. However, the functions being averaged do not fit into the traditional GMM framework. In equations ((ref)) and ((ref)) the terms in the curly brackets depend on the entire sample and not just the data for index $i$ and the parameters. The terms in the curly brackets will converge at ${O}_p \left( n^{-1/2} \right)$ and must be considered jointly with the asymptotic distributions of $ (\hat{\beta}_{IV, \tau n}(\hat{\alpha})', \hat{\alpha})'$.
The asymptotic distribution of the ridge path estimator can be determined with the GMM framework using the parameterization $\theta = \left[
\right]'$ where ${\rm vec}( \cdot)$ stacks the elements from a matrix into a column vector and ${ \rm vech}( \cdot)$ stacks the unique elements from a symmetric matrix into a column vector. The population parameter values are $$ \theta_0 = \left[
\right]'. $$ The ridge path estimator is part of the parameter estimates defined by the just identified system of equations
where the training and test samples are determined with the indicator function $$ {\bf 1}_{\tau n}(i) = \left\{
\right. $$
Three assumptions are sufficient to obtain asymptotic distribution for the ridge path estimator.
{\bf Assumptions 1} and {\bf 2} imply $E[h_i(\theta_0)] = 0$ and $\sqrt{n} H_n(\theta_0)$ satisfies the CLT.
First consider the tuning parameter. Even though it is empirically selected using the training and samples, its limiting value and rate of convergence are familiar.
Proofs are given in the appendix.
Lemma 1 implies that the population parameter value for the tuning parameter is zero, $ \alpha_0 = 0$, which is on the boundary of the parameter space. This results in a nonstandard asymptotic distribution which can be characterized by appealing to {\bf Theorem 1} in \citeA{andrews2002generalized}. The approach in \citeA{andrews2002generalized} requires the root-$n$ convergence of the parameters. Lemma 1, traditional 2SLS and method of moments establishes this for all the parameter in $\theta$. Equation ((ref)) puts the ridge path estimator in the form of the first part of equation (14) from \citeA{andrews2002generalized}. Because the system is just identified, the weighting matrix does not affect the estimator and is set to the identity matrix. The scaled GMM objective function can be expanded into a quadratic approximation about the centered and scaled population parameter values
The first term does not depend on $\theta$ and the last term converges to zero in probability. This suggests selecting $\hat{\theta}$ to minimize $H_n(\theta)' H_n(\theta)$ will result in the asymptotic distribution of $ \sqrt{n}( \hat{\theta} - \theta_0) $ being the same as the distribution of $\lambda \in \Lambda \equiv \left\{ \lambda \in R^{ m(m+1) + 2 km + k+1}: \lambda_{ m(m+1)/2 + km + k + 1} \geq 0 \right\} $ where $ ( {\cal Z} - \lambda)' M_0' M_0 ({\cal Z} - \lambda) $ takes its minimum, where the random variable is defined as
and $$ M_0 = E \left[ \frac{\partial H_n(\theta_0)}{\partial \theta'} \right] . $$ This indeed is the result by {\bf Theorem 1} of \citeA{andrews2002generalized}. The needed assumptions are given in \citeA{andrews2002generalized}. The estimator is defined as $$ \hat{\theta} = \operatorname*{arg\,min}_{\theta \in \Theta } \hspace{.1in} H_n(\theta)' H_n(\theta). $$
The objective function can be minimized at a value of the tuning parameter in $(0, \infty)$ or possibly at $\alpha = 0.$ The asymptotic distribution of the tuning parameter will be composed of two parts, a discrete mass at $\alpha = 0 $ and a continuous function over $(0,\infty)$. The asymptotic distribution over the other parameters can be thought of as being composed of two parts, the distribution conditional on $\alpha = 0$ and the distribution over $ \alpha > 0.$
In terms of the framework presented in \citeA{andrews2002generalized}, the random sample is used to create a random variable. This is then projected onto the parameter space, which is a cone. The projection onto the cone results in the discrete mass at $\alpha = 0$ and the continuous mass over $(0, \infty)$. As noted in \citeA{andrews2002generalized}, this type of a characterization of the asymptotic distribution can be easily programmed and simulated.
To investigate the small sample performance, linear IV models are simulated and estimated using 2SLS and the ridge path estimator. The model is given in equations (1) to (4) with $k=2$ and $m=3$. To standardize the model, set $z_i \sim \mbox{iid} N(0, I_3)$ and $\beta_{0}$ = (0, 0)'. Endogeneity is created with $$ \left[
\right] \sim iid N\left( 0, \left[
\right] \right). $$ The strength of the instrument signal is controlled by the parameter\footnote{Similar results are obtained via other specifications of $\Gamma_0$. These are included as part of supplementary material for the paper, available from the authors on request.} $\delta$ in $$ \Gamma_0 =
. $$ To judge the behavior of the estimator, three different dimensions of the model are adjusted.
We simulate a total of $48$ model specifications corresponding to $4$ sample sizes $n$, $4$ values of the precision parameter $\delta$ and $3$ values of the prior $\beta^p$. Each specification is simulated $10,000$ times and both 2SLS and ridge path estimator are estimated. We compare estimated $\beta_0$ values on bias, variance and MSE. For the ridge path estimator we use $\tau = .7$ to split the sample between training and test samples.
The regularization parameter $\alpha$ is selected in two steps -- first, we search in the log-space going from $10^{-5}$ to $10^6$; second, we perform a grid search\footnote{We consider a linear grid of $10,000$ points in the the second step.} in a linear space around the value selected in the first step. A final selected value of $\hat{\alpha} = 0$ in the second step corresponds to a “no regularization" scenario which implies the ridge path estimator ignores the prior in favor of the data and the value $\hat{\alpha} = 10^7$ corresponds to an “infinite regularization" scenario which implies the ridge path estimator ignores the data in favor of the prior.
Tables (ref) and (ref) compare the performance of the 2SLS estimator with the ridge path estimator for different precision levels and sample sizes when the prior is fixed at $\beta^p = (\frac{1}{\sqrt{2}}, \frac{1}{\sqrt{2}})'$ and $\beta^p = (\frac{3}{\sqrt{2}}, \frac{3}{\sqrt{2}})'$ respectively. Recall, our parameter of interest is $\beta_0 = (\beta_1, \beta_2)' = (0,0)'$. We compare the estimators based on a) bias, b) standard deviation of the estimates, c) MSE values of the estimates and d) sum of MSE values of $\hat{\beta_1}$ and $\hat{\beta_2}$. In both tables, the 2SLS estimator performs as expected -- both bias and standard deviation of estimates fall as sample size increases and as instrument signal strength increases. In smaller samples, the 2SLS estimators exhibit some bias, which confirms that 2SLSL estimators are consistent but not unbiased. Table (ref) presents a scenario where the prioir for the ridge path estimator is one standard deviation away from the true parameter estimate. We note that in the low precision setting of $\delta= 0.1$ the ridge path estimator has lower MSE for all sample sizes considered in the simulations. However as precision improves, we note that for larger sample sizes the 2SLS estimator has lower MSE. Table (ref) describes a scenario where the ridge path estimator does not have any particular advantage since it is biased to a prior which is $3$ standard deviations away from the true parameter value. However, even when prior values are far from true parameter values, there are a number of scenarios where the ridge path estimator outperforms the 2SLS estimator in terms of MSE. In particular, in small samples and low precision settings, the ridge path estimator leads to smaller MSE. When $\delta = 0.1$, the ridge path estimator leads to lower MSE values for all sample sizes except $n = 500$. When $\delta = 1$ and the model has high precision, the ridge path estimator has higher MSE than 2SLS. Thus as the signal strength improves and low precision issues subside, 2SLS dominates. The bias-variance trade-off is at work here. Consider the results corresponding to $n = 25$ and $\delta = 0.25$. The ridge path estimator has higher bias compared to the 2SLS estimator for both parameters, however this is compensated by considerably smaller standard deviation values leading to smaller MSE. This table also demonstrates scenarios where for a given $\delta$ value, as the sample size increases the estimator with lower MSE changes from ridge path to 2SLS. For $\delta = 0.25$, the ridge path estimator performs better for sample sizes $n\leq 50$ whereas 2SLS performs better for $n\geq 250$. Similarly, for $\delta = 0.50$, the ridge path estimator outperforms 2SLS only for the smallest sample size of $n =25$.
Figures (ref) - (ref) present scatter plots of the estimates from 2SLS and ridge path estimator with different priors for the following cases: a) low precision, small sample size; b) low precision, large sample size; c) high precision, small sample size; d) high precision, large sample size. These figures demonstrate the influence of the priors. The prior pulls the ridge path estimates away from the population parameter values. For low precision models ($\delta = 0.1$), the variance associated with 2SLS estimates is larger than the ridge path estimates, even in larger sample sizes. The ridge path estimator is biased towards the prior which is demonstrated by the estimates not being distributed symmetrically around the true value. On the other hand, for high precision models ($\delta = 1$) the variance reduction from 2SLS for the ridge path estimator is not as dramatic. In fact, while the variance reduction appears substantial for the prior value of $\beta^p = (\frac{1}{\sqrt{2}} , \frac{1}{\sqrt{2}})'$, it is unclear at least visually if there is a reduction in variance for a poorly specified prior at $\beta^p = (\frac{3}{\sqrt{2}} , \frac{3}{\sqrt{2}})'$. In larger samples with high precision (Figure (ref)) the 2SLS estimates outperform the ridge path estimators which is demonstrated by larger clouds which are slightly off-center from the true parameter values. However, ridge path estimators using different priors are still competitive and don't lead to a drastically worse performance (as a reference compare the performance of the 2SLS estimates to the ridge path estimates in Figure (ref)).
Table (ref), summarizes the distribution of the estimated regularization parameter $\hat{\alpha}$ for different precision levels, sample sizes and prior values. Recall Theorem 1 implies the asymptotic distribution will be a mixed distribution with some discrete mass at $\alpha = 0.$ Table (ref) reports the proportion of cases which correspond to “no regularization" ($\hat{\alpha} = 0$), “infinite regularization" ($\hat{\alpha} = 10^7 \approx \infty$) and “some regularization" ($\hat{\alpha} \in (0, 10^7)$). In all cases, there is a substantial mass of the distribution concentrated at $\hat{\alpha} = 0$. On the other hand we note that except in the cases where the prior is located at the true parameter value, there is no mass concentrated at $\hat{\alpha} \approx \infty$. We see some interesting variations corresponding to different prior values. In low precision settings (particularly $\delta = 0.1$), keeping sample size fixed, as the prior moves away from the true value, the proportion of cases with “no regularization" increases whereas the proportion of cases with “some regularization" falls. Similarly for high precision settings (particularly $\delta = 1$), as the sample size increases, the proportion of cases with “no regularization" increases whereas the proportion of cases with “some regularization" falls. In this table we also present results for large sample sizes of $n = 10,000$, which demonstrate that the mass at $\hat{\alpha} = 0$ approaches $50\%$ asymptotically, as predicted by Theorem 1. Distributions of $\hat{\alpha}$ for large sample sizes of $n = 10,000$ via histograms are presented in Figure (ref).
Table (ref) presents summaries of the smallest singular value of the matrix\footnote{This corresponds to the estimate of $E\left[\frac{\partial g_i(\beta)}{\partial \beta'} \right]$ where $g_i(\beta) = (y_i-x_i\beta)z_i$.} $\left( \frac{-X'Z}{n} \right)$ for different values of $\delta$ and $n$. The estimated asymptotic standard deviation is inversely related to the smallest singular value, or equivalently smaller singular values are associated with flatter objective functions at their minimum values. As the precision parameter increases from $\delta = 0.1$ to $\delta = 1$, the mean of the smallest singular value increases. As the sample size increases, the variance of the smallest singular values decreases.
This paper addresses the problem of poor precision in linear IV estimation which occurs in samples where the objective function is flat in some dimension(s) at its minimum. This results in imprecise estimates with high variances. S-sets and K-sets can be used to help address this problem, but without giving point estimates. The main contribution of this paper is a method to obtain point estimates that can provide lower MSE than traditional 2SLS estimates when this problem occurs. The regularized point estimates presented are based on strong identification but address a practical gap in the literature where a point estimate is needed and hence the weak identification framework is inappropriate.
A second contribution is the incorporation of a non-zero prior in the ridge path estimator. In the existing regularization literature within structural econometrics, the prior is typically fixed at the origin (following the machine learning literature). However, in structural econometric models, parameters have meaningful interpretations. Penalizing the discount factor and the risk aversion parameter towards zero is inappropriate and suggests the need to incorporate prior information. We show via simulations how a) the choice of prior affects the MSE and b) even poorly specified priors may outperform traditional 2SLS estimators in low precision or small sample size settings.
A third contribution is the characterization of the nonstandard asymptotic distribution for the ridge path estimator. This new approach incorporates the empirically selected tuning parameter into the asymptotic distribution.
The chief benefit of these estimators is better small sample performance. Simulations demonstrate the trade-off of sample size and accuracy of the prior in determining the estimators small sample performance. The general message from the simulations is that for low precision models, particularly with small samples, the ridge path estimator is superior to the 2SLS estimator. If the model has high precision and the sample size is large, then the 2SLS estimator is best. Fortunately, in these settings the ridge path estimator is competitive with the 2SLS estimator. If the prior is very close to, or at, the population parameter value then the ridge path estimator perform best in all simulations, including those with larger sample sizes. If the prior is away from the population parameter value, then the ridge path estimator's performance suffers; however even with a poorly defined prior the ridge path estimator may lead to lower MSE values, in low precision and small sample size settings,
Open questions for future research include characterizing the behavior of the ridge path estimator with alternative types of models, such as weak instrument, or nearly weak instruments. Another important area for future research is extending the asymptotic proof technique to other empirical model selection rules such as k-fold cross validation.
\nocite{*}