EconBase
← Back to paper

Testing the Exogeneity of Instrumental Variables and Regressors in Linear Regression Models Using Copulas

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

110,676 characters · 9 sections · 47 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing the Exogeneity of Instrumental Variables and Regressors in Linear Regression Models Using Copulas

abstractWe provide a Copula-based approach to test the exogeneity of instrumental variables in linear regression models. We show that the exogeneity of instrumental variables is equivalent to the exogeneity of their standard normal transformations with the same CDF value. Then, we establish a Wald test for the exogeneity of the instrumental variables. We demonstrate the performance of our test using simulation studies. Our simulations show that if the instruments are actually endogenous, our test rejects the exogeneity hypothesis approximately 93% of the time at the 5% significance level. Conversely, when instruments are truly exogenous, it dismisses the exogeneity assumption less than 30% of the time on average for data with 200 observations and less than 2% of the time for data with 1,000 observations. Our results demonstrate our test's effectiveness, offering significant value to applied econometricians.

Introduction

Suppose that we want to estimate the following linear regression model: $ Y_t=\beta_0+\beta^T X_t+\alpha P_t+\epsilon_t,~t=\{1,...,T\},$ where $X_t$ is the $k \times 1$ vector of exogenous regressors that are uncorrelated with the error term $\epsilon_t$, while $P_t$ is a scalar variable that is endogenous; i.e., it is potentially correlated with the error term. We are particularly interested in estimating the coefficient of the endogenous variable $\alpha$. Using instrumental variables is a popular approach to address this endogeneity issue (angrist1996identification). Suppose that we have access to $m \times 1$ vector of instrumental variables denoted by $Z_t$, where $Z_t$ does not include any variable in $X_t$. There are two fundamental conditions in the approach of instrumental variables: first, the instruments $Z_t$ should be correlated with the exogenous variable (relevance), and second, the instruments $Z_t$ should not be correlated with the error term (exogeneity or validity). Evaluation of the relevance condition is relatively straightforward, but testing the instrument exogeneity condition is not (wooldridge2010econometric). Usually, researchers justify this condition using economic-theoretical arguments. However, in many situations, these arguments could be considered subjective beliefs that could not be supported by the data. On the other hand, it has been shown in the existing literature (kiviet2020testing, dufour2003identification and bound1995problems) that a violation of this non-testable condition can not only lead to loss of efficiency, but can also result in even a higher estimation bias in comparison to the regular OLS.

To alleviate these concerns, there have been two streams of literature. The first stream does not address testing the instrument exogeneity condition. It explores the effect of violation of this condition on the estimate of interest and the resulting inference (ashley2009assessing, kraay2012instrumental and nevo2012identification). In other words, these works assess the robustness of different inferential conclusions to uncertainty in the validity of instruments. The second stream attempts to test the exogeneity of instruments with some caveats. sargan1958estimation and hansen1982large study formal testing of over-identification restrictions contingent on the validity of the initial just-identifying set of instruments. In other words, these works' methods rely on a non-testable assumption that the initial set of just-identifying instruments satisfies the exogeneity condition. As a result, as illustrated in deaton2010instruments and parente2012cautionary, these overidentification tests cannot reliably test the exogeneity for all instruments.

Taking a different direction, a more recent stream of work evaluates the instrument exogeneity condition using an indirect approach (kiviet2020testing and kripfganz2021kinkyreg). In kiviet2020testing, the authors adopt nonorthogonal moment conditions in regression by developing an instrument-free approach to estimate the parameters. The authors assume that the correlation between the endogenous variable and the error term is known. Then, they show that the instrument exogeneity condition is satisfied if the correlation value between the endogenous variable and the error term, which rejects the instrument exogeneity condition, is not in an admissible range. However, the issue is that the admissible range for the correlation coefficient of the endogenous variable and the error term is unknown from the data. Therefore, a researcher must rely on expert knowledge of the admissible degree of endogeneity. Relying on expert knowledge induces subjectivity to the approach in kiviet2020testing. In kitagawa2015test, the author tests a necessary condition for the exogeneity of the instruments based on ordering the conditional joint distribution of the outcome and the treatment depending on the instrument. However, the method in kitagawa2015test applies only to cases with binary treatments and discrete instruments, which limits its applicability.

Despite the lack of reliable methods to test instrument exogeneity in the existing literature, the issue remains critical due to the vast number of papers that leverage the instrumental variable approach. To address this issue, we propose a Copula-based method to test the exogeneity of instruments from the error term. Copulas have been used as an instrument-free approach to address endogeneity issues in the estimation of linear regression models (park2012handling and yang2022addressing). However, our work is the first to use this approach to test the exogeneity of instruments from the error term.

To model the endogeneity of $P_t$, we consider the reduced-form equation for it given by $ P_t=\delta_0+\delta^T X_t+\gamma^T Z_{t}+\eta_t$. Given that $X_t$ is exogenous, the endogeneity of $P_t$ implies the potential endogeneity of $Z_t$ and the error term $\eta_t$. Therefore, we model the endogeneity of $P_t$ by characterizing the joint distribution of the error term in the original regression equation $\epsilon_t$, $Z_t$ and $\eta_t$ by a Gaussian copula. The Gaussian copula is a multivariate normal distribution on the normal transformations of variables. The normal transformation of a variable in $\epsilon_t$, $Z_t$, and $\eta_t$ is a variable from the standard normal distribution with the same CDF value.

sloppyparWe assume that the error term has a normal distribution, which is a widely used assumption in the literature (kleibergen2003bayesian, ebbes2005solving, rossi2003bayesian, park2012handling and yang2022addressing). However, our simulation studies show that our approach is also robust to non-normal error terms. Following the normality assumption of the error term, we show that the exogeneity of any variable from the error term $\epsilon_t$ is equivalent to the exogeniety of its normal transformation from the error term. This is not a trivial result because the endogeneity issue is the lack of Pearson correlation, which is different from statistical independence. Therefore, testing for the exogeneity of a variable is tantamount to verifying whether the correlation between the standard normal transformations of that variable and the error term is zero. Hence, testing the lack of correlation between the standard normal transformations of that variable and the error term is a necessary and also sufficient condition for testing the exogeneity of instruments.

Next, using the Gaussian copula as the joint distribution of $\epsilon_t$, $Z_t$ and $\eta_t$, we decompose the error term $\epsilon_t$ into two components. The first is a linear combination of the standard normal transformations of $Z_t$ and $\eta_t$, to capture the endogeneity of $P_t$. The second component is a normal, independent error term essential for the consistent estimation of the regression model. We add the normal transformations of $Z_t$ and $\eta_t$ as added regressors to the original regression equation and estimate the model. We establish that the test for the exogeneity of the instruments is equivalent to a Wald test, where a linear transformation for the coefficients of the standard normal transformations of the instruments $Z_t$ should be equal to zero.

Next, by simplifying the reduced-form equation for $P_t$ and setting $\eta_t=P_t$, we establish an instrument-free approach to test for the exogeneity of a regressor. This is a valuable approach given that in the existing literature, the regressor endogeneity test boils down to comparing the coefficient of the regressor under OLS and TSLS or testing if the coefficient of the error term from the first stage is not statistically significant in the second stage (durbin1954errors, hausman1978specification, wu1973alternative, wooldridge2010econometric). However, both of these methods rely on having an exogenous instrument from the error term. In other words, the traditional approach to test for the endogeneity of a regressor is not helpful if the researcher does not have access to an instrument. In addition, this approach can lead to a wrong conclusion if the instrument used in the test is actually not exogenous to the error term. In particular, if the instruments used in the Hausman endogeneity test are not exogenous to the error term, this test will reject the exogeneity hypothesis even if the regressor is actually exogenous. In a recent work, caetano2015test develops an instrument-free approach to test the endogeneity of regressors in models with bunching. However, their method is based on two assumptions: 1) the outcome variable is continuous in the endogenous regressor, and 2) there is at least one unobservable confounder that is discontinuously distributed with respect to the endogenous variable. These two assumptions may not hold in various situations. Hence, the applicability of the method in caetano2015test could be limited.

Through a series of simulation studies, we evaluated our test's efficacy, yielding the following key findings:

itemize• Instrument Exogeniety Test: In cases where instruments are truly exogenous, our test dismisses the exogeneity assumption on average 5.8% of the times for $T=200$ and 7% of the times for $T=1,000$ at the 5% significance level. Conversely, if instruments are actually endogenous, our test rejects the exogeneity hypothesis on average 76.6% of the times for $T=200$ and 98.8% of the times for $T=1,000$ at the 5% significance level. In other words, the type-I error of our test is less than 10% and the type-II error is less than 30%, which can be decreased to 3% for larger data sets. Notably, our test remains robust even when the error term deviates from a normal distribution. • Regressor Exogeneity Test: Our method demonstrates superior accuracy compared to the Hausman test for endogeneity, even when an exogenous instrument is available for analysis. Importantly, while the Hausman test performance significantly deteriorates if the instrument used in the test is not exogenous, our test's effectiveness remains unaffected due to its instrument-free nature.

A shortcoming of our instrument exogeneity test is that if an instrument is normally distributed and highly correlated with the endogenous variable $P_t$, its normal transformation will be highly correlated with $P_t$ creating the multicollinearity issue. Moreover, similar to park2012handling, if the endogenous variable $P_t$ has a normal distribution, its normal transformation will be $P_t$ divided by its standard deviation, creating the multicollinearity issue in the regressor endogeneity test. In these cases, collecting a data set with a higher number of observations can mitigate the efficiency loss caused by multicollinearity.

Furthermore, we apply our test empirically in a Two-Stage Least Squares (TSLS) setting in Angrist's study on compulsory education, angristcompulsory. Our test successfully validates the endogeneity of the education variable. Additionally, it confirms the exogeneity of the instruments used in this study, which include interactions between the year of birth and the quarter of birth. We also show that after adding the demographic variables as regressors the endogeneity of the education variable disappears. This empirical application underscores the practical effectiveness of our test in real-world TSLS scenarios.

To the best of our knowledge, our work is the first one that provides a rigorous approach to test instruments and regressor exogeneity with minimal assumptions and straightforward implementation. Our method holds significant promise for practitioners who previously had to depend on untestable arguments to substantiate their methodologies and findings. We are confident that our work will greatly improve the rigor and reliability in the field of IV regression, offering substantial benefits to those in applied econometrics and related disciplines.

Testing the Exogeneity of Instrumental Variables Using Copulas

We aim to estimate the following linear regression model:

equation[equation omitted — 92 chars of source]

where $X_t$ represents a $k\times 1$ vector of exogenous regressors and $P_t$ is a scalar endogenous regressor. The term $\epsilon_t$ is the error component of the model. Here, $t$ indexes either time or cross-sectional units. We have at our disposal $m$ instrumental variables excluded from the original regression model to address the endogeneity of $P_t$. Denote by $Z_{t}$ the $(m\times 1)$ vector of instrumental variables, where $Z_{it}$ is the $i$th instrument. It is hypothesized that these instrumental variables may exhibit correlation with the error term $\epsilon_t$. Furthermore, for any variable $\kappa_t$, we denote its marginal CDF, its mean, and its standard deviation by $F_{\kappa}(\cdot)$, $\mu_{\kappa}$ and $\sigma_{\kappa}$, respectively.

Consider the reduced-form equation for the endogenous variable expressed as follows:

equation[equation omitted — 87 chars of source]

Given the correlation between $P_t$ and the error term in the main regression equation, and assuming the exogeneity of $X_t$, the potential correlation of $Z_t$ and $\eta_t$ with the error term $\epsilon_t$ becomes the focal point of consideration. We operate under the assumption that $Z_t$ and $\eta_t$ (the error term in the reduced-form equation) are uncorrelated.

According to Sklar's Theorem (sklar1959fonctions), any multivariate joint distribution can be expressed by a copula function alongside the marginal distributions of the involved variables. Based on this theorem, we model the endogeneity of $Z_t$ and $\eta_t$ by assuming that a Gaussian copula can represent the joint distribution of $Z_t$, $\eta_t$ and $\epsilon_t$.

assumptionA Gaussian copula represents the joint distribution of $Z_t$, $\eta_t$ and $\epsilon_t$.

Suppose that $H(Z_t, \eta_t,\epsilon_t)$ denotes the joint Cumulative Distribution Function (CDF) of the instrumental variables, the error term in the reduced-form equation for the endogenous variable, and the error term in the original regression equation. We can express this relationship as:

equation[equation omitted — 129 chars of source]

where $\Psi$ represents the CDF of a standard multivariate normal distribution. The variables $Z_t^* = [Z_{it}^*]_{i=1, \ldots, m}$, $\eta_t^*$, and $\epsilon_t^*$ are the standard normal transformations of $Z_t$, $\eta_t$, and $\epsilon_t$, respectively. For any variable $\kappa_t \in \{Z_t, \eta_t, \epsilon_t\}$, its standard normal transformation $\kappa_t^*$ is defined as a standard normal random variable corresponding to the same CDF value. The standard normal transformation of a variable $\kappa_t$ follows a standard normal distribution. The normal transformation $\kappa_t^*$ is defined as follows:

itemize• For continuous $\kappa_t$: We define $\kappa_t^* = \Phi^{-1}(F_{\kappa}(\kappa_t))$, where $\Phi$ is the CDF of the standard normal distribution. • For discrete $\kappa_t$: Assuming $\kappa_t$ takes on $n$ possible values $a_1 < a_2 < \ldots < a_n$ with $p(\kappa_t = a_i) = q_i$, and $F_{\kappa}(a_i) = \Sigma_{j=1}^{i} q_i$. For $\kappa_t = a_i$, we set $\kappa_t^* = \Phi^{-1}(u_i^*)$, where $u_i^*$ is a uniform random variable between $F_{\kappa}(a_i)$ and $F_{\kappa}(a_{i+1})$. We define $a_0 = -\infty$ and $a_{n+1} = \infty$. In the discrete case, $\kappa_t^*$ may vary for identical values of $\kappa_t = a_i$ since $u_i^*$ is a randomly drawn variable each time.

We represent the Pearson correlation coefficient between two variables, $\kappa_t$ and $\chi_t$, from the set $\{Z_t, \eta_t, \epsilon_t\}$ as $\rho_{\kappa\chi}$. Similarly, the Pearson correlation between their corresponding standard normal transformations, denoted by $\kappa_t^*$ and $\chi_t^*$, is expressed as $\rho_{\kappa^*\chi*}$.

We assume that the error term \(\epsilon_t\) has a normal distribution, which is a widely used assumption in the literature (kleibergen2003bayesian, ebbes2005solving, rossi2003bayesian, park2012handling and yang2022addressing). This assumption leads to the following relationships: \(F_{\epsilon}(\epsilon_t) = \Phi(\sigma_{\epsilon} \epsilon_t^*)\) and \(\epsilon = \sigma_{\epsilon} \epsilon_t^*\).

assumptionThe error term \(\epsilon_t\) has a mean-zero normal distribution \(N(0,\sigma_{\epsilon}^2)\).

In the subsequent analysis, we establish that for any variable $\kappa_t$, the condition $\rho_{\kappa \epsilon} = 0$ holds if and only if $\rho_{\kappa^* \epsilon^*} = 0$. Put differently, the exogeneity of $\kappa_t$ is synonymous with the absence of correlation between its standard normal transformation $\kappa_t^*$ and $\epsilon_t^* = \epsilon_t / \sigma_{\epsilon}$. This result will be the basis for our exogeneity test.

We emphasize the non-triviality of this result, particularly in distinguishing between exogeneity and independence conditions. The former merely necessitates the absence of correlation with the error term, as opposed to the latter's requirement of independence. Consider a continuous variable $\kappa_t \in \{Z_t, \eta_t, \epsilon_t\}$. We have $\kappa_t^*=\Phi^{-1}(F_{\kappa}(\kappa_t))$, indicating that $\kappa_t^*$ is a monotone transformation of $\kappa_t$. Consequently, if $\kappa_t$ is independent of $\epsilon_t$, it directly follows that $\kappa_t^*$ is independent of the transformed error term $\epsilon_t^*=\epsilon_t/\sigma_{\epsilon}$, leading to a correlation $\rho_{\kappa^* \epsilon^*} = 0$. However, the mere absence of correlation, $\rho_{\kappa \epsilon} = 0$, does not imply independence. Moreover, the standard normal transformation of $\kappa_t$ given by $\Phi^{-1}(F_{\kappa}(\kappa_t))$ is a non-linear transformation. Therefore, $\rho_{\kappa \epsilon} = 0$ resulting in $\rho_{\kappa^* \epsilon^*} = 0$ is not a trivial argument.

This result will be demonstrated first for variables $\kappa_t$ that are continuous, followed by an examination of discrete $\kappa_t$. We begin by presenting a proposition that articulates the relationship between $\rho_{\kappa \epsilon}$ and $\rho_{\kappa^* \epsilon^*}$ for a continuous variable $\kappa_t$.

propositionLet \(\epsilon_t\) be a continuous mean-zero normal random variable, and \(\kappa_t\) a continuous random variable with CDF, mean, and standard deviation denoted by \(F_{\kappa}(\cdot)\), \(\mu_{\kappa}\), and \(\sigma_{\kappa}\) respectively. The correlation between \(\kappa_t\) and \(\epsilon_t\), denoted by \(\rho_{\kappa\epsilon}\), is related to the correlation factor of their standard normal transformations, denoted by \(\rho_{\kappa^* \epsilon^*}\), as follows: \begin{equation} \rho_{\kappa\epsilon}=\rho_{\kappa^* \epsilon^*} \frac{\int \nu F_{\kappa}^{-1}(\Phi(\nu)) \phi(\nu) d\nu }{\sigma_{\kappa}}. \end{equation} Furthermore, the relationship is bounded as follows: \begin{equation} 0 < |\frac{\int \nu F_{\kappa}^{-1}(\Phi(\nu)) \phi(\nu) d\nu }{\sigma_{\kappa}}| \leq {\sqrt{1+(\frac{\mu_{\kappa}}{\sigma_{\kappa}})^2}} \end{equation}
proofWe have \begin{equation} E(\kappa_t \epsilon_t)=\rho_{\kappa\epsilon}\sigma_{\kappa}\sigma_{\epsilon}+E(\kappa_t) E(\epsilon_t) . \end{equation} On the other hand, we can write, \begin{align} E(\kappa_t \epsilon_t)=&\int \int \kappa_t \epsilon_t d\Psi(\Phi^{-1}(F_{\kappa}(\kappa_t)),\Phi^{-1}(F_{\epsilon}(\epsilon_t))),\\ &\int \int F_{\kappa}^{-1}(\Phi(\kappa_t^*)) F_{\epsilon}^{-1}(\Phi(\epsilon_t^*)) \phi(\kappa_t^*,\epsilon_t^*, \rho_{\kappa\epsilon*}) d\kappa_t^* d\epsilon_t^*,\ \end{align} where $\phi(\kappa_t^*,\epsilon_t^*, \rho_{\kappa\epsilon*})$ is the PDf of a bivariate normal distribution with the correlation factor $\rho_{\kappa^*\epsilon*}$. Using (ref), (ref), $E(\epsilon_t)=0$, and $F_{\epsilon}^{-1}(\Phi(\epsilon_t^*))=\sigma_{\epsilon}\Phi^{-1}(\Phi(\epsilon_t^*))=\sigma_{\epsilon} \epsilon^*$, we have \begin{equation} \rho_{\kappa\epsilon}\sigma_{\kappa}=\int F_{\kappa}^{-1}(\Phi(\kappa_t^*)) \Big( \int \epsilon_t^* \phi(\kappa_t^*,\epsilon_t^*, \rho_{\kappa\epsilon*}) d\epsilon_t^* \Big) d\kappa_t^*. \end{equation} From Theorem 5.10.4 in degroot2011probability, we can simplify the inside integral in (ref) as follows. \begin{equation} \int \epsilon_t^* \phi(\kappa_t^*,\epsilon_t^*, \rho_{\kappa\epsilon*}) d\epsilon_t^*=E(\epsilon_t^*|\kappa_t^*)=\kappa_t^* \rho_{\kappa^* \epsilon^*} \phi(\kappa_t^*). \end{equation} Using (ref)-(ref) and changing the variable $\kappa_t^*$ to $\nu$, we have \begin{equation} \rho_{\kappa\epsilon}=\rho_{\kappa^* \epsilon^*} \frac{\int \nu F_{\kappa}^{-1}(\Phi(\nu)) \phi(\nu) d\nu }{\sigma_{\kappa}}. \end{equation} This proves the first part of the proposition. For the second part, we can rewrite $\int \nu F_{\kappa}^{-1}(\Phi(\nu)) \phi(\nu) d\nu$ as follows by changing the variable to $\lambda=\Phi(\nu)$: \begin{equation} \int \nu F_{\kappa}^{-1}(\Phi(\nu)) \phi(\nu) d\nu=\int_{0}^1 F_{\kappa}^{-1}(\lambda) \Phi^{-1}(\lambda)d\lambda. \end{equation} By the Cauchy-Schwarz inequality, we have \begin{equation} |\int_{0}^1F_{\kappa}^{-1}(\lambda) \Phi^{-1}(\lambda)d\lambda| \leq \sqrt{\int_0^1 (F_{\kappa}^{-1}(\lambda))^2 d\lambda} \sqrt{\int_0^1 (\Phi^{-1}(\lambda))^2 d\lambda}. \end{equation} It is straightforward to show that $\int_0^1 (F_{\kappa}^{-1}(\lambda))^2 d\lambda=E(\kappa^2)=\sigma_{\kappa}^2+\mu_{\kappa}^2$, and $\int_0^1 (\Phi^{-1}(\lambda))^2 d\lambda=E_{\Phi}(\lambda^2)=1$. Hence, we can rewrite (ref) as follows. \begin{equation} |\int_{0}^1F_{\kappa}^{-1}(\lambda) \Phi^{-1}(\lambda)d\lambda| \leq \sqrt{\sigma_{\kappa}^2+\mu_{\kappa}^2}. \end{equation} Using (ref) and (ref), we have \begin{equation*} |\frac{\int \nu F_{\kappa}^{-1}(\Phi(\nu)) \phi(\nu) d\nu }{\sigma_{\kappa}}| \leq {\sqrt{1+(\frac{\mu_{\kappa}}{\sigma_{\kappa}})^2}}, \end{equation*} which is the right side of the inequality in (ref). For the left side of the inequality in (ref), given that $\nu \phi(\nu)=-1 \times (-\nu) \phi(-\nu)$ and $\Phi(-\nu)=1-\phi(\nu)$, we have the following. \begin{align} \int \nu F_{\kappa}^{-1}(\Phi(\nu)) \phi(\nu) d\nu &= \int_0^\infty \nu \phi(\nu) [F_{\kappa}^{-1}(\Phi(\nu))-F_{\kappa}^{-1}(\Phi(-\nu))] d\nu, \notag \\ &=\int_0^\infty \nu \phi(\nu) [F_{\kappa}^{-1}(\Phi(\nu))-F_{\kappa}^{-1}(1-\Phi(\nu))] d\nu. \end{align} For $\nu>0$, we have $\nu \phi(\nu)>0$, and $\Phi(\nu)>1-\Phi(\nu)$. Given that $F_{\kappa}(\cdot)$ is a strictly increasing function, its inverse is a strictly increasing function as well. Therefore, for $\nu>0$ we have $F_{\kappa}^{-1}(\Phi(\nu))-F_{\kappa}^{-1}(1-\Phi(\nu)) >0$. Consequently, we have $\nu \phi(\nu) [F_{\kappa}^{-1}(\Phi(\nu))-F_{\kappa}^{-1}(1-\Phi(\nu))] >0$. This results in $\int \nu F_{\kappa}^{-1}(\Phi(\nu)) \phi(\nu) d\nu=\int_0^\infty \nu \phi(\nu) [F_{\kappa}^{-1}(\Phi(\nu))-F_{\kappa}^{-1}(1-\Phi(\nu))] d\nu >0$, which gives the left-hand side of the inequality in (ref). This concludes the proof.

The following proposition follows from Proposition (ref).

propositionConsider the scenario where \(\epsilon_t\) is a continuous mean-zero normal random variable, and \(\kappa_t\) is a continuous random variable. In this setting, the condition \(\rho_{\kappa \epsilon} = 0\) (indicating no correlation between \(\kappa_t\) and \(\epsilon_t\)) holds if and only if \(\rho_{\kappa^* \epsilon^*} = 0\) (implying no correlation between their standard normal transformations \(\kappa_t^*\) and \(\epsilon_t^*\)).
proofProposition (ref) elucidates that for a continuous variable \(\kappa_t\), the correlation \(\rho_{\kappa \epsilon}\) can be expressed as \(\rho_{\kappa^* \epsilon^*} \times C\), where \(C\) is a non-zero bounded constant. Consequently, this implies that if \(\rho_{\kappa \epsilon} = 0\), then it necessarily follows that \(\rho_{\kappa^* \epsilon^*} = 0\), and conversely, if \(\rho_{\kappa^* \epsilon^*} = 0\), then \(\rho_{\kappa \epsilon} = 0\) as well.

Proposition (ref) establishes the same result as Proposition (ref) for discrete $\kappa_t$.

propositionConsider \(\epsilon_t\) as a continuous mean-zero normal random variable and \(\kappa_t\) as a discrete random variable, where $\kappa_t$ takes on $n$ possible values $a_1 < a_2 < \ldots < a_n$ with $p(\kappa_t = a_i) = q_i$, and $F_{\kappa}(a_i) = \Sigma_{j=1}^{i} q_i$. The correlation \(\rho_{\kappa \epsilon} = 0\) if and only if \(\rho_{\kappa^* \epsilon^*} = 0\).
proofFor \(\kappa_t = a_i\), \(\kappa_t^* = \Phi^{-1}(u_i^*)\), where \(u_i^*\) is uniformly distributed between \(F_{\kappa}(a_i)\) and \(F_{\kappa}(a_{i+1})\), with \(a_0 = -\infty\) and \(a_{n+1} = \infty\). Knowing \(\kappa_t^*\) uniquely determines \(\kappa_t\), as distinct \(\kappa_t\) values lead to different \(\kappa_t^*\) values. Thus, \(E(\epsilon_t | \kappa_t^*, \kappa_t) = E(\epsilon_t | \kappa_t^*)\). \(\kappa_t^*\) is derived based solely on \(\kappa_t = a_i\), making \(\epsilon_t\) independent of \(\kappa_t^*\) given \(\kappa_t = a_i\). This implies \(E(\epsilon_t | \kappa_t^*, \kappa_t) = E(\epsilon_t | \kappa_t)\). Hence, $E(\epsilon_t | \kappa_t)=E(\epsilon_t | \kappa_t^*)$. With \(\epsilon_t^* = \epsilon_t / \sigma_{\epsilon}\), it follows that $\sigma_{\epsilon}E(\epsilon_t^* | \kappa_t)=\sigma_{\epsilon}E(\epsilon_t^* | \kappa_t^*)$. Suppose that $\rho_{\kappa^* \epsilon^*}=0$. Given that both $\epsilon_t^*$ and $\kappa_t^*$ are normal random variables, zero correlation means that they are independent. The independence of \(\epsilon_t^*\) and \(\kappa_t^*\) implies \(E(\epsilon_t^* | \kappa_t^*) = E(\epsilon_t^*) = 0\). As \(E(\epsilon_t | \kappa_t) = \sigma_{\epsilon}E(\epsilon_t^* | \kappa_t^*)\), it follows that \(E(\epsilon_t | \kappa_t) = 0\). Therefore, \(Cov(\epsilon_t, \kappa_t) = E(\epsilon_t \kappa_t) - E(\epsilon_t)E(\kappa_t) =E(E(\epsilon_t \kappa_t|\kappa_t))= E(\kappa_t E(\epsilon_t |\kappa_t))=0\), leading to \(\rho_{\kappa \epsilon} = 0\). Conversely, suppose that $\rho_{\kappa \epsilon}=0$. Therefore, $E(\epsilon_t \kappa_t)-E(\epsilon_t )E(\kappa_t)=0$. Given $E(\epsilon_t)=0$, we have $E(\epsilon_t \kappa_t)=0$. We can write the following. \begin{itemize} • \(0 = E(\epsilon_t) = \Sigma_{i=1}^n q_i E(\epsilon_t | a_i)\). (b1) • \(0 = E(\epsilon_t \kappa_t) = \Sigma_{i=1}^n q_i a_i E(\epsilon_t | a_i)\). (b2) \end{itemize} Given that $a_i$s are distinct numbers, (b1) and (b2) show that two weighted sums of $E(\epsilon_t|a_i)$s with different weights are zero. This can happen only if $E(\epsilon_t|a_i)=0$ for all $a_i$s. In other words, $E(\epsilon_t |\kappa_t)=E(\epsilon_t)=0$. This results in $E(\epsilon_t^* | \kappa_t^*)=E(\epsilon_t^*)=0$, which leads $E(\epsilon_t^* \kappa_t^*)=E(E(\kappa_t^* \epsilon_t^* | \kappa_t^*))=E(\kappa_t^* E( \epsilon_t^* \\ | \kappa_t^*))=0$. Therefore, we have $\rho_{\epsilon^* \kappa *}=0$.

Building upon the findings of Propositions (ref), (ref), and (ref), we formulate a methodology to test the exogeneity of the instrumental variables \(Z_t\). This test hinges on examining the correlation between the standard normal transformation of \(Z_t\), denoted as \(Z_t^*\), and the normalized error term \(\epsilon^*\). Specifically, the test seeks to ascertain whether \(\rho_{Z_i^* \epsilon^*} = 0\). The details and formal structure of this exogeneity test are outlined in Theorem (ref).

theorem[Testing Exogeneity of the Instrumental Variables Using the Gaussian Copula] Denote by $\Sigma_{Z^*}$ the correlation matrix for the standard normal transformation vector of the instrumental variables $Z_t^*$. Testing the exogeneity of the instruments $Z_t$ for the endogenous variable $P_t$ in $Y_t=\beta^T X_t+\alpha P_t+\epsilon_t$ is equivalent to testing for $\Sigma_{Z^*} \theta_{Z_t^*}=0$ in the following OLS regression \begin{equation} Y_t=\beta_0+\beta^T X_t+\alpha P_t+\theta_{Z_t^*}^T Z_t^*+\theta_{\eta_t^*} \eta_t^*+\xi_t, \end{equation} where $\eta_t^*$ is the standard normal transformation of the error term from the reduced-form equation of the endogenous variable given by \begin{equation} P_t=\delta_0+\delta^T X_t+\gamma^T Z_{t}+\eta_t. \end{equation}
proofUnder the Gaussian copula framework, the joint distribution of the standard normal transformations of the instrumental variables $Z_t^*$, the error term in the reduced-form equation for the endogenous variable $\eta_t^*$, and the error term in the original regression equation $\epsilon_t^*$ is characterized by a mean-zero multivariate normal distribution. This joint distribution encapsulates the interdependencies among these variables, as captured by the Gaussian copula, and is formalized as follows: \begin{equation} \left[\begin{array}{l} Z_t^{*} \\ \eta_t^{*} \\ \epsilon_t^{*} \end{array}\right]\sim N\left([0]_{(m+2) \times 1},\Sigma_*\right). \end{equation} The covariance matrix $\Sigma_*$ is given by \begin{equation} \Sigma_*=\left[\begin{array}{ccc} \Sigma_{Z^*} & [0]_{(m \times 1)} & \rho_{Z^* \epsilon^*}^T \\ {[0]}_{(1 \times m)} & 1 & \rho_{\eta^* \epsilon^*} \\ \rho_{Z^* \epsilon^*} & \rho_{\eta^* \epsilon^*} & 1 \end{array}\right], \end{equation} where $\Sigma_{Z^*}$ is the covariance matrix of $Z_{t}^*$, $\rho_{\eta^* \epsilon^*}$ is the correlation between $\eta_{t}^*$ and $\epsilon_t^*$, and $\rho_{Z^*\epsilon^*}=[\rho_{Z_{it}^* \epsilon_t^*}]_{(i=1 \ldots m)}$, where $\rho_{Z_{it}^* \epsilon_t^*}$ is the correlation between $Z_{it}^*$ and $\epsilon_t^*$. Suppose that Cholesky decompositions of $\Sigma_*$ and $\Sigma_{Z^*}$ are given by $\Sigma_*=L L^T$ and $\Sigma_{Z^*}=L_Z L_Z^T$. We have the following. \begin{equation} L=\left[\begin{array}{ccc} L_{Z^*} & {[0]}_{(m \times 1)} & {[0]}_{(m \times 1)} \\ {[0]}_{(1 \times m)} & 1 & 0 \\ V_{(1 \times m)}& \tau & \zeta \end{array}\right], \end{equation} and \begin{equation} L^{T}=\left[\begin{array}{ccc} L_{Z^*}^T & {[0]}_{(m \times 1)} & V^{T} \\ {[0]}_{(1 \times m)} & 1 & \tau \\ {[0]}_{(1 \times m)} & 0 & \zeta \end{array}\right]. \end{equation} The relationships \(\tau = \rho_{\eta^* \epsilon^*}\) and \(\zeta = \sqrt{1 - \tau^2 - \sum_{i=1}^m V_i^2}\) can be readily established. Considering the Cholesky decomposition of \(\Sigma_*\), we obtain the following: \begin{equation} \left[\begin{array}{l} Z_t^{*} \\ \eta_t^{*} \\ \epsilon_t^{*} \end{array}\right]=L \times \left[\begin{array}{l} \Omega \\ \omega_{m+1} \\ \omega_{m+2} \end{array}\right] where \quad \Omega=\left[\omega_i\right]_{i=1\ldots m}, \end{equation} and all $\omega_i$ are standard normal random variables. From (ref), it is straightforward to see that $\omega_{m+1}=\eta_t^*$ Multiplying the third row of $L$ and the first column of $L^T$ gives us the vector $\rho_{Z^* \epsilon^*}$. Therefore, we have $V L_{Z^*}^T=\rho_{Z^* \epsilon *}$. Multiplying both sides by $(L_{Z^*}^T)^{-1}$, we have the following. \begin{equation} V=\rho_{Z^* \epsilon^*} (L_{Z^*}^T)^{-1}. \end{equation} From (ref) and (ref), we can write $Z_t^*=L_{Z^*} \Omega$. Therefore, we have the following. \begin{equation} \Omega=(L_{Z^*})^{-1} Z_t^*. \end{equation} Moreover, from (ref) and (ref), we can write \begin{equation} \epsilon_t^*=V \Omega+\rho_{\eta^* \epsilon^*} \omega_{m+1}+\zeta \omega_{m+2}. \end{equation} Substituting $V$ and $\Omega$ from equations (ref) and (ref), we can rewrite $\epsilon_t^*$ as follows \begin{align} \epsilon_t^* &=\rho_{Z^* \epsilon^*} (L_{Z^*}^T)^{-1} (L_{Z^*})^{-1} Z_t^*+\rho_{\eta^* \epsilon^*} \eta_t^*+\zeta \omega_{m+2} \notag\\ &=\rho_{Z^* \epsilon^*}(L_{Z^*}L_{Z^*}^T)^{-1}Z_t^*+\rho_{\eta^* \epsilon^*} \eta_t^*+\zeta \omega_{m+2} \\ &=\rho_{Z^* \epsilon^*}(\Sigma_{Z^*})^{-1}Z_t^*+\rho_{\eta^* \epsilon^*} \eta_t^*+\zeta \omega_{m+2}.\notag \end{align} Given that $\epsilon_t=\sigma_{\epsilon} \epsilon_t^*$, using (ref), we can rewrite the original OLS equation as follows: \begin{equation} Y_t=\beta^T X_t+\alpha P_t+\sigma_{\epsilon}\rho_{Z^* \epsilon^*}(\Sigma_{Z^*})^{-1}Z_t^*+\sigma_{\epsilon}\rho_{\eta^* \epsilon^*} \eta_t^*+\sigma_{\epsilon}\zeta \omega_{m+2}. \end{equation} Because $\omega_{m+2}$ is not correlated with any of the regressors in (ref), we can consistently estimate this OLS equation. Assuming $\xi_t=\sigma_{\epsilon}\zeta \omega_{m+2}$, $\theta_{\eta_t^*}=\sigma_{\epsilon}\rho_{\eta^* \epsilon^*}$, and $\theta_{Z_t^*}=\sigma_{\epsilon}(\Sigma_{Z^*})^{-1}\rho_{Z^* \epsilon^*}^T$, equations (ref) and (ref) are equivalent. Note that because $\Sigma_{Z^*}$ is a symmetric matrix, we have $\Sigma_{Z^*}^T=\Sigma_{Z^*}$. From Corollary (ref), we know that the exogeneity condition is equivalent to the test for $\rho_{Z^* \epsilon^*}=0$. Given that $\Sigma_{Z^*} \theta_{Z_t^*}= \sigma_{\epsilon}\rho_{Z^* \epsilon^*}$, testing for the exogeneity of the instruments is the same as testing for $\Sigma_{Z^*} \theta_{Z_t^*}=0.$ This can be done by a Wald test to verify the condition $\Sigma_{Z^*} \theta_{Z_t^*}=0.$.

Discussion

Theorem (ref) demonstrates that testing the exogeneity of the instrumental variables essentially reduces to a Wald test for \(\Sigma_{Z^*} \theta_{Z_t^*} = 0\). Furthermore, the exogeneity of the error term in the reduced-form equation for the endogenous variable \(\eta_t^*\) can be assessed by verifying \(\theta_{\eta_t^*} = 0\). Notably, \(\theta_{\eta_t^*} = 0\) implies \(\rho_{\eta^* \epsilon^*} = 0\), which subsequently leads to \(\rho_{\eta \epsilon} = 0\).

A potential scenario that may adversely affect the performance of our exogeneity test involves multicollinearity in the estimation of (ref). Specifically, if \(P_t\) is highly correlated with \(Z_t^*\), it leads to multicollinearity, which in turn causes inefficiency in regression estimates and the Wald test. Such high correlation may emerge under two conditions:

enumerate• An instrumental variable \(Z_{it}\) has a strong correlation with \(P_t\). • The instrument follows a normal distribution, resulting in \(Z_{it}^* = Z_{it} / \sigma_{Z_{it}}\).

A larger sample size may be required to mitigate the inefficiency issue in scenarios satisfying these conditions,.

Furthermore, considering that \(\Sigma_{Z^*} \theta_{Z_t^*} = \sigma_{\epsilon}\rho_{Z^* \epsilon^*}\), we can utilize the estimated coefficients ($\widehat{\theta_{Z_t^*}}$) and the Root Mean Square Error (RMSE) from the regression equation ($\widehat{\sigma_{\epsilon}}$) in (ref) to derive an estimate for \(\widehat{\rho_{Z^* \epsilon^*}} = [\widehat{\rho_{Z_i^* \epsilon^*}}]\). Subsequently for continuous instrumental variables, employing the results from Proposition (ref), we are able to calculate an estimate \(\widehat{\rho_{Z \epsilon}} = [\widehat{\rho_{Z_i \epsilon}}]\). This approach enables us to gauge the extent of endogeneity for each instrument, thereby aiding in more informed instrument selection. Instruments exhibiting higher levels of correlation can be excluded, and we also gain insight into the potential bias introduced in the Two-Stage Least Squares (TSLS) approach. This methodology not only refines instrument selection but also enhances the overall reliability of the econometric analysis.

In the next section, we delineate the application of Theorem (ref) for testing the endogeneity of regressors in the absence of instrumental variables. This approach leverages the theorem's results to test for endogeneity directly, bypassing the traditional reliance on instruments.

An Instrument-Free Approach to Test for Regressor Exogeneity Using Copulas

Suppose that we want to estimate the regression model in (ref), where \(P_t\) is a potentially endogenous variable. Traditional methods to test endogeneity, such as those proposed by durbin1954errors, hausman1978specification, wu1973alternative, and wooldridge2010econometric, require access to an exogenous instrument. In scenarios devoid of such instruments, these methods offer no reliable means to ascertain the exogeneity of \(P_t\). This section introduces an instrument-free approach, leveraging Theorem (ref), to test the exogeneity of a regressor.

Theorem (ref) provides a framework to test the exogeneity of instruments. Since \(\theta_{\eta_t^*}\) in (ref) equals \(\sigma_{\epsilon}\rho_{\eta^* \epsilon^*}\), and in light of Propositions (ref) and (ref), testing the exogeneity of \(\eta_t\) is analogous to verifying \(\theta_{\eta_t^*} = 0\). To assess the exogeneity of \(P_t\), we modify the reduced-form equation in (ref) by setting \(P_t^* = \eta_t\). Given that \(P_t^*\) follows a standard normal distribution, \(\eta_t^* = P_t^*\). This allows us to apply the results of Theorem (ref) under the assumption that there are no instruments and \(\eta_t^* = P_t^*\). The ensuing corollary formalizes this test:

corollary[A Test for Regressor Exogeneity Using Copulas] Testing the exogeneity of the regressor \(P_t\) in the regression model \(Y_t = \beta_0 + \beta^T X_t + \alpha P_t + \epsilon_t\) is equivalent to testing \(\theta_{P_t^*} = 0\) in the following OLS regression: \begin{equation} Y_t = \beta_0 + \beta^T X_t + \alpha P_t + \theta_{P_t^*} P_t^* + \xi_t, \end{equation} where \(P_t^*\) is the standard normal distribution transformation of \(P_t\).
proofPer Theorem (ref), \(\theta_{P_t^*} = \sigma_{\epsilon} ~\rho_{P^* \epsilon^*}\). Propositions (ref) and (ref) establish that \(\rho_{P \epsilon} = 0 \Leftrightarrow \rho_{P \epsilon^*} = 0\). Hence, testing for the exogeneity of \(P_t\) (i.e., \(\rho_{P \epsilon} = 0\)) is equivalent to testing \(\theta_{P_t^*} = 0\) in (ref).

Simulation Studies

In this section, we conduct simulation studies to demonstrate the efficacy of our copula-based tests in assessing the exogeneity of both instruments and regressors.

Performance of the Exogeneity Test for the Instrumental Variables

We examine a scenario where both \(X_t\) and \(P_t\) are \(1 \times 1\) vectors, with three available instrumental variables for the endogenous variable \(P_t\). The data-generating process adheres to the Gaussian copula model described in (ref). Initially, \(Z_t^*\), \(\eta_t^*\), \(\epsilon_t^*\), and \(X_t^*\) are generated using a multivariate normal distribution with a correlation matrix described below. Next, for each variable \(\kappa\) in the set \(\{Z_t, \eta_t, \epsilon_t, X_t\}\), we define \(\kappa_t = F_{\kappa}^{-1}(\kappa_t^*)\), where the CDF function $F_{\kappa}(\cdot)$ for each variable corresponds to the distributions specified below . The simulation parameters are as follows:

itemize• Except for the first instrument \(Z_{1t}\), which follows a Student-t distribution with two degrees of freedom, and the error term that may have a non-normal distribution, the rest of the variables ($X_t, P_t, \eta_t, Z_{2t}, Z_{3t}$) follow a standard normal distribution. • The following distributions are considered for the error term \(\epsilon_t\): $N(0,1)$, a student-t with $df=2$, a uniform distribution between -0.5 and 0.5, an exponential distribution with a mean of 1, and a Beta distribution with both shape and scale parameters equal to 0.5. • Two sample sizes are evaluated: \(T = 200\) and \(T = 1,000\). • Given the exogeneity of $X_t$, we assume $\rho_{X^* \epsilon^*}=0$. • Four scenarios are considered for the endogeneity of the instrumental variables: \begin{itemize} • Scenario 1: \(\rho_{Z^*\epsilon^*} = [0.0, 0.0, 0.0]\). • Scenario 2: \(\rho_{Z^*\epsilon^*} = [0.0, 0.5, 0.0]\). • Scenario 3: \(\rho_{Z^*\epsilon^*} = [0.3, 0.5, 0.0]\). • Scenario 4: \(\rho_{Z^*\epsilon^*} = [0.3, 0.5, 0.7]\). \end{itemize} • The covariance matrix \(\Sigma_{Z^*}\) is defined as: \[\Sigma_{Z^*} = \left[\begin{array}{ccc} 1 & 0.2 & 0.3 \\ 0.2 & 1 & 0.4 \\ 0.3 & 0.4 & 1 \end{array}\right].\] • A correlation of \(\rho_{\eta^* \epsilon^*} = 0.5\) is assumed. • Correlation between the exogenous variable and the instruments is modeled as \(\rho_{X^* Z^*} = [0.2, 0.2, 0.2]\). • The endogenous variable is set as \(P_t = 1 + 0.1 X_t + 0.1 Z_{1t} + 0.2 Z_{2t} + 0.3 Z_{3t} + \eta_t\). • The dependent variable is modeled as \(Y_t = 1 + 0.3 X_t + P_t + \epsilon_t\).

For each scenario, we conduct 100 simulations and apply the Wald test for the correlation factor of each instrument as stipulated in Theorem (ref). The proportion of instances where the Wald test rejects the null hypothesis (that the correlation between the instrument and error term is zero) is recorded. Additionally, the correlation between \(P_t\) and \(\epsilon_t\) is analyzed to assess the degree of endogeneity. The results are summarized in Tables (ref) and (ref) at the 5% significance level for $T=200$ and $T=1,000$, respectively. We summarize the main insights as follows:

itemize• If the instrument is not correlated with the error term, our test rejects the exogeneity hypothesis on average 5.8% of the times for $T=200$ and 7% of the times for $T=1,000$ at the 5% significance level. In other words, the type-I error of our test is less than 6% for $T=200$ and 7% for $T=1,000$. • If the instrument is correlated with the error term, our test rejects the exogeneity hypothesis on average 76.6% of the times for $T=200$ and 98.8% of the times for $T=1,000$ at the 5% significance level. In other words, the type-II error of our test (not rejecting the exogeneity hypothesis when the instruments are correlated with the error term) is less than 30% for $T=200$ and less than 2% for $T=1,000$. • Even though our model is developed based on the normality assumption for the error term, our test provides robust performance for non-normal error terms.

We present the simulation results for the 1% significance level in Appendix (ref). The insights are similar to the case with the 5% significance level.

table[table omitted — 9,950 chars of source]
table[table omitted — 9,951 chars of source]

Performance of the Exogeneity Test for the Regressors

To evaluate the performance of our test for regressor exogeneity, we generate the standard normal transformations of \(P_t\), \(X_t\), \(\epsilon_t\), and \(Z_t\) based on a specified correlation matrix described below using a multivariate normal distribution, subsequently creating the variables using their inverse CDF functions. We consider only one instrumental variable. This instrumental variable may have a potential correlation with the error term. We utilize this instrument to conduct the Hausman endogeneity test, comparing its outcomes with our instrument-free copula-based endogeneity test. The simulation setup is detailed as follows:

itemize• We assume that the endogenous variable $P_t$ has a student-t distribution with $df=2$. • The exogenous variable \(X_t\) and the instrument \(Z_t\) follow standard normal distributions. • The following distributions are considered for the error term \(\epsilon_t\): $N(0,1)$, a student-t with $df=2$, a uniform distribution between -0.5 and 0.5, an exponential distribution with mean of 1, and a beta distribution with both shape and scale parameters equal to 0.5. • Two scenarios are evaluated for the correlation between the instrument and the error term: \(\rho_{Z^* \epsilon^*} \in \{0, 0.2\}\). • Nine different levels of endogeneity for \(P_t\) are considered: \(\rho_{P^* \epsilon^*} \in \{-0.5, -0.25, -0.1, -0.05, 0\\, 0.05, 0.1, 0.25, 0.5\}\). • Correlations of \(\rho_{X^*P^*} = 0.2\) and \(\rho_{X^* Z^*} = 0.2\) are assumed to model potential correlations between the exogenous variable and the instrument. Given the exogeneity of $X_t$, we assume $\rho_{X^* \epsilon^*}=0$. • The dependent variable is set as \(Y_t = 1 + 0.3 X_t + P_t + \epsilon_t\).

Tables (ref) and (ref) show the regressor endogeneity results for the case with $\rho_{Z^* \epsilon^*}=0$ at the 5% significance level for $T=200$ and $T=1,000$, respectively, for various scenarios using both the instrument-free copula-based method and the Hausman approach. For brevity, we present the results for the case with $\rho_{Z^* \epsilon^*}=0$ at the 1% significance level for $T=200$ and $T=1,000$, and the case with $\rho_{Z^* \epsilon^*}=0.2$ at the 5% significance level for $T=1,000$ in Appendix (ref).

table[table omitted — 16,154 chars of source]
table[table omitted — 16,156 chars of source]

Key insights drawn from these simulation results are summarized below:

itemize• Relative to the traditional Hausman approach, the copula-based method exhibits greater accuracy. Specifically, it shows a higher frequency of correctly rejecting the exogeneity hypothesis when the regressor is not exogenous and a similar or lower frequency of incorrectly rejecting the hypothesis when the regressor is exogenous. • Both the copula-based and Hausman approaches demonstrate significantly improved performance with an increased number of observations (\(T = 200\) versus \(T = 1,000\)). • Our test shows a robust performance even for non-normal error terms. • As illustrated in Table (ref) in Appendix (ref), the effectiveness of the Hausman approach diminishes in scenarios where the instrument correlates with the error term, often erroneously rejecting the exogeneity hypothesis even for exogenous regressors. Conversely, as the copula-based approach is not reliant on instruments, its performance remains unaffected by the endogeneity of the instrument. • The type-I error of our approach (rejecting the exogeneity hypothesis when the regressor is exogenous) is less than 3% for both $T=200$ and $T=1,000$. • The type-II error of our approach (not rejecting the exogeneity when the regressor is endogenous) depends on the degree of endogeneity as follows: \begin{itemize} • For $|\rho_{Z^* \epsilon^*}=0.5|$, it is on average less than 21% for $T=200$ and less than 5% for $T=1,000$. • For $|\rho_{Z^* \epsilon^*}=0.25|$, it is on average less than 70% for $T=200$ and less than 6% for $T=1,000$. • For $0<|\rho_{Z^* \epsilon^*}| \leq 0.1$, it is on average less than 93% for $T=200$ and less than 74% for $T=1,000$. \end{itemize}

We emphasize that even though the type-II error of the copula based test is high for low levels of endogeneity, still our copula based method outperforms the Hausman endogeneity test approach.

Empirical Application: Instrumental Variables in Education and Earnings

This section evaluates the performance of our exogeneity tests in the context of Angrist's seminal paper on instrumental variables (angristcompulsory). The paper examines the impact of education duration on earnings, addressing the endogeneity issue arising from the unobserved ability of a student influencing both education duration and wage. The authors employ a two-stage-least-square (TSLS) method to assess the effect of education (\(E_i\)) on the log of wage (\(\ln W_i\)), using quarter of birth dummies interacted with year of birth dummies as instruments.

We apply our exogeneity tests to the endogenous variable (education length, \(EDUC\)) and the 30 instrumental variables (\(QTR120-QTR129\), \(QTR220-QTR229\), and \(QTR320-QTR329\)) for men born between 1920 and 1929. Our analysis mirrors the TSLS results presented in Table IV of angristcompulsory, specifically for Cases 2, 4, 6, and 8, described as follows:

itemize• Case 2: \(X_i\) includes year of birth. • Case 4: \(X_i\) includes year of birth, age, and age-squared. • Case 6: \(X_i\) includes year of birth, race dummies, a dummy for residence in SMSA, a marital status dummy, and eight region-of-residence dummies. • Case 8: \(X_i\) includes year of birth, age and age-squared, race dummies, a dummy for residence in SMSA, a marital status dummy, and eight region-of-residence dummies.

Considering the discrete nature of \(EDUC\) and the instrumental variables, and the randomness in the normal distribution transformation for discrete variables, we execute our test 100 times with different random draws for the normal distribution transformations. We then calculate the frequency, out of 100 trials, at which the exogeneity hypothesis is rejected at the 5% significance level. The results are shown in Table (ref). We provide the results for the 1% significance level in Appendix (ref).

table[table omitted — 3,782 chars of source]

As evidenced in Table (ref), for cases 2 and 4, the exogeneity hypothesis of the endogenous variable \(EDUC\) is rejected in over 77% of the simulations, strongly indicating the endogeneity of the education length variable. However, this apparent endogeneity measure diminishes to 6% and 2% in cases 6 and 8, where demographic variables are included, suggesting that these demographic factors account for the portions of the error term correlated with \(EDUC\).

Furthermore, Table (ref) indicates that for all instrumental variables across all cases, the exogeneity hypothesis is rejected on average 5% of the time. This consistently low rejection rate provides robust evidence supporting the exogeneity of the instruments.

Conclusion

In conclusion, our study addresses a pivotal challenge in causal inference using instrumental variables – the testing of the exogeneity condition, which has long been a contentious and unresolved issue in applied econometrics. We propose a novel Copula-based method, a significant departure from traditional practices that largely relied on economic-theoretical justifications or untestable assumptions. Our approach, grounded in modeling the joint distribution of error terms and variables using a Gaussian copula, marks a substantial advancement in verifying the exogeneity of both instruments and regressors.

Our findings, derived from extensive simulation studies, demonstrate the robustness and accuracy of our test. Notably, the instrument exogeneity test shows high efficacy, correctly rejecting exogenous or endogenous instruments in most cases, even under non-normal error term distributions. Furthermore, our regressor exogeneity test outperforms the Hausman test, providing more reliable results without the need for exogenous instruments.

The empirical application of our test in a TSLS setting, using Angrist's study on compulsory education, further validates its practical utility. We successfully confirmed the endogeneity of the education variable sing an instrument-free approach. We also confirmed and the exogeneity of the birth year and quarter interaction instruments, exemplifying the test's applicability in real-world scenarios.

Our work makes a significant contribution to the field of applied econometrics by introducing a reliable and easy-to-implement approach for testing exogeneity. This advancement enhances the rigor and reliability of econometric analyses and offers substantial benefits to researchers and practitioners in the field. It opens new avenues for conducting more accurate and credible empirical research, paving the way for informed econometric analysis across various applied contexts.