Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
50,110 characters · 11 sections · 45 citation commands
Inference on panel data models with a generalized factor structure
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
{\it Keywords:} Panel data, interactive fixed effects, two-way effects, semiparametric techniques, consistent specification test, EU Emissions Trading System.
\spacingset{1.8}
To study the effect of a vector of variables $x$ on output $y$, a very popular model in the panel data literature is
where $i$ is the cross-section index and $t$ the time series index, $y_{it}$ is a real-valued response variable, $x_{it}$ is a $d_x$-vector of regressors of different typology (i.e., individual-varying, $x_{1i}$, time-varying, $x_{2t}$, and individual-specific covariates, $x_{3it}$), and $\beta$ is a $d_x\times 1$ vector of parameters. The error term comprises two components: $m_0\left(\cdot,\cdot\right)$ is typically a known real-valued function that depends on the unobserved fixed effects, $\lambda_i \in \mathbb{R}^{d_{\lambda}}$, and a vector of unobserved common factors, $f_t \in \mathbb{R}^{d_{f}}$, whereas $u_{it}$ is the idiosyncratic error term. Both vectors, $\lambda_i$ and $f_t$, are allowed to be correlated with the observed regressors $x_{it}$, and $u_{it}$ is supposed to be uncorrelated with $x_{it}$. For instance, $m_0\left(\lambda_{i},f_{t}\right) = \lambda^{\prime}_if_t$ and $m_0\left(\lambda_{i},f_{t}\right) = \lambda_i+f_t$ give the well known interactive fixed effects and the two way fixed effects model, respectively. As in practice it is hard to know the correct specification of $m_0\left(\cdot,\cdot\right)$ one may prefer to think of it as an unknown smooth function.
In fact, model ((ref)) encompasses several panel data models that are of interest in macroeconomics, microeconomics, and finance, among other fields. In macroeconomics, the economic growth of countries is influenced by both individual-specific observable factors ($x_{3it}$) (i.e., physical capital investment, consumption, and population growth), as well as time-varying factors ($x_{2t}$) (i.e., global environmental policies or technological advancements). Furthermore, unobserved common factors ($f_t$) (i.e., global financial crises and fluctuations in world oil prices) exert an impact on all countries through trade and finance linkages with different intensities across countries ($\lambda_i$), as it is noted in CHUDIK_MOHADDES_PESARAN_RAISSI:2017. In microeconomics, the individual's wage is determined by a combination of observable characteristics ($x_{3it}$) (i.e., age and working experience) and individual-varying covariates ($x_{1i}$) (i.e., years of education and gender). Furthermore, several unobserved factors, such as innate ability ($\lambda_i$), may be correlated with the observable characteristics, and their causal effects are not constant over time ($f_t$).
For $m_0\left(\lambda_{i},f_{t}\right) = \lambda^{\prime}_if_t$, the currently dominant estimation methods for the parameters of interest in ((ref)) are the common correlated effect (CCE) estimator of PESARAN:2006 and the iterative least squares (ILS) estimator proposed in BAI:2009 and extended in MOON_WEIDNER:2015. The CCE approach proposes to approximate these unobserved factors using linear combinations of cross-sectional averages of the dependent and explanatory variables. Limitations of this method are discussed in WESTERLUND_URBAIN:2013. The ILS approach estimates the parameters of interest through a non-convex optimization problem. Recently, least squares estimation with nuclear norm penalization has been proposed to overcome this problem, but the resulting estimator exhibits a fairly slow convergence rate (see moon2025nuclearnormregularizedestimation and beyhum2019squarerootnuclearnormpenalized). See CHUDIK:2015, KARABIYIK_PALM_URBAIN:2019 and SARAFIDIS_WANSBEEK:2012 for recent surveys. For $m_0\left(\lambda_{i},f_{t}\right) = \lambda_i+f_t$, different estimators have been developed; see, for example, Chapter 3 in HSIAO:2014 and references therein.
On the one side, identification and inference of the parameters of interest in ((ref)) depend crucially on the specification of $m_0\left(\cdot,\cdot\right)$. On the other side, when $m_0\left(\cdot,\cdot\right)$ is unknown, identification and estimation of $\beta$ in ((ref)) is more cumbersome. A popular solution is to introduce a conditional mean independence restriction between $(\lambda_i,f_t)$ and $x_{it}\in \mathbb{R}^{d_x}$. This approach is in spirit related to the nonparametric control function approach (see, for instance, NEWEY_POWELL_VELLA:1999 in a different context), and even more to the extended Mundlak device for mixed-effects models (see LombardiaEtal:2012). Under conditional mean independence, model ((ref)) takes the form of a partial linear model HARDLE_LIANG_GAO:2000 with some advantages in terms of identification, inference and validation.
In this way, we can propose a consistent specification test that checks whether this assumption is credible. The test relies on ideas of ZHENG:1996, SU_JIN_ZHANG:2015, and CAI_FANG_XU:2020, among others, combining the methodology of conditional moments tests and nonparametric estimation techniques. Using degenerate and nondegenerate theories of U-statistics we can show the asymptotic distribution of our test under the null, and that it diverges under the alternative at a rate of order $\sqrt{NTh^{d/2}}$, where $h$ is a smoothing parameter and $d\equiv d_x+d_w$. If the conditional mean independence assumption is fulfilled, by using semiparametric estimation techniques for partial linear models, such as the profile least squares approach proposed in FAN_HUANG:2005, one can obtain relatively straight-forward but efficient estimators for $\beta$.
In sum, the introduction of a conditional mean independence restriction between $(\lambda_i,f_t)$ and $x_{it}$ converts model ((ref)) to a simple partially linear model for panel data. Such semiparametric specification has important advantages, mainly: i) We obtain consistent estimators of the parameters of interest, $\beta$, with optimal rates of convergence; ii) We can identify the parameters of interest even in the presence of individual-varying or time-varying covariates since the method is not based on quasi-differencing techniques; iii) Our asymptotic results are obtained under two different scenarios: $N,T \rightarrow \infty$ and for fixed $T$, $N\rightarrow \infty$; and iv) we can derive a specification test for the conditional mean independence assumption.
We employ our methods to assess the effect of the European Union Emissions Trading System (EU-ETS) on the economic development of EU countries. This topic is of great interest since this policy is the cornerstone of the EU's strategy to decarbonize the economy, but there is a certain reluctance due to its potentially adverse economic impact. We find that the postulates of the Environmental Kuznets Curve are corroborated, together with the significant effect of population and energy intensity on environmental degradation. Moreover, our results show the inability of the typically imposed functional forms for fixed effects interactions to capture the effect of unobserved common factors in a proper way, which results in counter-intuitive, misleading conclusions.
Section (ref) proposes our modeling approach with estimation procedures, and establishes the asymptotic distributions of our estimators. Section (ref) presents a test to check the conditional independence assumption, and introduces a bootstrap method to perform inference in finite samples. Section (ref) studies the finite sample performance of the proposed estimator and test statistic via several Monte Carlo experiments. Section (ref) studies the effects of the EU ETS. Section (ref) concludes. The proofs of the main theorems are shown in Appendix A. The remaining theoretical lemmas with proofs with additional simulation studies are provided in the Supplementary Material.
\setcounter{equation}{0}
Consider model ((ref)) with $m_0$ being an unknown smooth function. Following the discussion in the Introduction, for identification of $\beta$ we introduce a conditional independence assumption along a set of observed variables $w_{it} \in \mathbb{R}^{d_w}$, such that
This condition implies exogeneity of $u_{it}$ with respect to the pair $(x_{it},w_{it})$, i.e., $E\left(\left.u_{it}\right|x_{it},w_{it}\right)=0$ and that both factors and factor loadings are exogenous to the error term $u_{it}$, i.e., $E\left(\left.u_{it}\right| \mathbb{D}\right)=0$. Next, we introduce an assumption about the relationship between factors, factor loadings, and covariates
This implies that conditionally on $w_{it}$, $(\lambda_i, f_t)$ is mean independent of $x_{it}$. As said, it is a frequently used strategy for the identification of $\beta$, as in the case of control variables, control functions, or the Mundlak device. In practice, popular examples for (the elements of) $w_{it}$ are $\bar{x}_i$, $\bar{x}_t$ or their product. Depending on the context, other examples can be the inflation rate or, as in our application, the stock of debt liabilities.
Taking conditional expectations on both sides of ((ref)) with respect to $(x_{it},w_{it})$, and applying Assumptions (ref) and (ref), equation ((ref)) becomes
where $\epsilon_{it} = y_{it} - E\left(\left.y_{it}\right|x_{it},w_{it}\right) = u_{it} + \eta_{it}$ and $\eta_{it} = m_0\left(\lambda_i,f_t\right)-g\left(w_{it}\right)$.
Here, $x_{it}$ and $w_{it}$ represent different features: $x_{it}$ contains the observed regressors of interest, and $w_{it}$ observed or constructed covariates that may facilitate the control for unobserved factors.
To obtain a consistent estimator for $\beta$, we propose a profile least-squares estimation procedure following ideas of FAN_HUANG:2005 since among the known estimation procedures for partial linear models, it turned out to adapt quite well to our problem. More precisely, for a given value of $\beta$, we can propose a local linear least squares kernel estimator of the nonparametric function $g(\cdot)$ at a point $w\in \mathbb{R}^{d_w}$, $\widehat{a}$, as the solution of
where $K_{h_w}(\cdot)$ is a kernel function. For multivariate $w$ we use a product kernel $K_{h_w}(u)=\prod_{l=1}^{d_w}k_{h_w}(u_{l})$ with $u=(u_1,\ldots,u_{d_w})^{\prime}$ and $k_{h_w}(w_{it}-w)=h_w^{-1}k((w_{it}-w)/h_w)$.
For any given $\beta$, the minimizer of ((ref)) leads to the infeasible estimator for $g(\cdot)$
where $Y=(y_{11},\ldots,y_{NT})$ is a $NT\times 1$ vector, $\mathcal{W}\inI\!\!R^{(1+d_w)}$ whose $it$-th element is $[1,(w_{it}-w_0)^{\prime}]$, and $X=(x_{11},\ldots,x_{NT})$ is a $NT\times d_x$ matrix. In addition, $K_{w}=diag\{K_{h_w}(w_{11}-w),\ldots,K_{h_w}(w_{NT}-w)\}$ is a $NT\times NT$ diagonal matrix and $0_{d_w}$ a $d_w$-vector of zeros.
Writing $\widetilde{g}(w; h_w)$ in terms of the kernel projection $S\in I\!\!R^{NT \times NT} $ defined below, and plugging the resulting expression into model ((ref)), we obtain a transformed regression model that can be written in vectorial form as
where $\widetilde{Y}=(I_{NT}-S)Y$ and $\epsilon^*$ are $NT\times 1$ vectors, whereas $\widetilde{X}=(I_{NT}-S)X$ with $X$ and $\widetilde{X}$ being $NT\times d_x$ matrices. Also, it is not hard to see that $\epsilon_{it}^*=\epsilon_{it}+\widetilde{g}(w;h_w)-g(w_{it})$ and
Then, from ((ref)), the least squares estimator proposed for $\beta$ is
Finally, while $\widetilde{g}(w;h_w)$ is an infeasible estimator because it depends on the unknown parameter $\beta$, plugging ((ref)) in ((ref)) gives a feasible estimator of the form
Let us consider the two standard scenarios: (i) $N\rightarrow\infty$ and $T$ fixed; (ii) $N\rightarrow\infty$ and $T\rightarrow\infty$. For the latter case, we first recall the definition of a strongly mixing sequence.
Let $\{\zeta_t\}$ be a strictly stationary process and $\mathcal{F}_{t'}^t$ denotes a $\sigma$-algebra of events generated by the random variables $(\zeta_{t'},\ldots,\zeta_t)$ for $t'\leq t$. Following ROSENBLATT:1956, a process is said to be strongly mixing or $\alpha$-mixing if \[ \alpha(\tau)=\sup_{t'\in\mathcal{N}}\{|P(A\cap B)-P(A)P(B)|:A\in\mathcal{F}_{-\alpha}^{t'},B\in\mathcal{F}_{t'+\tau}^{-\infty}\}\rightarrow0,\quad\textrm{as} \quad T\rightarrow\infty \ . \]
For the sake of presentation, we introduce now the following notation: $B_x(w_0)=E[x_{it}|w_{it}=w_0]$, $\Phi_{\epsilon}(\chi_{it},\chi_{it'})=E(\epsilon_{it}\epsilon_{it'}|\chi_{it},\chi_{it'})$, and $\Omega_x=E[\{x_{it}-B_x(w_{it})\}\{x_{it}-B_x(w_{it})\}^{\prime}]$, for $\chi_{it}=(x_{it},w_{it})$. For the univariate kernel $k(\cdot)$, we denote $\mu_2=\int u^2k(u)du$, $\nu_0=\int k^2(u)du$, and $\nu_2=\int u^2k^2(u)du$, where $\mu_2$, $\nu_0$, and $\nu_2$ are scalars different from zero. Further, for a real matrix A, let $\|A\|=tr^{1/2}(A^{\prime}A)$. For a vector $\textbf{v}$, let $\|\textbf{v}\|$ denote its Euclidean norm. Also, let $D_g(w)$ and $\mathcal{H}_g(w)$ be the $d_w\times 1$ first-order derivative vector and the $d_w\times d_w$ Hessian matrix of $g(\cdot)$ with respect to $w$, respectively. Similarly, let $D_{\rho}(w)$ be the $d_w\times1$ first-order derivative vector of $\rho(w)$.
Then, to show $\sqrt{N}$-consistency of $\widehat \beta$ as $N$ tends to infinity and $T$ is fixed, we suppose
These assumptions are fairly standard for local linear estimation in panel data models with $T$ fix. In particular, Assumptions (ref), (ref) a), and (ref) are sufficient to show the uniform convergence of a kernel-type regression estimator (see MACK_SILVERMAN:1982).
We can also show the $\sqrt{NT}$-consistency of $\widehat \beta$ when both $N$ and $T$ tend to infinity. In order to do so, we replace Assumptions (ref), (ref), and (ref) by the following ones:
To obtain the asymptotic distribution of the nonparametric estimator ((ref)) we add
The following theorems collect the main asymptotic properties of the estimator of the parameter of interest, $\beta$. Detailed proofs of them can be found in the Appendix A.
\setcounter{equation}{0}
Above we have shown that Assumptions (ref) and (ref) are sufficient to identify the parameters of interest in model ((ref)). Since Assumption (ref) is the basic one, it would be desirable to have a testing device that checks the credibility of the crucial second one.
Testing Assumption (ref) is equivalent to testing the hypothesis
This is because by Assumption (ref), $E\left(\left.u_{it}\right|\chi_{it}\right)=0$, but if Assumption (ref) is not fulfilled, $E\left(\left.\eta_{it}\right|\chi_{it}\right) \ne 0$. Thus, the alternative encompasses all the possible departures from the null model. Let $\rho_{\chi}\left(\chi_{it}\right)$ be the probability density function of $\chi \in \mathbb{R}^d$, where $d=d_x+d_w$. Under $H_0$, since $E\left(\left. \epsilon_{it} \right|\chi_{it}\right) = 0$, we have
whereas under $H_1$, since $E\left(\left.\epsilon_{it}\right|\chi_{it}\right) = E\left[\left.m_0\left(\lambda_i,f_t\right)\right|\chi_{it}\right] - g(w_{it}) \ne 0$, we have
A sample analogue of $E\left[\epsilon_{it}E\left(\left.\epsilon_{it}\right|\chi_{it}\right)\rho_{\chi}\left(\chi_{it}\right)\right]$ is
where again we use a product kernel, namely \[ K_{it,js} = K\left(\frac{\chi_{it}-\chi_{js}}{h}\right); \quad K(v) = \prod^m_{l=1}k(v_l) \] with $h$ being the bandwidth, and $\widehat{\epsilon}_{it}$ the residuals, i.e., $\widehat{\epsilon}_{it} = y_{it} - x^{\prime}_{it}\widehat{\beta}- \widehat{g}(w_{it};h_w)$.
To derive asymptotic properties of $V_{NT}$, the following assumptions are used.
Then the asymptotic distribution under the null and the power of our test are given by the following theorems.
While on a first glimpse this seems to be an extremely strong result, indicating one could fully validate by a data-driven specification test the key identification condition, some caution is recommended. In fact, it is known that this type of identification conditions cannot be fully tested; instead one resorts to tests that give or rest credibility to the condition in question. In our case, for instance, it is not hard to see that under $H_1$, the power of our test hinges on the parts of $\eta$ that are either nonlinear in $x$ or reflect interactions between $x$ and $w$. Specifically, the stronger these are emphasized, the easier we reject, else we may not. This tells us how and when we can detect violations of Assumption (ref).
The asymptotic results for our estimators and test statistic are very helpful in showing consistency and convergence rate, and to better understand its statistical behavior. Nevertheless, it is less helpful for doing further inference in finite samples because estimates of those first-order asymptotic bias and variance are quite poor approximates in practice. A commonly employed remedy are resampling methods. This suggests using the so-called wild bootstrap, see MAMMEN:1992. Apart from being consistent for statistics based on local estimators like ours, it accounts for potential heteroscedasticity and allows us to accommodate nonparametric dependence structures in the error term. Recall that we only require independence for given $t$ but not between errors referring to the same individual.
First calculate $\widehat{\beta}$ and $\widehat{g}(w_{it})$ for all $w_{i,t}$ of the sample, obtaining residuals $\widehat{\epsilon}_{it}$. Next, for the given sample of regressors $\{ \chi_{it} = (x'_{it},w'_{it})' \}_{i=1,t=1}^{N,T}$, generate bootstrap output samples $\{ y_{it}^{*,b} \}_{i=1,t=1}^{N,T}$ by drawing $\vartheta_i^b \stackrel{i.i.d.}{\sim} N(0,1)$ for $b=1,\ldots,B$ and setting
It is easy to see that $E[\widehat{\epsilon}_{it} \vartheta_i^b] =0$, $Cov[\widehat{\epsilon}_{it} \vartheta_i^b, \widehat{\epsilon}_{js} \vartheta_j^b] = Cov[\widehat{\epsilon}_{it},\widehat{\epsilon}_{js}]$, for all $i$ and $j=1,\ldots,N$, $t$ and $s=1,\ldots,T$, i.e., the first two moments and dependence is maintained in the bootstrap. Third, calculate $\widehat{\beta}^{*,b}$ and $\widehat{g}^{*,b}(w_{it})$ for all bootstrap samples $b=1,\ldots,B$. Their sample means and variances give the bootstrap estimates of bias and variance of our estimators.
In the same way one can calculate bootstrap p-values for the test by comparing $V_{NT}/\widehat{\upsilon}_0$ with its bootstrap analogs $V_{NT}^{*,b}/\widehat{\upsilon}_0^{*,b}$. Note that the bootstrap samples are generated under the null hypothesis such that the bootstrap analogs follow approximately the finite sample distribution of $V_{NT}/\widehat{\upsilon}_0$ under $H_0$. Consequently, the p-value is the proportion of bootstrap statistics bigger than the original one.
For bootstrap tests, it is often recommended to employ in ((ref)) residuals $\widehat{\epsilon}_{it}$ obtained under the alternative. On the one hand, it is supposed this increased the power of the test, but in practice, this tends to produce over-rejection SPERLICH:2014. In our case, it is not even clear what a good predictor of residuals under the alternative could be. Therefore we stick to the former, more common practice of using the residuals obtained under $H_0$.
Regarding the bandwidth choice for estimation, it is to be noted that from a theoretical point of view, there does not exist a 'generally optimal' bandwidth since what is 'optimal' depends on the specific objective. More precisely, for estimating $\widehat{\beta}$ a different bandwidth is optimal than for estimating $\widehat{g}(w_{it})$, and both are different from the optimal testing bandwidths, not to mention the bandwidths optimal for generating the bootstrap samples SPERLICH:2014. If the null hypothesis is non- or semiparametric, as it is in our case, then it can become quite tedious to calibrate the test along all those different bandwidth choices as it is noted in RODRIGUEZ-POO_SPERLICH_VIEU:2015. In practice, we should be pragmatic and prefer a method that delivers reasonable estimates and provides a well-functioning test procedure. Consequently, we abstain here from the search for optimal bandwidths but propose a simple and computationally attractive solution. Then, it is helpful to recall that typically, the main interest is in $\beta$, not in function $g(\cdot)$. With this focus in mind, it turns out (see our simulations) that Silverman's rule-of-thumb, though invented for density estimation, is very useful in this context. When using for $k(\cdot )$ the Epanechnikov kernel, then the proposed bandwidth is $h_{w}=2.345\widehat{\sigma}_w(NT)^{-1/5}$, with $\widehat{\sigma}_w$ being the sample standard deviation of $w_{it}$. For more sophisticated bandwidth selection procedures, e.g., to estimate $g(\cdot )$ in an optimal way, we refer to the review of KOEHLER_SCHINDLER_SPERLICH:2014.
The aim is to illustrate the performance of our proposed estimators and test in finite samples using simulated data. We further compare our estimator with alternatives proposed in related panel data literature (i.e., the CCE estimator of PESARAN:2006 and the principal component approach (PCA) of BAI:2009). Hence, we employ the following data-generating processes (DGP)
for $i=1,\ldots,N$, $t=1,\ldots,T$. Note that $m_0\left(\lambda_i,f_t\right)=-w_{it}^2+2w_{it}+\xi_{m,it}$, where $\xi_{m,it}\sim IIDN(0,1)$, and each experiment has been replicated $1000$ times for $N=\{20,30,50,100\}$ and $T$ to be either $\{20,30,40,50,100\}$. In Section S1.1 of the Supplement, we specify how to generate the regressors $(x_{it},z_t)$, unobserved factors $(f_t)$, factor loadings $(\lambda_i)$, and individual-specific errors $(u_{it})$.
Although different choices of kernels and bandwidths would be feasible, we use the Epanechnikov kernel $k(u)=0.75(1-u^2)\mathbbm{1}\{|u|\leq1\}$ together with bandwidth $h_w=I_p h_{w}$, where $h_{w}$ is the Silverman's rule-of-thumb bandwidth as proposed above. For evaluation of the performance of our estimators, we use the bias and the root mean squared errors (RMSE) for the slope parameters, while the median of the RMSE is computed for the regression functions. For model ((ref)) we collect the results for $\beta_1$ in Tables (ref)-(ref). The results for $\beta_2$ are very similar and therefore not reported here.
Table (ref) tells us that the proposed estimator for $\beta_1$ seems to perform quite well in finite samples. As was expected from the results in Theorem (ref), the RMSEs of the estimator decrease as both $N$ and $T$ grow. Similar behavior is observed in Table (ref) for the nonparametric estimator proposed for $g(w_{it})$ corroborating the results in Theorem (ref).
To assess size and power of our test consider DGP ((ref)) with $m_0(\lambda_i,f_t)$ under $H_0$ and with $m_0^*(\lambda_i,f_t)=m_0(\lambda_i,f_t)+\Delta_N(4x_{it}^2+z_t^3-3w_{it})$ under $H_1$, for $0\leq\Delta_N\leq1$. This specification of $m_0^*(\lambda_i,f_t)$ enables us to evaluate the testing power over different values of $\Delta_N$ when it departs from the null ($\Delta_N=0$), by increasing $\Delta_N$ to $1$ at step length of $0.05$.
Table (ref) shows the obtained size for the proposed test at $1\%$, $5\%$, and $10\%$ significance levels, where critical values were calculated by wild bootstrap. Also, in Figure (ref) we plot the power functions of the test statistic against $\Delta_{N}$ for sample sizes $N1=20$ and $N4=100$ at $5\%$ significance level. Note that the results for the other percentile values (i.e., $10\%$ and $1\%$) and the power functions for the different values of $N$ are collected in the Supplement.
Looking at the results in Table (ref), we see that our test statistic $V_{NT}$ holds reasonably well for all sample sizes considered and for all percentile values of the null distribution of our test statistic. Also, Figure (ref) shows an impressive power of our test already for small and moderate samples, increasing rapidly with $\Delta_N$ even for the smallest value of $N$.
\setcounter{equation}{0}
We employ our methods to assess the effect of the European Union Emissions Trading System (EU-ETS) to decarbonize the economy and mitigate environmental degradation. As it is well-known, climate change is today's greatest environmental challenge and social concern. In recent years, the European Commission has been one of the most decisive institutions in leading the global energy transition, supporting the achievement of a low-carbon economy through targets and regulatory policies. The implementation of the EU-ETS in $2005$ has become the cornerstone of the European Union's strategy to decarbonize the economy and mitigate climate change BORGHESI_FRANCO_MARIN_2020. Nevertheless, despite the efforts made recently towards decarbonization, the burning of fossil fuels (carbon, oil, and gas) related to countries' economic activity continues to increase CO$_2$ emissions.
The relationship between economic development and environmental quality is an issue that has long puzzled economists, and there is a long tradition of employing the Environmental Kuznets Curve (EKC) to shed light on this issue. The EKC is based on the concept proposed by KUZNETS:1955 which posits an inverted U-shape relationship between income and environmental degradation GROSSMAN_KRUEGER:1993,GROSSMAN_KRUEGER:1995. In other words, the EKC postulates that environmental degradation rises with income during the initial phases of economic growth when income is relatively low since industry development causes great damage to the environment's quality. However, after passing a certain income threshold, this relationship reverses and becomes a negative one since higher levels of development are associated with a change in the economic structure in favor of industry and services that are more efficient and environmentally friendly. Since the pioneering work of GROSSMAN_KRUEGER:1993 several studies have tried to corroborate this EKC hypothesis, but so far there is no consensus on this relationship (see DINDA:2004, GALEOTTI_LANZA_PAULI:2006, or KAIKA_ZERVAS:2013a,KAIKA_ZERVAS:2013b, among others). This lack of consensus may be due to several misspecification errors related to the omission of relevant variables that can lead to inconsistent estimators and misleading inferences.
In this context, we propose to assess the effect of EU-ETS on CO$_2$ emissions through an extension of the conventional EKC specification based on the Stochastic Impacts by Regression on Population, Affluence and Technology (STIRPAT) model for evaluating environmental change that tries to overcome the above shortcomings DIETZ_ROSA:1997. On the one hand, we augment the model with a common stochastic covariate, $z_t$, to control for the aggregate effect of EU-ETS carbon pricing on CO$_2$. On the other hand, we propose to control for relevant omitted variables by allowing interactive fixed effects (i.e., $m_0\left(\lambda_i,f_t\right)$, where $f_t$ denotes a vector of unobserved common factors and $\lambda_i$ are the corresponding factor loadings that are allowed to be heterogeneous among countries $i$). Hence, the resulting panel data model to consider would be
where $CO_{2it}$ denotes the emissions of fossil CO$_2$ of country $i$ at time $t$, $gdp$ is the gross domestic product, $enit$ denotes technology which is proxied by energy intensity to capture technology's damaging effect on the environment, $pop$ denotes the population size, $z_t$ is the price of the carbon emissions set by the EU-ETS to all the country's members, and $u_{it}$ captures the innovation. Note that all the variables in ((ref)) are expressed in natural logarithms, so the estimated coefficients are all interpreted as elasticities.
The data used for this study covers EU27 countries plus the UK and Norway over the period $2005$-$2021$. The variables are derived from official sources and the specific definitions are detailed in Section S.2 of the Supplement. As control variables $w_{it}$ we use the natural logarithm of the total stock of debt liabilities as share of DGP that is collected by the International Monetary Fund and the age dependency ratio from the World Bank database.
The $\beta$ estimates are reported in Table (ref). We also provide estimates that we obtained using the CCE approach of PESARAN:2006 and naive estimates when unobserved factors were simply ignored. This was done to assess the potential misspecification problems related to the functional form of the interactive fixed effects. Estimates using the approach of BAI:2009 are not provided since the policy variable that we are evaluating in this paper is individual-invariant, and the proposed bootstrap in BAI:2009 does not converge.
The results in the first column of Table (ref) corroborate most of the postulates of the empirical literature (see DINDA:2004, LANTZ_FENG:2006, CHURCHILL_INEKWE_IVANOVSKI_SMYTH:2018, WANG_ZHANG_LI:2023, among others). On the one hand, the EKC hypothesis (i.e., a positive coefficient of $gdp$ and a negative one for $gdp^2$) for the income-pollution relationship is obtained, with the turning point at $gdp^* = \exp(-\beta_1/2\beta_2)$. On the other hand, the significant positive effect of population and energy intensity on environmental degradation is proven.
Finally, the results in Table (ref) enable us to endorse the inability of pre-established standard functional forms for interactive effects to capture certain non-linearities of the effects of unobserved common factors on individuals. More precisely, looking at the results in column (ii) we can see that the EKC hypothesis does not hold when we completely ignore the unobserved common factors, and the $gdp^2$ is not statistically relevant. In addition, when we follow PESARAN:2006's approach to control for the unobserved common factors, the effects of $gdp$ and $gdp^2$ are not the expected ones and are not statistically significant. Moreover, p-value of the test statistic related to the specification test that is proposed to check the crucial modeling assumption of the common factors is $0.740$. Therefore, the $H_0$ is not rejected supporting thereby our modeling proposal.
In this paper, identification, inference, and validation of a linear panel data model is considered that allows for an unspecified factor structure. This avoids the strong parametric restrictions that usually appear in the error term of traditional panel data models. The introduction of a conditional mean independence restriction between factors, factor loadings, and covariates converts the model into a partial linear model. A specification test to verify the crucial identification assumption made is introduced. For the parameters that attract most of the attention, consistent estimators that are asymptotically normal at the optimal rate are derived. Th specification test relies on combining the methodology of conditional moment tests and nonparametric estimation techniques. Using degenerate and nondegenerate theories of U-statistics we can show its convergence, asymptotic distribution under the null, and its divergence under the alternative at a rate arbitrarily close to $\sqrt{NT}$. The good performance of our estimators and test is confirmed by Monte Carlo experiments. Finally, the proposed approach is used to assess the effect of the EU ETS on CO2 emissions and the economic development of EU countries. The findings exhibit the practical relevance of our approach to avoid misleading conclusions, and they show the inability of existing popular methods to control well for the confounding effects of unobserved common factors.
Additional supporting information may be found in the Supplement of this article at the publisher's website. Section S1 includes the extended Mont Carlo experiment aiming to analyze the finite-sample performance of the proposed estimator and test statistic. Section S2 describes the data set used for the empirical application. Section S3 presents and proves some lemmas needed to show the main results.