EconBase
← Back to paper

Analysis of Multiple Long-Run Relations in Panel Data Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

221,238 characters · 28 sections · 1 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

-0.8inAnalysis of Multiple Long-Run Relations in Panel Data Models

abstractThe literature on panel cointegration is extensive but does not cover data sets where the cross section dimension, $n$, is larger than the time series dimension $T$. This paper proposes a novel methodology that filters out the short run dynamics using sub-sample time averages as deviations from their full-sample counterpart, and estimates the number of long-run relations and their coefficients using eigenvalues and eigenvectors of the pooled covariance matrix of these sub-sample deviations. We refer to this procedure as pooled minimum eigenvalue (PME). We show that PME estimator is consistent and asymptotically normal as $n$ and $T$ $\rightarrow \infty $ jointly, such that $T\approx n^{d}$, with $d>0$ for consistency and $d>1/2$ for asymptotic normality. Extensive Monte Carlo studies show that the number of long-run relations can be estimated with high precision, and the PME estimators have good size and power properties. The utility of our approach is illustrated by micro and macro applications using Compustat and Penn World Tables.

Keywords: Multiple long-run relations, Pooled Minimum Eigenvalue (PME) estimator, eigenvalue thresholding, panel data, cointegration, interactive time effects, financial ratios, Penn World Table

JEL Classification: C13, C23, C33, G30.

\thispagestyle{empty}

\pagenumbering{arabic}

Introduction

\onehalfspacing

This paper provides a new methodology for the analysis of multiple long-run relations in panel data models where the cross section dimension, $n$, is large relative to the time series dimension, $T$. While there is an extensive literature that considers multiple long-run (cointegrating) relations for time series models, for panel data models with large $n$ researchers have mainly focussed on a single long-run relation with known long-run causal links. The panel literature that does consider multiple long-run relations assumes $n$ is fixed as $T\rightarrow \infty ,$ or adopt sequential asymptotics whereby $T\rightarrow \infty $ first followed by $ n\rightarrow \infty $, effectively requiring $T$ to be large relative to $n$ and do not cover many applications of interest in economics and finance that involve many cross section units, such as firms and countries, observed over relatively short time spans. One example is empirical corporate finance, which investigates the stability of long-run relations, including financial ratios, using accounting data, such as Compustat, where thousands of firms are observed over relatively few time periods. \citeN{ColesLi2023} provide examples from a number of sub-fields of corporate finance, including: \ propensity to pay dividends, leverage, \ investment policy, and firm performance. Another example is cross country empirical growth studies that use data sets such as the Penn World Tables that provide annual data on a range of macro variables for as many as $n=183$ countries over different time periods, with a maximum time span of $T=70$. For both panel data sets one would expect multiple long-run relations between the variables, some of which may be the mean reverting ratios discussed in the macroeconomic and finance literatures.

This paper proposes an estimation and testing strategy that applies to panel data models with $n$ possibly much larger than $T$. We consider an $m\times 1 $ vector $\mathbf{w}_{it}$ for units $i=1,2,...,n$ over the time periods $ t=1,2,...,T$, with an unknown number, $r_{0}\in \left\{ 0,1,...,m-1\right\} $ , of linear combinations that are stationary. We refer to such linear combinations as long-run relations. If it is known that all elements of $ \mathbf{w}_{it}$ are $I\left( 1\right) $ then the stationary relations can be viewed as cointegrating relations. The focus of our analysis is to estimate $r_{0}$, the number of common long-run relations, and their coefficients, when $r_{0}\geq 1$. We filter out the short run dynamics by means of $q$ ($\geq 2$) non-overlapping sub-sample time averages, $\mathbf{ \bar{w}}_{i\ell },$ $\ell =1,2,...,q,$ as deviations from their full-sample counterpart, $\mathbf{\bar{w}}_{i\circ }$, namely $\mathbf{\bar{w}}_{i\ell }- \mathbf{\bar{w}}_{i\circ }$, and then construct a pooled sample covariance matrix of these deviations which we denote by $\mathbf{Q}_{\bar{w}\bar{w}}$. The number of long-run relations and their coefficients are estimated using the eigenvalues and eigenvectors of $\mathbf{Q}_{\bar{w}\bar{w}}$. We refer to this procedure as pooled minimum eigenvalue (PME), and note that it is simple to implement, extends readily to unbalanced panels, and is shown to be robust to stationary interactive time effects. It is semi-parametric since it does not require modelling the short run dynamics and applies to general linear process, thus allowing for moving average processes and is not confined to vector autoregressions (VAR). Most importantly, the PME approach does not require knowing long-run causal linkages that might exist amongst the variables under consideration. To our knowledge, no other panel estimation procedure exists for such a setting.

Denoting the first $r_{0}$ eigenvectors of $\mathbf{Q}_{\bar{w}\bar{w}}$ by $ \mathbf{\hat{{\Greekmath 010C}}}_{j0},$ for $j=1,2,...,r_{0}$, we then consider structural estimation of the long-run relations assuming they are subject to $r_{0}\times r_{0}$ exact identifying restrictions. Assuming $r_{0}$ is known, we derive the asymptotic distribution of the exactly identified long-run relations and propose consistent estimators for their covariance matrices that does not require estimation of the dynamics of individual $ \mathbf{w}_{it}$ processes; thus allowing us to test restrictions on the elements of the exactly identified long-run relations with relatively short $ T$. The identified long-run relations are shown to be consistent and asymptotically normally distributed as $n$ and $T$ $\rightarrow \infty $ jointly such that $T\approx n^{d}$. For consistency only $d>0$ is required, but for asymptotic normality a faster relative rate of $d>1/2$ is required. Many panel time series estimation and inference procedures require $d>1$ ($ n/T\rightarrow 0$). Our requirement $d>1/2$ indicates that the procedure should work well with large $n$ and moderate $T$ allowing one to estimate the coefficients of multiple long-run relations in such cases, without estimating short-run dynamics.

We propose to estimate $r_{0}$ by the number of eigenvalues of $\mathbf{Q}_{ \bar{w}\bar{w}}$ that fall below a given threshold $C_{T}=CT^{-{\Greekmath 010E} }$, for some $C>0$ and ${\Greekmath 010E} >0$. One could use cross validation procedures to set $C$ and ${\Greekmath 010E} $, but based on extensive Monte Carlo experiments we have found that setting $C=1$ works well if we base our selection procedure on the eigenvalues of the correlation matrix, $\mathbf{R}_{_{\bar{w }\bar{w}}}=\left[ diag\left( \mathbf{Q}_{\bar{w}\bar{w}}\right) \right] ^{-1/2}\mathbf{Q}_{\bar{w}\bar{w}}\left[ diag\left( \mathbf{Q}_{\bar{w}\bar{w }}\right) \right] ^{-1/2}$. Such an estimator can be written conveniently as $\tilde{r}=\sum_{j=1}^{m}\mathcal{I}\left( \tilde{{\Greekmath 0115}}_{j}<T^{-{\Greekmath 010E} }\right) ,$ where $\tilde{{\Greekmath 0115}}_{j},$ $j=1,2,...,m$ are the eigenvalues of $\mathbf{R}_{_{\bar{w}\bar{w}}}$, and $\mathcal{I}\left( \mathcal{A} \right) =1$ if $\mathcal{A}$ is true and zero otherwise. In the Monte Carlo experiments and the empirical applications we report results for ${\Greekmath 010E} =(1/4,1/2)$, and find that overall setting ${\Greekmath 010E} =1/4$ works well.

Monte Carlo experiments show near-perfect performance of $\tilde{r}$ as an estimator of $r_{0}$, when ${\Greekmath 010E} =1/4$, for all $\,n=50$, $500$, $1,000$, $ 3,000$ and $T=20$, $50$, $100$ sample size combinations and across a large number of VAR and VARMA data generating processes, with and without interactive time effects, non-Gaussian errors, generalized autoregressive conditional heteroskedasticity (GARCH), threshold autoregressions (TAR), and for different patterns of long-run causal ordering. $\tilde{r}$ performs almost equally well for the smallest sample sizes of $T=20$ and $n=50$, as well as for the largest $T=100$ and $n=3,000$. As an alternative approach we considered Johansen's trace tests applied to each cross section unit separately, and then estimated $r_{0}$ by the simple average of these individual estimates. This was done purely for comparison since to the best of our knowledge there are no other methods that apply to panels in the literature.\ We found that this average type estimator performed reasonably well when $T$ was large, but still fell short as compared to the thresholding estimator.

The finite sample performance of PME estimator of the coefficients of the long-run relations\ is found to be satisfactory with inference based on PME estimator with $q=2$ (sub-sample time averages) generally more accurate in terms of empirical size of the tests, compared with $q=4$, in line with intuition that suggest a larger choice of $q$ is likely to result in a larger finite-sample bias. We also considered a simple VAR(1) design with $ m=2$, $r_{0}=1$ and one-way long-run causality to see how PME performs compared to the many single equation estimators proposed in the literature (and cited below). We found that in this simple case the PME\ estimator is less efficient in terms of root mean square errors only when $T=100$. However, PME with $q=2$ proved to be less biased and performed much better in terms of size than the single equation approaches for all sample size combinations.

To illustrate the utility of PME procedure we present one micro and one macro application. The micro application considers a number of key financial variables (in logs) and investigates if they are cointegrated, and whether financial ratios can be regarded as stationary variables. To this end we used accounting data for individual firms from CRSP/Compustat on their book value (BV), market value (MV), short-term debt (SD), long-term debt (LD), total assets (TA) and total debt outstanding (DO). The panels involving these variables are unbalanced and cover the period $1950-2021$. We consider firms with at least $20$ years of data, with $n$ varying between about $ 1,000 $ and $2,500$. The variables are grouped into three sets, where we have prior expectations about possible cointegration and identification. The first set considered has just two variables: the logarithm of total debt outstanding and logarithm of total assets: \{$DO_{it}$, $TA_{it}$\}. The ratio of total debt outstanding to total assets is often used as a measure of leverage, which suggests a single hypothesized long-run relation. The other two variable sets are: the logarithms of short and long term debt and total assets, \{$SD_{it}$, $LD_{it}$, $TA_{it}$\}; and the logarithms of total debt outstanding, book value and market value, \{$DO_{it}$, $BV_{it}$, $MV_{it}$\}. We expect two hypothesized long-run relations in these sets with three variables. For each set of variables we provide estimates for the full sample 1950-2021 as well as for a shorter sample that ends in 2010. The estimates provide strong evidence of one long-run relation when we consider two variables, and, with one exception, two long-run relations when we consider panels with three variables. In the case of panels with $m=2$, we illustrate that the PME estimates are invariant to normalization, which is in contrast to the estimates obtained using panel regressions that depend on which way the regression is run. For the relation between logarithms of debt and total assets, we find the estimates of the long-run coefficients are close to one in all cases, ranging from $1.113$ to $1.143$, and precisely estimated. In the case of panels with $m=3$ we find $\tilde{r}=2$, and the null hypothesis that long-run coefficients are equal to unity is not rejected in about a quarter of the panel estimates. These results provide partial support for use of logarithm of financial ratio in corporate finance. In cases where the use of log ratio is not supported, one could use the PME estimates of long-run relations in second stage regressions on stationary variables that also include short run dynamics as well as other stationary variables.

The macro application investigates long-run relations using unbalanced cross country macroeconomic time series data from the Penn World Tables, featuring up to $n=177$ countries over the years $1950-2019$. This dataset has a much smaller cross-section dimension and a larger average time dimension compared with the micro application. We focus on four key macro variables: per capita real merchandise exports ($ex_{it}$) and imports ($im_{it}$), real labour productivity per hour worked ($prod_{it}$), and real wages per hour worked ($ wage_{it}$). The choice of these variables was motivated by two widely maintained hypotheses. Firstly, real wages and productivity should balance for steady state growth to be feasible. Secondly export and imports should balance for international solvency, though the constraint may not be binding for reserve-currency countries such as the US. These hypotheses are largely confirmed for emerging economies, and, with notable departures from unit long-run elasticities, also for advanced economies. In addition, when we consider all the four variables together we uncover cross country evidence on the long-run relation between exports and productivity without making any assumption about the direction of causality between these variables.

Related literature: We first discuss the literature for a single time series process, which could be viewed as the $p\times 1$ ($p=m$ $n$) stacked vector, $\mathbf{w}_{t}=(\mathbf{w}_{1t}^{\prime },\mathbf{w} _{2t}^{\prime },...,\mathbf{w}_{nt}^{\prime })^{\prime }$. Our approach is related to that of \citeANP{PhillipsOuliaris1990} (\citeyearNP{PhillipsOuliaris1988}, \citeyearNP{PhillipsOuliaris1990}) in that they also start from general linear processes. They propose testing the null of no cointegration using the smallest eigenvalues of the spectral density of $\Delta \mathbf{w}_{t}$ evaluated at zero frequency. However, it is difficult to obtain reasonably precise estimates of the spectral density, particularly in the presence of high persistence in first differences. Attempting to eliminate the effects of the short run dynamics by using time averages of sub-samples of time series data is also widely used. \citeN{MuellerWatson2018} consider using sub-sample averages to estimate the long-run relation between two variables $(y_{t}$ and $x_{t})$. Their estimated long-run coefficient from regression of sub-sample averages of $y_{t}$\ on those of $x_{t}$ \ is not the same as the reciprocal of the estimate that will be obtained from the reverses regression. A panel version of their procedure can be considered, but will be subject to the same limitations, namely it can handle only one long-run relation and will require knowing the direction of long-run causality. In not requiring any assumptions regarding the direction of long-run causality our approach is comparable to the maximum likelihood approach pioneered by \citeANP{Johansen1988} (\citeyearNP{Johansen1988}, \citeyearNP{Johansen1991a}) that allows for multiple long-run relations without assuming any long-run causal orderings of the variables, but assumes a $VAR(s)$ specification in $ \mathbf{w}_{t}$ where $p$ and $s$ are fixed (and quite small) relative to $T$ . \citeANP{Onatski2018} (\citeyearNP{Onatski2018}, \citeyearNP{Onatski2019}) \ investigate the asymptotic properties of Johansen test when $ \mathbf{w}_{t}$ follows VAR(1) but allow $p,T\rightarrow \infty ,$ such that $p/T\rightarrow c\in (0,1]$. They provide theoretical arguments why Johansen's test of cointegration rank is likely to be severely over-sized even if $p$ takes moderate values. Extensions to higher order VARs are provided by Bykhovskaya2022. This is a promising approach which is yet to be fully developed for the analysis of multiple cointegrations across many units, which is the primary focus of this paper. Since the ordering of the variables in the VAR does not affect the Johansen's tests of the cointegration rank, without further restrictions the use of high-dimensional VARs in $\mathbf{w}_{t}$ does not distinguish between cointegration across units as compared to cointegration between the variables specific to the cross section units. Also, the condition $p/T=nm/T\rightarrow c\in (0,1]$ is unlikely to be met when $m>1$ and $n$ is of the same order of magnitude as $ T $.

Turning to the panel cointegration literature, most studies consider $I(1)$ variables with a single cointegrating vector where the direction of long-run causality is known. These estimators are typically generalizations of the time series procedures such as the panel Fully Modified OLS of \citeANP{Pedroni1996} (\citeyearNP{Pedroni1996}, \citeyearNP{Pedroni2001}, \citeyearNP{Pedroni2001ReStat}) , the Pooled Mean Group (PMG) estimator of \citeN{PesaranShinSmith1999} ,or the panel Dynamic OLS of \citeN{MarkSul2003} . There are panel generalizations of Johansen's approach, such as \citeN{GroenKleibergen2003} , and \citeN{LarssonLyhagen2007} , which can be used to test for the number of cointegrating relations and estimate their parameters. These are based on a vector error correction model, VECM, which can deal with multiple cointegrating vectors, but require $T$ to be large relative to $n.$ In their applications \citeN{LarssonLyhagen2007} have $m=3,$ $n=4.$ \citeN{Breitung2005} proposes a systems estimator, but that requires that every cross section unit cointegrate. \citeN{ChudikPesaranSmith2023} suggest system pooled mean group estimator for a single common long-run relation coefficient ${\Greekmath 0112} $, that can handle any long-run causal ordering and allow some units to fail to cointegrate, but again requires $T$ to be large relative to $n$. \

A large number of other topics have been examined within the context of a panel with a single cointegrating relation. These include: estimation with I(1) latent factors: \citeN{BaiKaoNg2009} and \citeN{KapetaniosPesaranYamagata2011} ; structural breaks: \citeN{BanerjeeCarrion2024} and \citeN{DitzenKaraviasWesterlund2025} ; and non-linear effects: \citeN{deJongWagner2025} . Further details can be found in the surveys by \citeN{Breitung2008} and \citeN{Choi2015a} that also cover testing for cointegration using residuals ( \citeANP{Westerlund2005}, \citeyearNP{Westerlund2005} ), and second generation panel unit root tests allowing for cross section dependence ( \citeANP{Pesaran2007}, \citeyearNP{Pesaran2007} ). As this brief overview indicates, none of the methods advanced in the literature consider multiple long-run relations when $n>>T$.

Outline of the paper: The rest of the paper is set out as follows: Section (ref) sets out the panel data model and introduces the PME estimator. Section (ref) introduces the assumptions and discusses the identification conditions. Section (ref) gives a formal description of the PME estimator and some of its asymptotic properties. Section (ref)\ considers identification and derives the asymptotic distribution of the exactly identified PME estimator. Section (ref) shows how $r_{0}$ can be estimated by eigenvalue thresholding. Section (ref) allows for interactive time effects. Section (ref) discusses the choice of the number of sub-sample averages, $q$, and how to set the parameters of the thresholding estimator of $r_{0}$. Section (ref) provides Monte Carlo evidence on the small sample properties of the PME estimators of $r_{0}$ and $\mathbf{{\Greekmath 010C} }_{j0}$, $j=1,2,...,r_{0}$. Section (ref) discusses the empirical applications, and Section (ref) provides some concluding remarks. The proofs of the propositions and theorems are provided in an appendix, with related lemmas given in a supplement. This supplement also includes sub-sections on extensions of PME to panels with interactive time effects, on how to implement the proposed estimator for unbalanced panels, and gives details of the data generating processes used in the Monte Carlo experiments, plus additional information on data sources and the construction of the variables used in the empirical applications.

Notations: Matrices are denoted by bold upper case letters and vectors are denoted by bold lower case letters. All vectors are column vectors. $\left\Vert \mathbf{x}\right\Vert $ denotes the Euclidean norm of a vector $\mathbf{x}$. $rank\left( \mathbf{A}\right) $ denotes the column rank of $\mathbf{A}$. $vec\left( \mathbf{A}\right) $ denotes vectorization of $ \mathbf{A}$. $tr(\mathbf{A})$ denotes the trace of a square matrix $\mathbf{A }$. Eigenvalues of $m\times m$ symmetric positive semi-definite real matrix $ \mathbf{A}$ sorted in ascending order are $0\leq {\Greekmath 0115} _{1}(\mathbf{A} )\leq {\Greekmath 0115} _{2}(\mathbf{A})\leq ...\leq {\Greekmath 0115} _{m}(\mathbf{A})$. $ \lVert \mathbf{A}\rVert $ is the spectral norm of $\mathbf{A}$. Small and large finite positive constants that do not depend on sample sizes $n$ and $ T $ are denoted by ${\Greekmath 010F} $ and $K$, respectively. These constants can take different values at different instances in the paper. $T_{n}\approx n^{d}$ if {there exist $n_{0}\geq 1$ and positive constants }${\Greekmath 010F} ${\ and }$K${, such that }$\inf_{n\geq n_{0}}\left( T_{n}/n^{d}\right) \geq {\Greekmath 010F} ${\ and $\sup_{n\geq n_{0}}\left( T_{n}/n^{d}\right) \leq K$.} {For simplicity of exposition we omit subscript }$n$ and write $T\approx n^{d}$. Convergence in probability and distribution are denoted by $\rightarrow _{p}$ and $\rightarrow _{d}$, respectively. In this paper $o_{p}\left( 1\right) $ is short for sequence of random variables, random vectors or random matrices that converge to zero in probability as $n\rightarrow \infty $ for all values of $d>1/2$. $\mathbf{A}_{n}=O_{p}\left( 1\right) $ if sequence $ \mathbf{A}_{n}$ is bounded in probability. $a_{n}=O(b_{n})$\ denotes the deterministic sequence $\left\{ a_{n}\right\} $\ is at most of order $ b_{n}$. Equivalence of asymptotic distributions is denoted by $ \overset{a}{\thicksim }$.

Preliminaries

Consider the following general linear model for $\mathbf{w}_{it}$

equation[equation omitted — 232 chars of source]

where $\mathbf{w}_{it}$ is an $m\times 1$ vector of outcomes, $\mathbf{a} _{i} $ is $m\times 1$ vector of fixed effects, $\mathbf{f}_{t}$ is a vector of stationary latent factors with associated loading matrices, $\mathbf{G} _{i}$. $\mathbf{s}_{it}$ is the partial sum process defined by

equation[equation omitted — 351 chars of source]

$\mathbf{u}_{it}$ is independently distributed over $i$ and $t$ with mean zero and the $m\times m$ positive definite matrix, $\mathbf{\Sigma }_{i}$, $ \mathbf{C}_{i}$ is an $m\times m$ matrix of fixed coefficients, and $\mathbf{ v}_{it}=\mathbf{C}_{i}^{\ast }(L)\mathbf{u}_{it}$, where $\mathbf{C} _{i}^{\ast }(L)=\sum_{\ell =0}^{\infty }\mathbf{C}_{i\ell }^{\ast }L^{\ell }$ . This model covers many specifications of interest such as vector autoregressions, error correction models, as well as first-differenced stationary models. It allows for interactive time effects which reduce to time effects under the so-called parallel trends assumption, namely setting $ \mathbf{G}_{i}=\mathbf{G}$ for all $i$. In stacked form the model for all $n$ units can be written as $\mathbf{w}_{t}=\mathbf{a+Gf}_{t}+\mathbf{Cs}_{t}+ \mathbf{C}^{\ast }(L)\mathbf{u}_{t},$ where $\mathbf{w}_{t}=(\mathbf{w} _{1t}^{\prime },\mathbf{w}_{2t}^{\prime },...,\mathbf{w}_{nt}^{\prime })^{\prime },$ $\mathbf{a=}\left( \mathbf{a}_{1}^{\prime },\mathbf{a} _{2}^{\prime },...,\mathbf{a}_{n}^{\prime }\right) ^{\prime }$, $\mathbf{G=(G }_{1}^{\prime },\mathbf{G}_{2}^{\prime },...,\mathbf{G}_{n}^{\prime })^{\prime }$, and $\mathbf{u}_{t}=(\mathbf{u}_{1t}^{\prime },\mathbf{u} _{2t}^{\prime },...,\mathbf{u}_{nt}^{\prime })^{\prime }$. Under our specification $\mathbf{C}$ and $\mathbf{C}_{i}^{\ast }(L)$ are assumed to be block-diagonal matrices with $\mathbf{C}_{i}$ and $\mathbf{C}_{i}^{\ast }(L)$ as their $i^{th}$ block, respectively. Such restrictions seem inevitable when $n$ is large relative to $T$, and seems plausible considering that we allow for cross-sectional dependence through the common factors, $\mathbf{f} _{t}$.

Assuming that $\mathbf{C}_{i}$ has rank $m-r_{0}>0$ for all $i$, we are interested in estimating $r_{0}$, and the associated stationary linear combinations defined by $\mathbf{{\Greekmath 010C} }_{j0}^{\prime }\mathbf{w}_{it},$ $ j=1,2,...,r_{0},$ where $\mathbf{B}_{0}=\left( \mathbf{{\Greekmath 010C} }_{10},\mathbf{ {\Greekmath 010C} }_{20},...,\mathbf{{\Greekmath 010C} }_{r_{0}0}\right) $ is the $m\times r_{0}$ matrix of long-run relations that are common across all $i$, and satisfies $ \mathbf{B}_{0}^{\prime }\mathbf{C}_{i}=0$. We also consider estimation of long-run relations subject to the exactly identifying restrictions that are motivated by the theory. The estimator we propose involves splitting the data for each unit into $q\geq 2$ sub-samples; taking time averages of these sub-samples and forming a pooled demeaned covariance matrix we label $ \mathbf{Q}_{\bar{w}\bar{w}}$. The eigenvalues of this matrix allow us to estimate $r_{0},$ and the eigenvectors corresponding to the first $r_{0}$ eigenvalues provide estimates of $\mathbf{{\Greekmath 010C} }_{j0}$. But to simplify the exposition and focus on the main contribution of the paper, initially we abstract from the interactive time effects, but return to this complication in Section (ref), where we show that our analysis remains valid so long as the latent factors are stationary.

It is possible to allow for non-linear features, such as GARCH and threshold autoregressions, so long as the effects of shocks to $\mathbf{u}_{it}$ decay exponentially fast. But to keep the theoretical analyses relatively simple, we only consider the robustness of our estimation and testing strategies to such non-linear effects using Monte Carlo experiments.

Assumptions and identification conditions

We directly work with ((ref)) and make the following assumptions:

assumption\ The error terms, $\mathbf{u}_{it}$, are distributed independently over $i=1,2,...,n$ and $t=1,2,....,T$ with $E(\mathbf{u}_{it})= \mathbf{0,}$ and the covariance matrix $\,E(\mathbf{u}_{it}\mathbf{u} _{it}^{\prime })=\mathbf{\Sigma }_{i}$, where $\mathbf{\Sigma }_{i}$ is a positive definite matrix, $\inf_{i}\mathbf{{\Greekmath 0115} }_{1}\left( \mathbf{ \Sigma }_{i}\right) >{\Greekmath 010F} $, $\sup_{i}\mathbf{{\Greekmath 0115} }_{m}\left( \mathbf{\Sigma }_{i}\right) <K$, and $\sup_{it}E\left\Vert \mathbf{u} _{it}\right\Vert ^{4+{\Greekmath 010F} }<K$, for some ${\Greekmath 010F} >0$.
assumptionThe coefficient matrices $\mathbf{C}_{i}$ and $\mathbf{C} _{i\ell }^{\ast }$, are non-stochastic constants such that $ \sup_{i}\left\Vert \mathbf{C}_{i}\right\Vert <K$ and $\sup_{i}\left\Vert \mathbf{C}_{i\ell }^{\ast }\right\Vert <K{\Greekmath 011A} ^{\ell }$, where ${\Greekmath 011A} $ lies in the range $0<{\Greekmath 011A} <1$. The $m\times m$ matrix $\mathbf{C}_{i}$ has rank $ m-r_{0}$, for $i=1,2,...,n.$
assumption(a) Let \begin{equation} \mathbf{\Psi }_{n}=n^{-1}\sum_{i=1}^{n}\mathbf{C}_{i}\mathbf{\Sigma }_{i} \mathbf{C}_{i}^{\prime }\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, and }\mathbf{\Psi }=lim_{n\rightarrow \infty }\mathbf{\Psi }_{n}. \end{equation} Then there exists $n_{0}$ such that for all $n>n_{0}$, $rank(\mathbf{\Psi } _{n})=rank(\mathbf{\Psi })=m-r_{0}>0$. (b) The orthonormalized eigenvectors associated with the $r_{0}$ zero eigenvalues of $\ \left( \frac{q-1}{6q} \right) \mathbf{\Psi }_{n}$\ are denoted by $\mathbf{{\Greekmath 010C} }_{j0}$, for $j=1,2,...,r_{0}$, and the orthonormalized eigenvectors associated with the ordered non-zero eigenvalues of $\left( \frac{q-1}{6q}\right) \mathbf{ \Psi }_{n}$, namely ${\Greekmath 0115} _{r_{0}+1}\leq {\Greekmath 0115} _{r_{0}+2}\leq ...\leq {\Greekmath 0115} _{m}$, by $\mathbf{{\Greekmath 010C} }_{j}$, for $r_{0}+1,r_{0}+2,...,m$. Specifically \begin{equation} \left( \frac{q-1}{6q}\right) \mathbf{\Psi }_{n}\mathbf{{\Greekmath 010C} }_{j0}=0\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, for }j=1,2,...,r_{0}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,} \end{equation} and \begin{equation} \left( \frac{q-1}{6q}\right) \mathbf{\Psi }_{n}\mathbf{{\Greekmath 010C} }_{j}={\Greekmath 0115} _{j}\mathbf{{\Greekmath 010C} }_{j}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, for }j=r_{0}+1,r_{0}+2,...,m\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{.} \end{equation}
remarkUnder Assumptions (ref)-(ref) $\left\{ \mathbf{v}_{it}\right\} $ has absolute summable autocovariances, $\sum_{h=0}^{\infty }\sup_{i}\left\Vert \mathbf{\Gamma }_{i}(h)\right\Vert <K$, where $\mathbf{ \Gamma }_{i}(h)=E\left( \mathbf{v}_{it}\mathbf{v}_{i,t-h}^{\prime }\right) $ is the autocovariance function of $\mathbf{v}_{it}$. See Lemma (ref) in the supplement, Section (ref).
remarkUnder Assumptions (ref), $\mathbf{{\Greekmath 010C} }_{j0}^{\prime }\mathbf{\Psi } _{n}\mathbf{{\Greekmath 010C} }_{j0}=0$, for $j=1,2,...,r_{0}$ and ${\Greekmath 0115} _{j}=\left( \frac{q-1}{6q}\right) \mathbf{{\Greekmath 010C} }_{j}^{\prime }\mathbf{\Psi }_{n}\mathbf{ {\Greekmath 010C} }_{j}>0$, for $j=r_{0}+1,r_{0}+2,...,m$. Nonzero eigenvalues\ ${\Greekmath 0115} _{j}$, for $j=r_{0}+1,r_{0}+2,...,m$, and the corresponding eigenvectors $\mathbf{{\Greekmath 010C} }_{j}$, for $j=r_{0}+1,r_{0}+2,...,m$, depend on $n$, but to simplify the notations we avoid using the subscript $n$.
remarkUnder part (b) of Assumption (ref) $\mathbf{\Psi }_{n}$ and $\mathbf{\Psi }$ can be written as $\mathbf{\Psi }_{n}=\mathbf{P} _{n}^{\prime }\mathbf{P}_{n}$, and $\mathbf{\Psi =P}^{\prime }\mathbf{P,}$ where $\mathbf{P}_{n}$ and $\mathbf{P}$ are $(m-r_{0})\times m$ full rank matrices, with $rank\left( \mathbf{P}_{n}\right) =rank\left( \mathbf{P} \right) =m-r_{0}$, for all $n>n_{0}$.
remarkCondition $rank\left( \mathbf{C}_{i}\right) =m-r_{0}$ for all $i=1,2,...,n$ in Assumption (ref) can be relaxed to allow for no stochastic trends for some cross section units, so long as the rank condition $rank\left( \mathbf{\Psi }_{n}\right) =rank\left( \mathbf{\Psi }\right) =m-r_{0}$ holds. Specifically, suppose that $\mathbf{C}_{i}=\mathbf{0}$ for $i=1,2,...,n_{1}$ , but the cointegration rank condition holds for the remaining units. Then \begin{equation*} \mathbf{\Psi }_{n}=n^{-1}\sum_{i=1}^{n}\mathbf{C}_{i}\mathbf{\Sigma }_{i} \mathbf{C}_{i}^{\prime }\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi=(1-{\Greekmath 0119} _{n})\left( \frac{1}{n-n_{1}} \sum_{i=n_{1}+1}^{n}\mathbf{C}_{i}\mathbf{\Sigma }_{i}\mathbf{C}_{i}^{\prime }\right) ,\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi \end{equation*} where ${\Greekmath 0119} _{n}=n_{1}/n$ is the proportion of units without stochastic trends. Then the rank requirement continues to hold if ${\Greekmath 0119} _{n}>0$, and $ \frac{1}{n-n_{1}}\sum_{i=n_{1}+1}^{n}\mathbf{C}_{i}\mathbf{\Sigma }_{i} \mathbf{C}_{i}^{\prime }$ tends to a matrix having rank $m-r_{0}$. But for clarity of exposition we maintain Assumption (ref) without loss of generality.

Estimation of long-run relations

Introducing sub-sample time averages

We base our estimation procedure on non-overlapping sub-sample time averages of $\mathbf{w}_{it}$. For the ease of exposition, suppose the panel data under consideration is balanced, $T$ is divisible by $q$, and consider $q$ $ (\geq 2)$ non-overlapping time averages of equal length $T_{q}$ defined by

equation[equation omitted — 278 chars of source]

where $T_{q}=T/q$. To simplify the exposition we abstract from interactive effects and apply the above time average operator to ((ref)) to obtain \footnote{ Sections (ref) and (ref) of the supplement consider unbalanced panels and models with interactive effects.}

equation[equation omitted — 236 chars of source]

where $\mathbf{\bar{v}}_{i\ell }=T_{q}^{-1}\sum_{t=(\ell -1)T_{q}+1}^{\ell T_{q}}\mathbf{v}_{it}$ and $\mathbf{\bar{s}}_{i\ell }=T_{q}^{-1}\sum_{t=(\ell -1)T_{q}+1}^{\ell T_{q}}\mathbf{s}_{it}.$ We now use standard de-meaning procedure and eliminate $\mathbf{a}_{i}$ from ((ref)) to obtain

equation[equation omitted — 322 chars of source]

where $\mathbf{\bar{w}}_{i\circ }=q^{-1}\sum_{\ell =1}^{q}\mathbf{\bar{w}} _{i\ell }$, and similarly $\mathbf{\bar{s}}_{i\circ }=q^{-1}\sum_{\ell =1}^{q}\mathbf{\bar{s}}_{i\ell }$ and $\mathbf{\bar{v}}_{i\circ }=q^{-1}\sum_{\ell =1}^{q}\mathbf{\bar{v}}_{i\ell }$. Consider now the $ m\times m$ pooled sample covariance matrix

equation[equation omitted — 114 chars of source]

where

equation[equation omitted — 242 chars of source]

The limiting value of $\mathbf{Q}_{\bar{w}\bar{w}}$, as $n,T\rightarrow \infty $, plays a critical role in our approach to estimation of long-run relations. Using ((ref)) in ((ref)) we first note that

equation[equation omitted — 300 chars of source]

where $T\mathbf{Q}_{\bar{s}_{i}\bar{s}_{i}}=q^{-1}\sum_{\ell =1}^{q}\left( \mathbf{\bar{s}}_{i\ell }-\mathbf{\bar{s}}_{i\circ }\right) \left( \mathbf{ \bar{s}}_{i\ell }-\mathbf{\bar{s}}_{i\circ }\right) ^{\prime },$ $T\mathbf{Q} _{\bar{s}_{i}\bar{v}_{i}}=q^{-1}\sum_{\ell =1}^{q}\left( \mathbf{\bar{s}} _{i\ell }-\mathbf{\bar{s}}_{i\circ }\right) \left( \mathbf{\bar{v}}_{i\ell }- \mathbf{\bar{v}}_{i\circ }\right) ^{\prime }=T\mathbf{Q}_{\bar{v}_{i}\bar{s} _{i}}^{\prime },$ and $T\mathbf{Q}_{\bar{v}_{i}\bar{v}_{i}}=q^{-1}\sum_{\ell =1}^{q}\left( \mathbf{\bar{v}}_{i\ell }-\mathbf{\bar{v}}_{i\circ }\right) \left( \mathbf{\bar{v}}_{i\ell }-\mathbf{\bar{v}}_{i\circ }\right) ^{\prime } $. Since $\left\{ \mathbf{v}_{it}\right\} $ is covariance stationary with absolute summable autocovariances and $\left\{ \mathbf{s}_{it}\right\} $ is a partial sum process it then follows that $\mathbf{\bar{v}}_{i\ell }- \mathbf{\bar{v}}_{i\circ }=$ $O_{p}\left( T^{-1/2}\right) $, and $\mathbf{ \bar{s}}_{i\ell }-\mathbf{\bar{s}}_{i\circ }=O_{p}(T^{1/2})$. Moreover, as established in Lemma (ref),

equation[equation omitted — 357 chars of source]

Pooled minimum eigenvalue (PME) estimator

Our proposed estimation procedure is based on eigenvalues and eigenvectors of $\mathbf{Q}_{\bar{w}\bar{w}}$, defined by ((ref)). Averaging $ \mathbf{Q}_{\bar{w}_{i}\bar{w}_{i}}$ in ((ref)) over all cross section units now yields:

equation[equation omitted — 437 chars of source]

The pooled minimum eigenvalue (PME) estimator of $\mathbf{{\Greekmath 010C} }_{j0},$ $ j=1,2,...,r_{0}$, is given by the $j^{th}$ orthonormalized eigenvector of $ \mathbf{Q}_{\bar{w}\bar{w}}$ , $\mathbf{\hat{{\Greekmath 010C}}}_{j}$, associated with its $r_{0}$ smallest eigenvalues, $\hat{{\Greekmath 0115}}_{1}\leq \hat{{\Greekmath 0115}} _{2}\leq ...\leq \hat{{\Greekmath 0115}}_{r_{0}}$. Specifically, for $j=1,2,....,m$, $ \mathbf{Q}_{\bar{w}\bar{w}}\mathbf{\hat{{\Greekmath 010C}}}_{j}\mathbf{=}\hat{{\Greekmath 0115}} _{j}\mathbf{\hat{{\Greekmath 010C}}}_{j}$, such that $\mathbf{\hat{{\Greekmath 010C}}} _{j}^{\prime }\mathbf{\hat{{\Greekmath 010C}}}_{j}=1$, and $\hat{{\Greekmath 0115}}_{j}=\mathbf{ \hat{{\Greekmath 010C}}}_{j}^{^{\prime }}\mathbf{Q}_{\bar{w}\bar{w}}\mathbf{\hat{{\Greekmath 010C}}} _{j}$. In matrix notations we have

equation[equation omitted — 288 chars of source]

and the PME estimator of $\mathbf{B}_{0}$ is given by $\mathbf{\hat{B}} _{0}=\left( \mathbf{\hat{{\Greekmath 010C}}}_{1},\mathbf{\hat{{\Greekmath 010C}}}_{2},...,\mathbf{ \hat{{\Greekmath 010C}}}_{r_{0}}\right) .$

Consistency of the PME estimator

Under Assumption (ref),$\ \mathbf{Q}_{\bar{s}_{i}\bar{s}_{i}}$ has the following exact moment (established in Lemma (ref))

equation[equation omitted — 162 chars of source]

Hence $n^{-1}\sum_{i=1}^{n}\mathbf{C}_{i}E\left( \mathbf{Q}_{\bar{s}_{i}\bar{ s}_{i}}\right) \mathbf{C}_{i}^{\prime }=\frac{(q-1)}{6}\left( \frac{1}{q}+ \frac{1}{T^{2}}\right) \mathbf{\Psi }_{n}\mathbf{,}$ where $\mathbf{\Psi } _{n}$ is defined by ((ref)), and by Assumption (ref) is assumed to have rank $m-r_{0}>0$. Furthermore, using $\sup_{i}\left\Vert \mathbf{C} _{i}\right\Vert <K$ and ((ref)) then

eqnarray[eqnarray omitted — 498 chars of source]

Using the above results in ((ref)) we have

equation[equation omitted — 286 chars of source]

Further $n^{-1}\sum_{i=1}^{n}\mathbf{C}_{i}\left[ \mathbf{Q}_{\bar{s}_{i} \bar{s}_{i}}-E\left( \mathbf{Q}_{\bar{s}_{i}\bar{s}_{i}}\right) \right] \mathbf{C}_{i}^{\prime }=q^{-1}\sum_{\ell =1}^{q}\mathbf{G}_{\ell }-\mathbf{G }_{0}$, where

equation*[equation* omitted — 473 chars of source]

and

equation*[equation* omitted — 381 chars of source]

Under Assumptions (ref) and (ref), $\mathbf{\bar{s}}_{i\ell }$ and $\mathbf{\bar{s}}_{i\circ }$ are cross-sectionally independent random variables and $\sup_{i}\left\Vert \mathbf{C}_{i}\right\Vert <K$. Further, $ T_{q}^{-1/2}\mathbf{\bar{s}}_{i\ell }$ and $T^{-1/2}\mathbf{\bar{s}}_{i\circ }$ are scaled partial sums of $u_{it}\,\ $and tend to bounded random variables. See, for example, result (d) of Proposition 17.1 in \citeN{Hamilton1994} . Therefore, $\mathbf{G}_{\ell }$ and $\mathbf{G}_{0}$ both converge at the rate of $n^{-1/2}$ to their means that are zero, by construction. Namely $ \mathbf{G}_{\ell }=O_{p}\left( n^{-1/2}\right) $ and $\mathbf{G} _{0}=O_{p}\left( n^{-1/2}\right) $. Hence

equation[equation omitted — 222 chars of source]

and using this result in ((ref)) yields

equation*[equation* omitted — 136 chars of source]

Therefore, for a fixed $q\left( \geq 2\right) $, $\mathbf{Q}_{\bar{w}\bar{w} }\rightarrow _{p}\frac{(q-1)}{6q}\mathbf{\Psi }$, as $n,T\rightarrow \infty $ jointly such that $T_{n}\approx n^{d}$ and $d>0$, where $\mathbf{\Psi } =lim_{n\rightarrow \infty }\mathbf{\Psi }_{n}$. This result is formally established in the following proposition.

propositionConsider the panel data model for $\mathbf{w}_{it}$ given by ( (ref)) without the interactive time effects ($\mathbf{G}_{i}=0$), and suppose Assumptions (ref) to (ref) hold. Consider the $m\times m$ pooled sample covariance matrix $\mathbf{Q}_{\bar{w}\bar{w}}$ defined by ( (ref)). Then$\ $ for $q\geq 2$ and a fixed $m$ we have \begin{equation} \mathbf{Q}_{\bar{w}\bar{w}}=\frac{(q-1)}{6q}\mathbf{\Psi }_{n}+O_{p}\left( n^{-1/2}\right) +O_{p}\left( T^{-1}\right) , \end{equation} \begin{equation} \mathbf{Q}_{\bar{w}\bar{w}}\mathbf{{\Greekmath 010C} }_{j0}=O_{p}\left( n^{-1/2}\right) +O_{p}\left( T^{-1}\right) \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, for }j=1,2,...,r_{0}, \end{equation} and $\mathbf{Q}_{\bar{w}\bar{w}}\rightarrow _{p}\frac{(q-1)}{6q}\mathbf{\Psi }$, as $n,T\rightarrow \infty $ jointly such that $T_{n}\approx n^{d}$ and $ d>0$, where $\mathbf{\Psi }$ and $\mathbf{\Psi }_{n}$ are defined in Assumption (ref).

For a known $r_{0}$ the PME estimator of $\mathbf{B}_{0}$, is given by $ \mathbf{\hat{B}}_{0}=\left( \mathbf{\hat{{\Greekmath 010C}}}_{1},\mathbf{\hat{{\Greekmath 010C}}} _{2},...,\mathbf{\hat{{\Greekmath 010C}}}_{r_{0}}\right) $, where $\mathbf{\hat{{\Greekmath 010C}}} _{j}$ , $j=1,2,...,m$ are the orthonormalized eigenvectors of $\mathbf{Q}_{ \bar{w}\bar{w}}$, as set out below equation ((ref)). Then

equation[equation omitted — 236 chars of source]

and $\mathbf{\hat{B}}_{0}^{\prime }\mathbf{Q}_{\bar{w}\bar{w}}\mathbf{\hat{B} }_{0}$ is asymptotically minimized when $\mathbf{\hat{B}}_{0}^{\prime } \mathbf{\Psi \mathbf{\hat{B}}}_{0}=\mathbf{0}$, and this occurs if $\mathbf{ \hat{B}}_{0}$ lies in the space spanned by the $r_{0}$ eigenvectors of $ \mathbf{\Psi }$ that are associated with its $r_{0}$ zero eigenvalues, namely if and only if $\mathbf{\hat{B}}=\mathbf{\mathring{B}}_{0}\mathbf{H}$ , for some $r_{0}\times r_{0}$ non-singular rotation matrix, $\mathbf{H}$. Hence, we have $\mathbf{\hat{B}}_{r}\mathbf{H}\rightarrow _{p}\mathbf{ \mathring{B}}_{0}$ as $n,T\rightarrow \infty $ jointly such that $T\approx n^{d}$ and $d>0$. A formal statement is provided in the following theorem with proofs in the Appendix.

theoremConsider the panel data model for the $m\times 1$ vector $ \mathbf{w}_{it}$ given by ((ref)) without the interactive time effects ( $\mathbf{G}_{i}=0$), and suppose that Assumptions (ref) to (ref) hold and the number of long-run relations, $r_{0}$, is known. Let $\mathbf{ \hat{B}}_{0}=\left( \mathbf{\hat{{\Greekmath 010C}}}_{1},\mathbf{\hat{{\Greekmath 010C}}}_{2},..., \mathbf{\hat{{\Greekmath 010C}}}_{r_{0}}\right) $ be the $m\times r_{0}$ matrix formed from the orthonormalized eigenvectors of $\mathbf{Q}_{\bar{w}\bar{w}}$, defined by ((ref)), associated with its $r_{0}$ smallest eigenvalues. Then for a fixed $m$ and $q$ $(\geq 2),$ $\mathbf{\hat{B}}_{0}\mathbf{ H\rightarrow }_{p}\mathbf{\mathring{B}}_{0}$ as $n,T\rightarrow \infty $ jointly such that $T_{n}\approx n^{d}$ and $d>0$, for any $r_{0}\times r_{0}$ non-singular matrix, $\mathbf{H}$.

Identification and asymptotic distribution of PME estimator

We focus on the case of exact identification of the long-run relations, noting that the estimation of $r_{0}$ is invariant on how exact identification is achieved.

Exact identifying conditions

We assume there exist $r_{0}^{2}$ a priori given and theoretically meaningful exact identifying restrictions on $\mathbf{B}_{0}$ given by

equation[equation omitted — 216 chars of source]

where $\mathbf{\mathring{B}}_{0}$ is the $m\times r_{0}$ matrix of identified long-run relations, $\mathbf{R}_{1}$,$\mathbf{R}_{2}$ and $ \mathbf{A}$ are $r_{0}\times r_{0},$ $r_{0}\times (m-r_{0})$ and $ r_{0}\times r_{0}$ matrices of known fixed constants, with $rank\left( \mathbf{A}\right) =rank\left( \mathbf{R}_{1}\right) =r_{0}<m$. \ Then it follows that $\mathbf{H}=\left( \mathbf{RB}_{0}\right) ^{-1}\mathbf{A}$, and the PME estimator of $\mathbf{\mathbf{\mathring{B}}_{0}}$ is given by

equation[equation omitted — 154 chars of source]

where $\mathbf{\hat{B}}_{0}=\left( \mathbf{\hat{{\Greekmath 010C}}}_{10},\mathbf{\hat{ {\Greekmath 010C}}}_{20},...,\mathbf{\hat{{\Greekmath 010C}}}_{r_{0},0}\right) $. The exact identifying restrictions, ((ref)), will often take the form $ \mathbf{\mathring{B}}_{0}=\left( \mathbf{I}_{r_{0}},\mathbf{\mathring{B}} _{0,2}^{\prime }\right) ^{\prime }$, where $\mathbf{\mathring{B}}_{0,1}= \mathbf{I}_{r_{0}}$\ is an identity matrix of order $r_{0}$. Without loss of generality, we consider this formulation and denote $\mathbf{\mathring{B}} _{0,2}$ by $\mathbf{\Theta }$. Once $\mathbf{\Theta }$ is estimated it is possible to estimate $\mathbf{\mathring{B}}_{0,1}$ under more general restrictions as $\mathbf{\mathring{B}}_{0,1}=\mathbf{R}_{1}^{-1}\left( \mathbf{A-R}_{2}\mathbf{\Theta }\right) .$

propositionConsider the $r_{0}^{2}$ exact identifying restrictions given by ((ref)), and suppose $m\times r$ matrix of long-run relations $ \mathbf{\mathring{B}}_{0}$ is normalized as $\mathbf{\mathring{B}} _{0}=\left( \mathbf{I}_{r_{0}},\mathbf{\Theta }^{\prime }\right) ^{\prime }$ , and Assumption (ref) holds. Partition $\mathbf{\Psi }$ conformably with $\mathbf{\mathring{B}}_{0}$ as $\mathbf{\Psi =}\left( \begin{array}{cc} \mathbf{\Psi }_{11} & \mathbf{\Psi }_{21}^{\prime } \\ \mathbf{\Psi }_{21} & \mathbf{\Psi }_{22} \end{array} \right) $. Then $\mathbf{I}_{r_{0}}+\mathbf{\Psi }_{11}$ and $\mathbf{\Psi } _{22}$ are respectively $r_{0}\times r_{0}$ and $(m-r_{0})\times (m-r_{0})$ positive definite matrices and $\mathbf{\Theta }$ is uniquely determined by $ \mathbf{\Theta =-\Psi }_{22}^{-1}\mathbf{\Psi }_{21},$ subject to the restrictions $\mathbf{\Psi }_{11}=\mathbf{\Psi }_{21}^{\prime }\mathbf{\Psi } _{22}^{-1}\mathbf{\Psi }_{21}$. A proof is provided in the Appendix.

Asymptotic distribution

To derive the asymptotic distribution of $\widehat{\mathbf{\mathring{B}}} _{0}-\mathbf{\mathring{B}}_{0}$, note from ((ref)) that

equation[equation omitted — 414 chars of source]

By Lemma (ref), $n^{-1}\sum_{i=1}^{n}T^{2}\mathbf{Q}_{\bar{v}_{i} \bar{v}_{i}}=O_{p}\left( n^{-1/2}\right) $ and since $\left\Vert \mathbf{ \mathring{B}}_{0}\right\Vert <K$, then

equation[equation omitted — 288 chars of source]

Further, write the first term as

eqnarray*[eqnarray* omitted — 487 chars of source]

and using $\sup_{i}\left\Vert E\left( \mathbf{Q}_{\bar{s}_{i}\bar{v} _{i}}\right) \right\Vert =O\left( T^{-2}\right) $ (established in Lemma (ref)))

equation[equation omitted — 154 chars of source]

Using the above results in ((ref)) we have

equation[equation omitted — 223 chars of source]

where $\mathbf{Z}_{i}=\mathbf{C}_{i}\left[ T\mathbf{Q}_{\bar{s}_{i}\bar{v} _{i}}-E\left( T\mathbf{Q}_{\bar{s}_{i}\bar{v}_{i}}\right) \right] \mathbf{ \mathring{B}}_{0}$. Recalling that $T\mathbf{Q}_{\bar{s}_{i}\bar{v} _{i}}=q^{-1}\sum_{\ell =1}^{q}\left( \mathbf{\bar{s}}_{i\ell }-\mathbf{\bar{s }}_{i\circ }\right) \left( \mathbf{\bar{v}}_{i\ell }-\mathbf{\bar{v}} _{i\circ }\right) ^{\prime }$, then $\mathbf{Z}_{i}$ can be written as

equation[equation omitted — 577 chars of source]

Writing ((ref)) in vec form

equation*[equation* omitted — 494 chars of source]

where $\mathbf{\bar{{\Greekmath 0118}}}_{iq}=q^{-1}\sum_{\ell =1}^{q}\mathbf{{\Greekmath 0118} }_{i\ell } $, and $\mathbf{{\Greekmath 0118} }_{i\ell }=vec\left[ \left( \mathbf{\bar{s}}_{i\ell }- \mathbf{\bar{s}}_{i\circ }\right) \left( \mathbf{\bar{v}}_{i\ell }-\mathbf{ \bar{v}}_{i\circ }\right) ^{\prime }\right] =\left( \mathbf{\bar{v}}_{i\ell }-\mathbf{\bar{v}}_{i\circ }\right) \mathbf{\otimes }\left( \mathbf{\bar{s}} _{i\ell }-\mathbf{\bar{s}}_{i\circ }\right) $. Under Assumption (ref), $ \mathbf{\bar{v}}_{i\ell }$, $\mathbf{\bar{s}}_{i\ell }$, $\mathbf{\bar{v}} _{i\circ }$ and $\mathbf{\bar{s}}_{i\circ }$ are distributed independently over $i$. Hence, $\mathbf{\bar{{\Greekmath 0118}}}_{iq}$ is distributed independently over $i$. In addition, $\sup_{i}\left\Vert \left( \mathbf{\mathring{B}} _{0}^{\prime }\mathbf{\otimes \mathbf{C}}_{i}\right) \right\Vert =\left\Vert \mathbf{\mathring{B}}_{0}\right\Vert \sup_{i}\left\Vert \mathbf{C} _{i}\right\Vert <K$ and using Lemma (ref), it follows that for $ (n,T)\rightarrow \infty \,,\ $jointly such that $T\approx n^{d}$ and $d>0$, we have

equation[equation omitted — 296 chars of source]

where

equation[equation omitted — 325 chars of source]

and

equation[equation omitted — 297 chars of source]

Using ((ref)) in ((ref)) yields asymptotic normality of $\widehat{ \mathbf{\mathring{B}}}$, which is formally established in the following theorem.

theoremConsider the panel data model for the $m\times 1$ vector $\mathbf{w }_{it},$ given by ((ref)) without interactive time effects ($\mathbf{G} _{i}=\mathbf{0}$). Suppose that Assumptions (ref) to (ref) hold, $ m $ and $q$ $(\geq 2)$ are fixed integers, and the number of long-run relations, $r_{0}$ ($m>r_{0}>0$) is known. Suppose further that the long-run relations, $\mathbf{\mathring{B}}_{0}$, of interest are subject to the exact identifying restrictions, $\mathbf{R}\mathbf{\mathring{B}}_{0}\mathbf{=A}$ , given by ((ref)), and consider the PME estimator of $ \mathbf{\mathring{B}}_{0}$\ given by \begin{equation*} \widehat{\mathbf{\mathbf{\mathring{B}}}}_{0}\mathbf{=\hat{B}}_{0}\left( \mathbf{R\hat{B}}_{0}\right) ^{-1}\mathbf{A,} \end{equation*} where $\mathbf{\hat{B}}_{0}=\left( \mathbf{\hat{{\Greekmath 010C}}}_{10},\mathbf{\hat{ {\Greekmath 010C}}}_{20},...,\mathbf{\hat{{\Greekmath 010C}}}_{r_{0},0}\right) $ are the first $ r_{0} $ orthonormalized eigenvectors of $\mathbf{Q}_{\bar{w}\bar{w}}$ defined by ((ref)). Then \begin{equation} \sqrt{n}T\left( \mathbf{I}_{r_{0}}\mathbf{\otimes Q}_{\bar{w}\bar{w}}\right) \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{$vec$}\left( \widehat{\mathbf{\mathring{B}}}_{0}-\mathbf{\mathring{B}} _{0}\right) \rightarrow _{d}N\left( \mathbf{0,\Omega }_{q}\right) , \end{equation} as $\left( n,T\right) \rightarrow \infty $, jointly such that $T\approx n^{d} $ for $d>1/2$, where $\mathbf{\Omega }_{q}$ is defined by ((ref)), and $\mathbf{Q}_{\bar{w}\bar{w}}\rightarrow _{p}\frac{(q-1)}{6q}\mathbf{\Psi }$. (see ((ref))). A proof is provided in the Appendix.

Theorem (ref) can be readily used to obtain asymptotic distribution of any linear combination of $\widehat{\mathbf{\mathring{B}}}_{0}$. One notable case of interest is to consider the exact identifying restrictions $\mathbf{ \mathring{B}}_{0,1}=\mathbf{I}_{r_{0}}$ discussed above. Under these restrictions $\mathbf{\mathring{B}}_{0,1}=\widehat{\mathbf{\mathring{B}}} _{0,1}=\mathbf{I}_{r_{0}}$, and $\widehat{\mathbf{\mathring{B}}}_{0}-\mathbf{ \mathring{B}}_{0}=\left(

array[array omitted — 92 chars of source]

\right) $. Partitioning $\mathbf{Q}_{\bar{w}\bar{w}}$ and $\mathbf{\Omega } _{z}$ accordingly, we have

equation*[equation* omitted — 432 chars of source]

where $\mathbf{Q}_{22,\bar{w}\bar{w}}\rightarrow _{p}\frac{(q-1)}{6q}\mathbf{ \Psi }_{22}$, and $\mathbf{\Psi }_{22}$ is the $(m-r_{0})\times (m-r_{0})$ lower right block of $\mathbf{\Psi }$ which is positive definite (see Proposition (ref)). We obtain the following corollary.

corollarySuppose assumptions of Theorem (ref) hold and consider the exact identifying restrictions $\mathbf{\mathring{B}}_{0,1}=\widehat{\mathbf{ \mathring{B}}}_{0,1}=\mathbf{I}_{r_{0}}$. Suppose $r_{0}$ is known, $q\geq 2$ , and let $\mathbf{\hat{\Theta}}$ and $\mathbf{\Theta }_{0}$ be the lower $ \left( m-r_{0}\right) \times r_{0}$ block of $\widehat{\mathbf{\mathring{B}}} _{0}$ and $\mathbf{\mathring{B}}_{0}$, respectively. Then, \begin{equation} \sqrt{n}T\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ $vec$}\left( \mathbf{\hat{\Theta}}-\mathbf{\Theta } _{0}\right) \rightarrow _{d}N\left( \mathbf{0,\Omega }_{{\Greekmath 0112} q}\right) , \end{equation} as $n,T\rightarrow \infty $, jointly such that $T\approx n^{d}$ for $d>1/2$ ,where \begin{equation} \mathbf{\Omega }_{{\Greekmath 0112} q}=\left( \frac{6q}{q-1}\right) ^{2}\left( \mathbf{I }_{r}\mathbf{\otimes \Psi }_{22}^{-1}\right) \mathbf{\Omega }_{q,22}\left( \mathbf{I}_{r}\mathbf{\otimes \Psi }_{22}^{-1}\right) \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,} \end{equation} $\mathbf{\Omega }_{q,22}$ is the $\left( m-r_{0}\right) \times \left( m-r_{0}\right) $ lower right block of $\mathbf{\Omega }_{q}$ given by ((ref)), and $\mathbf{\Psi }_{22}$ is the $(m-r_{0})\times (m-r_{0})$ lower right block of $\mathbf{\Psi }$.

Estimation of the asymptotic covariance of $\mathbf{\hat{\Theta}} $

To consistently estimate $\mathbf{\Omega }_{{\Greekmath 0112} q}$ we need a consistent estimator of $\mathbf{\Omega }_{q,22}$, since $\left[ (q-1)/6q\right] \mathbf{\Psi }_{22}$ can be consistently estimated by $\mathbf{Q}_{\bar{w} \bar{w},22}$. Consider $\mathbf{\Omega }_{q}$ given by ((ref)). Using ( (ref)) note that

eqnarray*[eqnarray* omitted — 1,081 chars of source]

Let $\mathbf{{\Greekmath 0110} }_{i\ell }=vec\left[ \left( \mathbf{\bar{w}}_{i\ell }- \mathbf{\bar{w}}_{i\circ }\right) \left( \mathbf{\bar{w}}_{i\ell }-\mathbf{ \bar{w}}_{i\circ }\right) ^{\prime }\mathbf{\mathring{B}}_{0}\right] $, and $ \mathbf{{\Greekmath 0111} }_{i\ell }=vec\left[ \left( \mathbf{\bar{v}}_{i\ell }-\mathbf{ \bar{v}}_{i\circ }\right) \left( \mathbf{\bar{v}}_{i\ell }-\mathbf{\bar{v}} _{i\circ }\right) ^{\prime }\right] $, where recall that $\mathbf{{\Greekmath 0111} } _{i\ell }=O_{p}(T^{-1})$ uniformly in $i$ and $\ell $. Then $\mathbf{{\Greekmath 0110} } _{i\ell }=\left( \mathbf{\mathring{B}}_{0}^{\prime }\mathbf{\otimes \mathbf{C }}_{i}\right) \mathbf{{\Greekmath 0118} }_{i\ell }+\left( \mathbf{\mathring{B}} _{0}^{\prime }\mathbf{\otimes I}_{m}\right) \mathbf{{\Greekmath 0111} }_{i\ell }$, and

equation*[equation* omitted — 555 chars of source]

where $\mathbf{\bar{{\Greekmath 0118}}}_{iq}=q^{-1}\sum_{\ell =1}^{q}\mathbf{{\Greekmath 0118} }_{i\ell } $, as before. It follows

equation*[equation* omitted — 406 chars of source]

and $E\left( \mathbf{\bar{{\Greekmath 0118}}}_{iq}\right) =O(T^{-1})$. Hence, as $ n,T\rightarrow \infty $ jointly such that $T\approx n^{d}$ for $d>1/2$,

equation*[equation* omitted — 604 chars of source]

Therefore, $\mathbf{\Omega }_{q}$ can be consistently estimated by the $ m^{2}\times m^{2}$ matrix

equation*[equation* omitted — 338 chars of source]

where

equation*[equation* omitted — 600 chars of source]

Let $\widehat{\mathbf{\mathring{B}}_{0}}^{\prime }\left( \mathbf{\bar{w}} _{i\ell }-\mathbf{\bar{w}}_{i\circ }\right) =\widehat{\mathbf{\bar{E}}} _{i\ell }$, the $r\times 1$ vector of error corrections for the sub-sample $ \ell $, then

equation[equation omitted — 452 chars of source]

Using the above results we now have

equation[equation omitted — 222 chars of source]

When $r_{0}=1$, estimate of the error correction term $\widehat{\mathbf{ \mathring{B}}}_{0}^{\prime }\left( \mathbf{\bar{w}}_{i\ell }-\mathbf{\bar{w}} _{i\circ }\right) =\widehat{\bar{e}}_{i\ell }$ is a scalar and we can write $ \widehat{\mathbf{\bar{{\Greekmath 0110}}}}_{iq}=q^{-1}\sum_{\ell =1}^{q}\left( \mathbf{ \bar{w}}_{i\ell }-\mathbf{\bar{w}}_{i\circ }\right) \widehat{\bar{e}}_{i\ell }$, and

eqnarray*[eqnarray* omitted — 788 chars of source]

which resembles the robust covariance matrix estimator that arises in estimation of panel data models with short $T$ and large $n$. Here $q$ plays the role of $T$.

Estimation of $r_{0}$ by eigenvalue thresholding

Under Assumption (ref), the true number of common long-run relations, $ r_{0}$, is defined by $rank(\mathbf{\Psi }_{n})=m-r_{0}>0$, where $\mathbf{ \Psi }_{n}=n^{-1}\sum_{i=1}^{n}\mathbf{C}_{i}\mathbf{\Sigma }_{i}\mathbf{C} _{i}^{\prime }$. Subject to this condition, $\mathbf{\Psi }_{n}\mathbf{{\Greekmath 010C} }_{j,0}=0$, for $j=1,2,...,r_{0}$, where $\mathbf{{\Greekmath 010C} }_{j,0}$ is the $ j^{th}$ long-run relation ($j\leq r_{0}$). The $m\times r_{0}$ matrix of long-run relations is denoted by $\mathbf{B}_{0}$. It is also worth bearing in mind that under Assumption (ref), $\mathbf{\Psi }_{n}\mathbf{{\Greekmath 010C} } _{j}\neq \mathbf{0}$, for $j=r_{0}+1,r_{0}+2,...m$, namely cannot be spanned by the $r_{0}$ columns of $\mathbf{B}_{0}$. See ((ref)) and ((ref)). There is a clear shift in the ordered eigenvalues of $\mathbf{ \Psi }_{n}$ from ${\Greekmath 0115} _{r_{0}}=$ $0$ to ${\Greekmath 0115} _{r_{0}+1}>0$, which allows us to propose a thresholding estimator of $r_{0}$ applied to the eigenvalues of $\mathbf{Q}_{\bar{w}\bar{w}}$, noting that under our assumptions $\mathbf{Q}_{\bar{w}\bar{w}}$ tends to $\left( \frac{q-1}{6q} \right) \mathbf{\Psi }$ as $n$ and $T\rightarrow \infty $. See result ((ref)) of Proposition (ref). Such an estimator can be written conveniently as

equation[equation omitted — 111 chars of source]

where $\hat{{\Greekmath 0115}}_{1}\leq \hat{{\Greekmath 0115}}_{2}\leq ....\leq \hat{{\Greekmath 0115}} _{m} $ are ordered eigenvalues of $\mathbf{Q}_{\bar{w}\bar{w}}$, and $ \mathcal{I}\left( \mathcal{A}\right) =1$ if $\mathcal{A}$ is true or zero otherwise, and $C_{T}=KT^{-{\Greekmath 010E} }$, for some ${\Greekmath 010E} >0$. This estimator is invariant to the ordering of the eigenvalues, but using the ordering helps with the exposition and the rationale behind the proofs.

To establish the consistency of $\hat{r}$ as an estimator of $r_{0}$, we first note that ((ref)) can be written equivalently as

equation[equation omitted — 211 chars of source]

which in turn yields:

equation[equation omitted — 222 chars of source]

Again noting the ordering of the eigenvalues, $\Pr \left( \hat{{\Greekmath 0115}} _{j}\geq C_{T}\right) \leq \Pr \left( \hat{{\Greekmath 0115}}_{1}\geq C_{T}\right) $, for $j=2,3,...,r_{0},$ and $\Pr \left( \hat{{\Greekmath 0115}}_{r_{0}+1}<C_{T}\right) \geq \Pr \left( \hat{{\Greekmath 0115}}_{j}<C_{T}\right) $, for $ j=r_{0}+2,r_{0}+3,...,m$. Using these results in ((ref)) we have

equation[equation omitted — 203 chars of source]

Similarly,

eqnarray[eqnarray omitted — 431 chars of source]

Also, by Markov inequality there exists ${\Greekmath 010F} >0$ such that

equation[equation omitted — 295 chars of source]

Again by Markov inequality $\Pr \left( \hat{{\Greekmath 0115}}_{1}\geq C_{T}\right) \leq C_{T}^{-1}E\left( \hat{{\Greekmath 0115}}_{1}\right) $, and using result ((ref) ) established in Lemma (ref), we have

equation[equation omitted — 128 chars of source]

Consider now $\Pr \left( \hat{{\Greekmath 0115}}_{r_{0}+1}<C_{T}\right) $, and recall that

equation[equation omitted — 467 chars of source]

with associated population values given by (see Assumption (ref)).

equation[equation omitted — 458 chars of source]

where $\mathbf{\Psi }_{n}=n^{-1}\sum_{i=1}^{n}\mathbf{C}_{i}\mathbf{\Sigma } _{i}\mathbf{C}_{i}^{\prime }\succeq 0$, and $\mathbf{{\Greekmath 010C} }_{j}^{\prime } \mathbf{{\Greekmath 010C} }_{j}=1$. Furthermore, ${\Greekmath 0115} _{j}>0$ and $\mathbf{\Psi }_{n} \mathbf{{\Greekmath 010C} }_{j}\neq 0,$ for $j>r_{0}$. But using ((ref)) and results established in Lemma (ref) and Proposition (ref) we have $ E\left( \mathbf{Q}_{\bar{w}\bar{w}}\right) =\frac{(q-1)}{6q}\mathbf{\Psi } _{n}+O\left( T^{-2}\right) $, and

equation*[equation* omitted — 370 chars of source]

Using the above results then ((ref)) can be written equivalently as $ E\left( \mathbf{Q}_{\bar{w}\bar{w}}\right) \mathbf{{\Greekmath 010C} }_{j}={\Greekmath 0115} _{j} \mathbf{{\Greekmath 010C} }_{j}+O\left( T^{-2}\right) $,$\,$\ and ${\Greekmath 0115} _{j}=\mathbf{ {\Greekmath 010C} }_{j}^{\prime }E\left( \mathbf{Q}_{\bar{w}\bar{w}}\right) \mathbf{ {\Greekmath 010C} }_{j}+O\left( T^{-2}\right) $. Therefore, together with ((ref)), we have

equation*[equation* omitted — 316 chars of source]

and rearranged as

eqnarray[eqnarray omitted — 705 chars of source]

Also, since $\mathbf{{\Greekmath 010C} }_{j}^{\prime }E\left( \mathbf{Q}_{\bar{w}\bar{w} }\right) ={\Greekmath 0115} _{j}\mathbf{{\Greekmath 010C} }_{j}^{\prime }+O\left( T^{-2}\right) ,$ then \

equation[equation omitted — 735 chars of source]

where $\mathbf{\tilde{Q}}_{\bar{w}\bar{w}}=O_{p}\left( n^{-1/2}\right) +O_{p}\left( T^{-1}\right) $. Pre-multiplying both sides of the above equations by $\sqrt{n}$ and stacking them in matrix notation we will have

equation*[equation* omitted — 720 chars of source]

where

equation*[equation* omitted — 875 chars of source]

It then follows that $\mathbf{p}_{nT}$ is of lower order as compared to$ \sqrt{n}\left( \hat{{\Greekmath 0115}}_{j}-{\Greekmath 0115} _{j}\right) $ and $\sqrt{n}\left( \mathbf{\hat{{\Greekmath 010C}}}_{j}-\mathbf{{\Greekmath 010C} }_{j}\right) $, and as $ n,T\rightarrow \infty $ jointly, such that $T\approx n^{d}$ for $d>1/2$, we have

equation*[equation* omitted — 499 chars of source]

where $\mathbf{\Omega }_{j}=\left(

array[array omitted — 205 chars of source]

\right) $. Using partitioned inverse it is easily seen that $\mathbf{\Omega } _{j}$ has an inverse if $\mathbf{\Upsilon }_{j}=E\left( \mathbf{Q}_{\bar{w} \bar{w}}\right) -{\Greekmath 0115} _{j}\mathbf{I}_{m}-2{\Greekmath 0115} _{j}\mathbf{{\Greekmath 010C} }_{j} \mathbf{{\Greekmath 010C} }_{j}^{\prime }$ is invertible. To check the invertibility of $ \mathbf{\Upsilon }_{j}$ we note that $\mathbf{\Upsilon }_{j}\mathbf{{\Greekmath 010C} } _{j}=-2{\Greekmath 0115} _{j}\mathbf{{\Greekmath 010C} }_{j}+O\left( T^{-2}\right) $, and since $ {\Greekmath 0115} _{j}>0$ for $j>r_{0}$ it then follows that $\mathbf{\Upsilon }_{j}$ must be invertible. Therefore, we can now solve for $\sqrt{n}\left( \hat{ {\Greekmath 0115}}_{j}-{\Greekmath 0115} _{j}\right) $, in terms of a linear combination of$ \sqrt{n}\mathbf{{\Greekmath 010C} }_{j}^{\prime }\mathbf{\tilde{Q}}_{\bar{w}\bar{w}} \mathbf{{\Greekmath 010C} }_{j}$ and $\sqrt{n}\mathbf{\tilde{Q}}_{\bar{w}\bar{w}}\mathbf{ {\Greekmath 010C} }_{j}$, and its asymptotic distribution can be derived accordingly. Consider the asymptotic distribution of these two terms, and note that since $\mathbf{Q}_{\bar{s}_{i}\bar{s}_{i}}=T^{-1}q^{-1}\sum_{\ell =1}^{q}\left( \mathbf{\bar{s}}_{i\ell }-\mathbf{\bar{s}}_{i\circ }\right) \left( \mathbf{ \bar{s}}_{i\ell }-\mathbf{\bar{s}}_{i\circ }\right) ^{\prime }$, then

equation*[equation* omitted — 516 chars of source]

where$\ {\Greekmath 0110} _{ij,\ell }=T^{-1/2}\mathbf{{\Greekmath 010C} }_{j}^{\prime }\mathbf{C} _{i}\left( \mathbf{\bar{s}}_{i\ell }-\mathbf{\bar{s}}_{i\circ }\right) $. Also by Minkowski inequality

equation*[equation* omitted — 500 chars of source]

Using Lemma (ref) in the supplement we have $\sup_{i,\ell }\left\Vert T^{-1/2}\mathbf{\bar{s}}_{i\ell }\right\Vert ^{4+{\Greekmath 010F} }<K$, and hence $ \left( E\left\vert {\Greekmath 0110} _{ij,\ell }\right\vert ^{4+{\Greekmath 010F} }\right) ^{1/4+{\Greekmath 010F} }<K$ and $\sqrt{n}\mathbf{{\Greekmath 010C} }_{j}^{\prime }\mathbf{\tilde{ Q}}_{\bar{w}\bar{w}}\mathbf{{\Greekmath 010C} }_{j}\rightarrow _{d}N(0,{\Greekmath 0124} _{j}^{2})$ , for some ${\Greekmath 0124} _{j}^{2}>0$. Similarly, it follows that all $m$ elements\thinspace of $\sqrt{n}\mathbf{\tilde{Q}}_{\bar{w}\bar{w}}\mathbf{ {\Greekmath 010C} }_{j}=n^{-1/2}\sum_{i=1}^{n}\mathbf{C}_{i}\left[ \mathbf{Q}_{\bar{s} _{i}\bar{s}_{i}}-E\left( \mathbf{Q}_{\bar{s}_{i}\bar{s}_{i}}\right) \right] \mathbf{C}_{i}^{\prime }\mathbf{{\Greekmath 010C} }_{j}$ are asymptotically normally distributed with zero means and finite variances. Therefore, it also follows that $\sqrt{n}\left( \hat{{\Greekmath 0115}}_{j}-{\Greekmath 0115} _{j}\right) \overset{a}{ \thicksim }N(0,{\Greekmath 0124} _{{\Greekmath 0115} _{j}}^{2})$, for some ${\Greekmath 0124} _{{\Greekmath 0115} _{j}}^{2}>0$. Using this result and setting $j=r_{0}+1$ we have

equation*[equation* omitted — 413 chars of source]

Since ${\Greekmath 0115} _{r_{0}+1}>0$, then $\Pr \left( \hat{{\Greekmath 0115}} _{r_{0}+1}<C_{T}\right) \rightarrow 0$, as $n$ $\rightarrow \infty $ if $ C_{T}<{\Greekmath 0115} _{r_{0}+1}$. Recall also from ((ref)) that $\Pr \left( \hat{{\Greekmath 0115}}_{1}\geq C_{T}\right) =O\left( C_{T}^{-1}n^{-1/2}T^{-2}\right) $ , and $\Pr \left( \hat{{\Greekmath 0115}}_{1}\geq C_{T}\right) \rightarrow 0$ if $ C_{T}^{-1}$ does not rise too fast with $T$. Setting $T=KT^{-{\Greekmath 010E} }$, and recalling that $n\approx T^{1/d}$, these two conditions on $C_{T}$ are met if $0<{\Greekmath 010E} <2+(1/2d)$. Using this result in ((ref)) and ((ref)), it follows that $\hat{r}\rightarrow _{p}r_{0}$ and $E\left( \hat{r }-r_{0}\right) ^{2}\rightarrow 0$, as $n$ and $T\rightarrow \infty $, if $ {\Greekmath 010E} $ is set close to zero such that $KT^{-{\Greekmath 010E} }<$ ${\Greekmath 0115} _{r_{0}+1}$ . \ These results also suggest that the probability of selecting too many long-run relations is more affected by $n$ than $T$, and the probability of selecting too few long-run relations tends to zero much faster with $T$ than with $n$. We need large $n$ for not selecting more than $r_{0}$ long-run relations. This latter probability, $\Pr \left( \hat{{\Greekmath 0115}} _{j}<C_{T}\right) $ for $j>r_{0}$, is also affected by the size of ${\Greekmath 0115} _{r_{0}+1}$ which measures the degree to which there is a transition from a stationary linear combination under\ which ${\Greekmath 0115} _{j}=0$ for $j\leq r_{0}$ , to ${\Greekmath 0115} _{j}>0$ for $j>r_{0}$.

Our theoretical derivations also provide some insight on how to set $K$ and $ {\Greekmath 010E} $ when choosing $C_{T}$. It is clear that ${\Greekmath 010E} $ need not be too large, so long as $K\thickapprox {\Greekmath 0115} _{r_{0}+1}$. In practice, this can be achieved approximately, by appropriate scaling of the observations, $ \mathbf{w}_{it}$, as discussed below.

remarkThe above derivations also suggest that our proposed selection/estimation procedure would be valid even if there were near stationary relations, namely if ${\Greekmath 0115} _{r_{0}+1}\thickapprox n^{-b}$ for some $b>0$. Then $\Pr \left( \hat{{\Greekmath 0115}}_{r_{0}+1}<C_{T}\right) \rightarrow 0,$ so long as $ b<1/2 $, which could be viewed as local-to-zero eigenvalue. Such cases will not be pursued in this paper, where we require ${\Greekmath 0115} _{r_{0}+1}>0$.

Allowing for interactive time effects

The model with interactive time effects is given by ((ref)), which we reproduce here for convenience:

equation[equation omitted — 231 chars of source]

where $\mathbf{f}_{t}$ is an $m_{f}\times 1$ vector of latent factors with $ \mathbf{G}_{i}$ the associated $m\times m_{f}$ matrix of factor loadings. We assume $\mathbf{f}_{t}$ is covariance stationary, and treat the factor loadings, $\mathbf{G}_{i}$,\ as nonstochastic, without placing any restrictions on them, besides being uniformly bounded. Hence, latent factors could be strong, semi-strong or weak. However, we do not allow for the possibility of unit roots in latent factors, and therefore do not consider the case of cointegration between $\mathbf{w}_{it}$ and latent factors. Specifically, we assume:

assumption(i) The $m_{f}\times 1$ vector of latent factors, $\mathbf{ f}_{t}$, is given by $\mathbf{f}_{t}=\mathbf{\Phi }_{f}(L)\mathbf{ {\Greekmath 0122} }_{ft}=\sum_{j=0}^{\infty }\mathbf{\Phi }_{f\ell }L^{\ell } \mathbf{{\Greekmath 0122} }_{f,t-\ell }$, for $t=...0,1,2,...,T$, where $\mathbf{ {\Greekmath 0122} }_{ft}$ is an $m_{f}\times 1$ vector of errors distributed independently over $t$ with $E(\mathbf{{\Greekmath 0122} }_{ft})=\mathbf{0,}$ and $ \sup_{it}E\left\Vert \mathbf{{\Greekmath 0122} }_{ft}\right\Vert ^{4+{\Greekmath 010F} }<K$ , for some ${\Greekmath 010F} >0$. $\mathbf{{\Greekmath 0122} }_{ft}$ is independently distributed of $\mathbf{u}_{it^{\prime }}$, for all $i,t,t^{\prime }$. (ii) The $m\times m_{f}$ coefficient matrices $\mathbf{\Phi }_{f\ell }$ are non-stochastic constants such that $\left\Vert \mathbf{\Phi }_{f\ell }\right\Vert <K{\Greekmath 011A} ^{\ell }$, where ${\Greekmath 011A} $ lies in the range $0<{\Greekmath 011A} <1$. (iii) $\mathbf{G}_{i}\,\ $are nonstochastic constants such that $ \sup_{i}\left\Vert \mathbf{G}_{i}\right\Vert <K$.

Under Assumption (ref), $E\left( \mathbf{G}_{i}\mathbf{f} _{t}\right) $ is time invariant, and together with Assumption (ref), $ E\left( \mathbf{w}_{it}\right) $ continues to be time invariant. Subtracting sub-sample time averages from the full sample time average, $\mathbf{\bar{w}} _{i\ell }-\mathbf{\bar{w}}_{i\circ }$, will therefore continue to remove unit-specific means.\ More specifically, under ((ref)) $\mathbf{Q}_{\bar{ w}_{i}\bar{w}_{i}}$ given by ((ref)) has the following expension

equation[equation omitted — 523 chars of source]

where $\mathbf{Q}_{\bar{f}_{i}\bar{s}_{i}}=\left( Tq\right) ^{-1}\sum_{\ell =1}^{q}\mathbf{G}_{i}\left( \mathbf{\bar{f}}_{\ell }-\mathbf{\bar{f}}_{\circ }\right) \left( \mathbf{\bar{s}}_{i\ell }-\mathbf{\bar{s}}_{i\circ }\right) ^{\prime }=\mathbf{Q}_{\bar{s}_{i}\bar{f}_{i}}^{\prime }$,

equation*[equation* omitted — 296 chars of source]

$\mathbf{Q}_{\bar{f}_{i}\bar{f}_{i}}=\left( Tq\right) ^{-1}\sum_{\ell =1}^{q} \mathbf{G}_{i}\left( \mathbf{\bar{f}}_{\ell }-\mathbf{\bar{f}}_{\circ }\right) \left( \mathbf{\bar{f}}_{\ell }-\mathbf{\bar{f}}_{\circ }\right) ^{\prime }\mathbf{G}_{i}^{\prime }$, and the terms not involving the latent factor are as before. By Lemma (ref) in the supplement we have

equation[equation omitted — 520 chars of source]

Pooling over $i$, $\mathbf{Q}_{\bar{w}\bar{w}}=n^{-1}\sum_{i=1}^{n}\mathbf{Q} _{\bar{w}_{i}\bar{w}_{i}}$, and using the above results together with those already established in ((ref)) and ((ref)) now yields

equation[equation omitted — 406 chars of source]

which is identical to results ((ref))-((ref)) in Proposition (ref) for models without interactive time effects.\ Similarly, inclusion of interactive time effects does not alter the convergence rate of the exactly identified PME estimator $\widehat{\mathbf{ \mathring{B}}}_{0}$, given by ((ref)), and its asymptotic distribution will remain correctly centered at zero. To see this consider the following expression for $\mathbf{Q}_{\bar{w}\bar{w}}\sqrt{n}T\left( \widehat{\mathbf{\mathring{B}}}_{0}-\mathbf{\mathring{B}}_{0}\right) $, which extends ((ref)) to panel data models with interactive time effects,

eqnarray*[eqnarray* omitted — 822 chars of source]

The new terms involve matrices $\mathbf{Q}_{\bar{s}_{i}\bar{f}_{i}}$, $ \mathbf{Q}_{\bar{f}_{i}\bar{f}_{i}}$ and $\mathbf{Q}_{\bar{f}_{i}\bar{v} _{i}}=\mathbf{Q}_{\bar{v}_{i}\bar{f}_{i}}^{\prime }$. Using the bounds in ( (ref)), we obtain $E\left\Vert n^{-1}\sum_{i=1}^{n}T^{2}\mathbf{ Q}_{\bar{f}_{i}\bar{f}_{i}}\right\Vert \leq n^{-1}\sum_{i=1}^{n}T^{2}E\left\Vert \mathbf{Q}_{\bar{f}_{i}\bar{f} _{i}}\right\Vert =O\left( T^{-2}\right) $, and similarly \newline $E\left\Vert n^{-1}\sum_{i=1}^{n}T^{2}\left( \mathbf{Q}_{\bar{f}_{i}\bar{v} _{i}}+\mathbf{Q}_{\bar{v}_{i}\bar{f}_{i}{}_{i}}\right) \right\Vert <K$. In addition, by Lemma (ref) recall that $n^{-1}\sum_{i=1}^{n}T^{2} \mathbf{Q}_{\bar{v}_{i}\bar{v}_{i}}=O_{p}\left( n^{-1/2}\right) $. Using these results and noting that $\left\Vert \mathbf{\mathring{B}} _{0}\right\Vert <K$, we have

equation[equation omitted — 480 chars of source]

To simplify the notations we set $\mathbf{{\Greekmath 0121} }_{it}=\mathbf{f}_{it}+ \mathbf{v}_{it}$, and note that

equation*[equation* omitted — 418 chars of source]

Then ((ref)) can be written as

equation*[equation* omitted — 302 chars of source]

Let $\mathbf{Z}_{i}^{\ast }=\mathbf{C}_{i}\left[ T\mathbf{Q}_{\bar{s}_{i} \bar{{\Greekmath 0121}}_{i}}-E\left( T\mathbf{Q}_{\bar{s}_{i}\bar{{\Greekmath 0121}}_{i}}\right) \right] \mathbf{\mathring{B}}_{0}$. Since $\mathbf{s}_{it}$ is independent of $\mathbf{f}_{t}$ and $E\left( \mathbf{s}_{it}\right) =\mathbf{0}$, we have $E\left( \mathbf{Q}_{\bar{s}_{i}\bar{f}_{i}}\right) =\mathbf{0}$, and using ((ref)) we obtain $\left\Vert n^{-1}\sum_{i=1}^{n}\mathbf{C} _{i}E\left( T^{2}\mathbf{Q}_{\bar{s}_{i}\bar{{\Greekmath 0121}}_{i}}\right) \mathbf{ {\Greekmath 010C} }_{0}\right\Vert <K$. It now follows that

equation[equation omitted — 229 chars of source]

where $\mathbf{Z}_{i}^{\ast }=\mathbf{C}_{i}\left[ T\mathbf{Q}_{\bar{s}_{i} \bar{{\Greekmath 0121}}_{i}}-E\left( T\mathbf{Q}_{\bar{s}_{i}\bar{{\Greekmath 0121}}_{i}}\right) \right] \mathbf{{\Greekmath 010C} }_{0}$. Writing ((ref)) in vec form

equation*[equation* omitted — 497 chars of source]

where $\mathbf{{\Greekmath 0118} }_{iq}^{\ast }$ is the $m^{2}\times 1$ vector given by $ \mathbf{{\Greekmath 0118} }_{iq}^{\ast }=q^{-1}\sum_{s=1}^{q}\left( \mathbf{\bar{{\Greekmath 0121}}} _{is}-\mathbf{\bar{{\Greekmath 0121}}}_{i}\right) \mathbf{\otimes }\left( \mathbf{\bar{s }}_{is}-\mathbf{\bar{s}}_{i}\right) $. Although $\mathbf{{\Greekmath 0118} }_{iq}^{\ast }$ is not independent over $i$, it is a martingale difference sequence, and \newline $n^{-1/2}\sum_{i=1}^{n}\left( \mathbf{\mathring{B}}_{0}^{\prime }\mathbf{ \otimes \mathbf{C}}_{i}\right) \left[ \mathbf{{\Greekmath 0118} }_{iq}^{\ast }-E\left( \mathbf{{\Greekmath 0118} }_{iq}^{\ast }\right) \right] $ converges to a normal distribution as $n,T\rightarrow \infty $, jointly. Therefore, $\widehat{ \mathbf{\mathring{B}}}$ will continue to be asymptotically normally distributed, with its asymptotic distribution correctly centered at zero. Even though the variance of the asymptotic distribution of PME estimator in general depends on the interactive time effects, the estimator of the asymptotic variance given by ((ref)) will continue to be consistent under Assumption (ref), and inference can be carried out in the same way as in panel data models without interactive time effects.

Overall, we find that the PME estimator is robust to interactive time effects so long as Assumption (ref) holds. Theorems and (ref) and (ref) in the supplement provide formal statements regarding consistency and the asymptotic normality of the PME estimators in presence of interactive time effects.

How to choose $q$ and $C_{T}$

To implement the estimation of $r_{0}$ and the associated long-run relations we need to decide on $q$, and $C_{T}=CT^{-{\Greekmath 010E} }$ that enter the thresholding estimation of $r_{0}$.

Choice of $q$

To ensure that $\mathbf{Q}_{\bar{w}\bar{w}}=n^{-1}T^{-1}q^{-1}\sum_{i=1}^{n} \sum_{\ell =1}^{q}\left( \mathbf{\bar{w}}_{i\ell }-\mathbf{\bar{w}}_{i\circ }\right) \left( \mathbf{\bar{w}}_{i\ell }-\mathbf{\bar{w}}_{i\circ }\right) ^{\prime }$ does not depend on the fixed effects, given by $E(\mathbf{w} _{it})=\mathbf{a}_{i}$, we need at least two sub-samples, namely $q\geq 2$. To reduce the variance of time series dependence of $\mathbf{\bar{w}}_{i\ell }-\mathbf{\bar{w}}_{i\circ }$ over the sub-samples we need $T/q$ to be reasonably large. An optimum choice of $q$ most likely depends on $T$, and not so much on $n$. To ensure that the mathematical derivations are manageable and transparent, so far we have assumed that $q$ is fixed as $ T\rightarrow \infty $. But to select $q$ we must allow $q$ to depend on $T$, denoted as $q_{T}$, and consider values of $q_{T}$ that rise with $T$, but at a slower rate such that $q_{T}/T\rightarrow 0$. Analogous to the problem of selecting the lag order in time series literature, we conjecture setting $ q_{T}$ to rise at the rate of $T^{1/3}$, and set $q_{T}$ to the lower integer part of $\max (2,T^{1/3})$. This would suggest that values of $q_{T}$ equal to $2$, $3$ and $4$ for values of $T=20,50$ and $100$, respectively, considered in our Monte Carlo simulations reported in Section (ref) below, where we consider the values of $2$ and $4$, to save space. For estimation of $r_{0}$, the choice of $q=2$ works perfectly well for all values of $T$ considered. But the higher value of $q=4$ does seem to perform slightly better than $q=2$ for estimation of long-run coefficients when $ T=100$.

Choices of $C$ and $\protect{\Greekmath 010E} $ for estimation of $r_{0}$

It is clear that eigenvalues of $\mathbf{Q}_{\bar{w}\bar{w}}$ depend on the scale of the observations, $\mathbf{w}_{it}$, and some form of scaling of data is required to reduce the sensitivity of the eigenvalues to scale. One could resort to cross validation procedures to set $C$ and ${\Greekmath 010E} $, but based on extensive Monte Carlo experiments, we have found that setting $C=1$ works well if we base our selection procedure on the eigenvalues of the following correlation matrix

equation[equation omitted — 226 chars of source]

Accordingly, our proposed estimator of $r_{0}$ is given by

equation[equation omitted — 135 chars of source]

where $\tilde{{\Greekmath 0115}}_{j}$, for $j=1,2,...,m$ are the eigenvalues of $ \mathbf{R}_{_{\bar{w}\bar{w}}}$.\footnote{ Using the correlation matrix, $\mathbf{R}_{_{\bar{w}\bar{w}}}$, has the advantage that it is unaffected by scaling, so long as the same scaling is used across all cross section units. It is not invariant if the scaling varies across units as well as across the variables. This issues is addressed in Section (ref) of the supplement where we investigate the small sample sensitivity of $\tilde{r}$ to differential scaling of the variables across the units. We find that the small sample performance of $ \tilde{r}$ as an estimator of $r_{0}$, is hardly affected by such differential scaling.}

Also, based on our theoretical derivations any value of ${\Greekmath 010E} $ close to $ zero$ should work. In the Monte Carlo experiments we consider the values of $ {\Greekmath 010E} =1/4$ and $1/2$ and conclude that ${\Greekmath 010E} =1/4$ is a good overall choice and cross-validation is not necessary for the implementation of our estimation strategy.

Monte Carlo Evidence

We investigate small sample properties of the proposed PME estimator with Monte Carlo experiments using both VARMA(1,1) and VAR(1) designs, with and without interactive time effects. The designs we consider are all special cases of the general linear model ((ref)). We set $m=3$, and generate the $3\times 1$ vector $\mathbf{w}_{it}$ as $I(1)$ variables under three scenarios: non-cointegration, $r_{0}=0$, and $r_{0}=1$ and $r_{0}=2$ cointegrating relations. We consider sample size combinations, $ T=(20,50,100) $ and $n=(50,500,1000,3000)$, and report results for $q=2$ and $q=4$ (sub-samples), which are in line with our conjecture of setting $q$ in line with the $\max (2,T^{1/3})$ rule. See Section (ref).

When $r_{0}=1$, there are a range of alternative estimators of the cointegrating relation in the literature that we can use for comparison, most of which assume the direction of long-run causality is known. To accommodate existing estimators, we distinguish between experiments based on long-run causal ordering. The PME estimator does not require the direction of long-run causality to be known. When $r_{0}>1$, to the best of our knowledge, there are no obvious alternative estimators of $r_{0}$ and the associated cointegrating relations in the panel cointegration literature that we can use. Accordingly, for the purpose of comparison, we report results using a mean group version of Johansen's maximum likelihood procedure, whereby we estimate $r_{0}$ and associated cointegrating vectors (if any), for all individual units in the panel separately, and report the frequency with which $r_{0}$ is selected by Johansen procedure across the $n$ units, and the mean group estimates of the cointegrating coefficients and their standard errors.

Subsection (ref) outlines the data generating processes (DGPs). Subsection (ref) gives the results for estimates of $r_{0}$. Our results show near perfect performance for our proposed estimator of $ r_{0}$, even for samples as small as $T=20$ and $n=50$. This is in line with the theory developed in Section (ref). Subsection (ref) reports the results for the coefficients of the long-run relations assuming $ r_{0}$ is known, which is justified considering the near perfect performance of our estimator of $r_{0}$. The MC results provide simulation evidence that the PME estimator performs well in panels with $n$ as large as $1,000$ and $ T $ as small as $20$.

Data generating processes

We consider experiments with and without long-run relations. In the experiments with long-run relations, $\mathbf{w}_{it}$ is generated as

equation[equation omitted — 161 chars of source]

for $i=1,2,...,n$, $t=1,2,...,T$, where $\mathbf{\Pi }_{i}=\mathbf{A}_{i} \mathbf{B}_{0}^{\prime },$ $\mathbf{A}_{i}$ is $m\times r_{0},$ $\mathbf{B} _{0}$ is $m\times r_{0}.$ We set $\mathbf{d}_{i}=\mathbf{\Pi }_{i}\mathbf{ {\Greekmath 0116} }_{iw}$ to ensure no linear trends in data. See, for example, Section 5.7 in \citeN{Johansen1995} .\ The elements of $\mathbf{{\Greekmath 0116} }_{iw}$ are generated as $IIDN(0,1)$ . We consider both VAR(1) and VARMA(1,1) designs. For VAR(1) we set $\mathbf{ \Theta }_{i}=\mathbf{0}$, and for VARMA(1,1) with set $\mathbf{\Theta } _{i}=diag({\Greekmath 0112} _{ij},j=1,2,...,m)$, and generate ${\Greekmath 0112} _{ij}$ as $IIDU \left[ -0.5,0.5\right] $. The errors $\mathbf{u}_{it}$, are generated following both Gaussian and chi-squared distributions.

We initially consider $m=3$ variables in $\mathbf{w}_{it}=\left( w_{it,1},w_{it,2},w_{it,3}\right) ^{\prime }$, with both one and two long-run relations. For $r_{0}=1$, we set $\mathbf{B}_{0}=\mathbf{{\Greekmath 010C} } _{1,0}=\left( 1,0,-1\right) ^{\prime }$ and $\mathbf{A} _{i}=(a_{i,11},a_{i,21},...,a_{i,31})^{\prime }$. In this case, the long-run relation is given by

equation[equation omitted — 273 chars of source]

with ${\Greekmath 010C} _{11,0}=1$, ${\Greekmath 010C} _{12,0}=0$ and ${\Greekmath 010C} _{13,0}=-1$. We identify the long-run relation by imposing ${\Greekmath 010C} _{11,0}=1$ and estimate $ {\Greekmath 010C} _{12,0}$ and ${\Greekmath 010C} _{13,0}$. When $r_{0}=2$, we set

equation*[equation* omitted — 452 chars of source]

In this case, the two long-run relations are given by

eqnarray[eqnarray omitted — 542 chars of source]

where ${\Greekmath 010C} _{11,0}={\Greekmath 010C} _{22,0}=1$, ${\Greekmath 010C} _{12,0}={\Greekmath 010C} _{21,0}=0$, and $ {\Greekmath 010C} _{13,0}={\Greekmath 010C} _{23,0}=-1$. We identify these long-run relations by imposing ${\Greekmath 010C} _{11,0}={\Greekmath 010C} _{22,0}=1$ and ${\Greekmath 010C} _{12,0}={\Greekmath 010C} _{21,0}=0$ , and we estimate ${\Greekmath 010C} _{13,0}$ and ${\Greekmath 010C} _{23,0}$.

We set the values of $\mathbf{A}_{i}$ to control the average speed of convergence towards the long-run relations. For example, in the case where $ r_{0}=1$ then $\mathbf{B}_{0}^{\prime }\mathbf{A}_{i}={\Greekmath 011A} _{i}=a_{i,11}-a_{i,31}$ and ${\Greekmath 011A} _{i}$, for $i=1,2,...,n$ are generated as $ IIDU[0.1,0.2]$ representing slow convergence, and as $IIDU[0.1,0.3]$ representing moderate convergence. We then set $a_{i,21}=0$, which leaves us with one free parameter in $\mathbf{A}_{i}$ which we use to set the system measures of the fit, $PR_{nT}^{2}=0.2$ and $0.3,$ defined as a pooled $R^{2}$ , given by equation ((ref)) in the supplement. Since $a_{i,11}$ and $ a_{i,31}$ are both nonzero, the long-run causality runs from $\left( w_{it,2},w_{it,3}\right) $ to $w_{it,1}$ as well as from $w_{it,1}$ to $ w_{it,3}$. For $r_{0}=2$ the rate of convergence will depend on the eigenvalues of $\mathbf{I}_{2}-\mathbf{B}_{0}^{\prime }\mathbf{A}_{i}$ and the details of how we generate the elements of $\mathbf{A}_{i}$ are given in the supplement.

We also consider experiments with interactive time effects, which are obtained by augmenting the solution of ((ref)) with $\mathbf{G}_{i} \mathbf{f}_{t}$, where $\mathbf{f}_{t}$ is and $m_{f}\times 1$ vector of latent factors. Each of these factors are generated as AR(1) process with a break in the AR coefficient, and the individual elements of the $m\times m_{f}$ matrix of factor loadings $\mathbf{G}_{i}$ are generated as $IIDU \left[ 0.0.4\right] $. We set $m_{f}=4$. Specifically, we augment the general linear process versions of the above VARMA(1,1) and VAR(1) specifications with $\mathbf{G}_{i}\mathbf{f}_{t}$\thinspace , namely

equation*[equation* omitted — 150 chars of source]

where $\mathbf{C}_{i}$ and $\mathbf{C}_{i}^{\ast }(L)$ are obtained from

equation*[equation* omitted — 227 chars of source]

and $\mathbf{\Psi }_{i}=\mathbf{I}_{3}-\mathbf{A}_{i}\mathbf{B}_{0}^{\prime } $. See the supplement for further details.

For the experiments with no long-run relations ($r_{0}=0$) we generate $ \Delta \mathbf{w}_{it}$ using the VAR(1) model in first-differences:

equation[equation omitted — 116 chars of source]

for $i=1,2,...,n$, $t=1,2,...,T$, where $\mathbf{u}_{it}\thicksim IIDN( \mathbf{0},\mathbf{\Sigma }_{ui})$. The elements of the covariance matrix $ \mathbf{\Sigma }_{ui}=\left( {\Greekmath 011B} _{i,\ell \ell ^{\prime }}\right) $ are generated as ${\Greekmath 011B} _{i,\ell \ell }=1$ for $i=1,2,...,n$ and $\ell =1,2,3$, and ${\Greekmath 011B} _{i,\ell \ell ^{\prime }}\thicksim IIDU(0,0.5)$, for $\ell \neq \ell ^{\prime }\,$, and $i=1,2,...,n$. We use a diagonal matrix for $\mathbf{ \Phi }_{i}=({\Greekmath 011E} _{i,\ell \ell ^{\prime }})$, with ${\Greekmath 011E} _{i,\ell \ell }$ elements on its diagonal, for $r=1,2,...,m$. We consider three options for $ {\Greekmath 011E} _{i,\ell \ell }$: ($i$) low values ${\Greekmath 011E} _{i,\ell \ell }\sim U[0,0.8]$, ($ii$) moderate values ${\Greekmath 011E} _{i,\ell \ell }\sim U[0.7,0.9]$, and ($iii$) high values ${\Greekmath 011E} _{i,\ell \ell }\sim U[0.80,0.95]$. $\mathbf{w}_{it}$ is then obtained by cumulating $\Delta \mathbf{w}_{it}$ from the initial value $ \mathbf{w}_{i,0}=0$. Similarly to the experiment with long-run relations given by ((ref)), the model ((ref)) is a special case of ((ref)). Specifically, ((ref)) leads to $\mathbf{w}_{it}=\mathbf{w} _{i0}+\mathbf{G}_{i}\mathbf{f}_{t}+\mathbf{C}_{i}\mathbf{s}_{it}+\mathbf{C} _{i}^{\ast }(L)\mathbf{u}_{it},$ where $\mathbf{C}_{i}=\left( \mathbf{I}_{m}- \mathbf{\Phi }_{i}\right) ^{-1}$, and $\mathbf{C}_{i}^{\ast }(L)=-\mathbf{ \Phi }_{i}\left( \mathbf{I}_{m}-\mathbf{\Phi }_{i}\right) ^{-1}\left( \mathbf{I}_{m}-\mathbf{\Phi }_{i}L\right) ^{-1}$. For experiments without interactive time effects we set $\mathbf{G}_{i}=\mathbf{0}$, for all $i$.

In addition to the designs described above, we also consider a data generating process taken from Section 3.1 of \citeN{ChudikPesaranSmith2021PB} . For this $m=2$, and $\mathbf{w}_{it}=\left( w_{1,it},w_{2,it}\right) ^{\prime }$ is generated as

eqnarray*[eqnarray* omitted — 264 chars of source]

where $a_{i,11}\sim IIDU\left[ 0.2,0.3\right] $, $\mathbf{u}_{it}=\left( u_{1,it},u_{2,it}\right) ^{\prime }$ is heteroskedastic and cross-sectionally independent, and $u_{1,it}$ is correlated with $u_{2,it}$. \footnote{$u_{1,it}$ $={\Greekmath 011B} _{1i}e_{1,it}$, $u_{2,it}={\Greekmath 011B} _{2i}e_{2,it}$ , ${\Greekmath 011B} _{1,i}^{2},{\Greekmath 011B} _{2,i}^{2}\sim IIDU\left[ 0.8,1.2\right] $,

equation*[equation* omitted — 535 chars of source]

See Section 3.1 of \citeN{ChudikPesaranSmith2021PB} for details.} We use this design to see how PME compares to single-equation estimators that correctly assume long-run causality runs from $w_{2,it}$ to $ w_{1,it}$. We would expect such estimators to perform reasonably well, particularly when $T$ is large relative to $n$, and provide a good baseline to evaluate the performance of PME in settings favorable to the single equation techniques advanced in the literature.

In short, we have 71 experiments. 6 experiments with $m=3$ variables and no long-run relations ($r_{0}=0$), given by 3 choices for the distribution of $ {\Greekmath 011E} _{ij}$, and interactive time effects are included or not. $64=2^{5}$ experiments feature $m=3$ variables with long-run relations, given by combinations of the choice of model (VAR(1) or VARMA(1,1)), $r_{0}=1$ or $2$ , Gaussian or chi-square error distributions, moderate or slow speed of convergence, system measure of fit $PR_{nT}^{2}$ $=0.2$ or $0.3$, and interactive time effects are included or not. In addition, we have one experiment with $m=2$ variables, $r_{0}=1$ long-run relation and one one-way long-run causality taken directly from \citeN{ChudikPesaranSmith2021PB} .

To save space, we report only summary results that are averages across a number of selected experiments, with the results for the individual experiments available from the authors upon request. Section (ref) of the supplement also provides further details of the Monte Carlo designs, how the processes are initialized, and the rationale behind the parameterization adopted. Additionally, Section (ref)\ of the supplement shows that the results for estimation of $r_{0}$ the associated long-run relations are robust to GARCH and threshold autoregressive effects.

Small sample evidence on estimation of $r_{0}$

We summarize the Monte Carlo findings for the estimation of $r_{0}$ by the eigenvalue thresholding estimator $\tilde{r}$ given by ((ref)), using $ {\Greekmath 010E} =1/4$ and $1/2$ in Tables 1-2. Table 1 reports selection frequencies of $\tilde{r}=0$, $1$, $2$, $3$ using VAR(1) based DGPs without interactive time effects, for the three cases of $r_{0}=0,1$ or $2$.\footnote{ These results are averaged over 3 VAR(1) experiments in first differences\ ($ r_{0}=0$), differeing in terms of autoregressive coefficients (low, medium and high values), and over 8 VAR(1) experiments in levels featuring $r_{0}=1$ and $2$ long-run relationships, differing in terms of the error distributions (Gaussian or chi-squared), fit (high or low), and speed of convergence towards long run (moderate or low).} The theory indicated very fast convergence of $\tilde{r}$ to $r_{0}$, and this is confirmed by the Monte Carlo simulations. The eigenvalue thresholding estimator $\tilde{r}$ correctly identifies $r_{0}$ in $100$ per cent of cases, except for the very small sample sizes considered. For $n=50$ and $T=20$, we see $5\%$ probability of $\tilde{r}$ overestimating the true number of long-run relations when the smaller exponent ${\Greekmath 010E} =1/4$ is used in experiments with $r_{0}=0$ in Table1. For comparison, we also report selection frequency for the Johansen procedure using the trace statistic and the conventional nominal level of 5 percent.\footnote{ We assume the true lag order of the VAR design is known.} Whereas there is a single estimate for the whole panel per replication using $\tilde{r},$ there are $n$ such estimates per replication using the Johansen procedure ($\hat{r} _{i}$, $i=1,2,...,n$). We simply use them all in calculating the selection frequency.\footnote{ There are a number of ways that one might choose $r$ for the panel from these $n$ tests, based on the modal selection or the average value of the test statistic for instance. We do not explore these avenues here.} Hence $n$ will not influence these results, but increasing $T$ does improve the frequency with which the correct value is selected using the trace statistic. For $T=100$, frequency of correctly estimating the number of long-run relations using the Johansen trace statistics is $69$ percent for $ r_{0}=0$, $82$ percent for $r_{0}=1$, and only $31$ percent for $r_{0}=2$ in Table 1.

Simulation results for the performance of $\tilde{r}$ as an estimator of $ r_{0}$ in the case of experiments with interactive time effects are presented in Section (ref) of the supplement, to save space. These results continue to show near perfect performance of $\tilde{r}$ similarly to the experiments summarized above for the case of panels without interactive time effects.

Overall, our findings are in line with the theoretical insights of a very fast convergence of $\tilde{r}$ to $r_{0}$. For this reason, we report next on the small sample performance of the identified PME estimator assuming $ r_{0}$ is known.

center[center omitted — 6,447 chars of source]
flushleft\singlespacing \scriptsize \singlespacing Notes: For $r_{0}=0,$ (panel A) there are 3 experiments: with low, medium and high serial correlation in first-difference VAR(1)\ model. For $r_{0}=1,$ (panel B) and $r_{0}=2$ (panel C) there are 8 VAR(1) experiments differing in terms of the error distributions (Gaussian or chi-square), fit (high or low), and speed of convergence towards long run (moderate or low). Lag order for the computation of Johansen's trace statistics is set equal to the true lag order for the VAR. All experiments based on $R=2,000$ MC replications. Results for individual Monte Carlo experiments are available from the authors upon request. A summary of the different MC designs is given in Subsection (ref), with a detailed account of the data generating processes provided in Section (ref) of the supplement.
center[center omitted — 4,593 chars of source]
flushleft\singlespacing \scriptsize Notes: For $r_{0}=1,$ (panel A) and $r_{0}=2$ (panel B) there are 8 VARMA(1,1) experiments differing in terms of the error distributions (Gaussian or chi-squared), fit (high or low), and speed of convergence towards long run (moderate or low). Lag order for the computation of Johansen's trace statistics is set equal to the integer part of $T^{-1/3}$. All experiments are based on $R=2,000$ MC replications.\ Results for individual Monte Carlo experiments are available from the authors upon request. A summary of the different MC designs is given in Subsection (ref), with a detailed account of the data generating processes provided in Section (ref) of the supplement.

\onehalfspacing

center[center omitted — 5,655 chars of source]
flushleft\scriptsize \singlespacing Notes: The long-run relations are given by $\mathbf{{\Greekmath 010C} }_{1,0}^{\prime } \mathbf{w}_{it}=w_{it,1}-w_{it,3}$ and $\mathbf{{\Greekmath 010C} }_{2,0}^{\prime } \mathbf{w}_{it}=w_{it,2}-w_{it,3}$, and identified using ${\Greekmath 010C} _{11,0}={\Greekmath 010C} _{22,0}=1$, and ${\Greekmath 010C} _{12,0}={\Greekmath 010C} _{21,0}=0$. The 8 experiments differ with respect to distribution (Gaussian or chi-squared), fit (high or low), and speed of convergence toward long run (moderate or low). Size and power are computed at the five percent nominal level. Reported results are based on $R=2,000$ Monte Carlo replications. Simulated power are computed under $H_{1}:$ ${\Greekmath 010C} _{13}=-0.97$, ${\Greekmath 010C} _{23}=-0.97$, as alternatives to $-1$ for both coefficients under the null. Results for individual Monte Carlo experiments are available from the authors upon request. A summary of the different MC designs is given in Subsection (ref), with a detailed account of the data generating processes provided in Section (ref) of the supplement.

\onehalfspacing

Estimation of coefficients of long-run relations

Comparison of PME with MG-Johansen in VAR(1) designs with multiple long-run relations

We first investigate how PME estimators of multiple long-run relations compare with the Mean Group estimator based on Johansen's estimator of individual cointegrating vectors (MG-Johansen). To this end, we focus on the VAR(1) design with multiple long-run relations ($r_{0}=2$ long-run relations among $m=3$ variables) and no interactive time effects. We assume the true lag order is known in the case of MG-Johansen. This is a set up which we believe is most favorable to Johansen procedure when applied to the individual units. But we note that the PME estimator does not require the knowledge of the true lag order. These long-run relations are identified according to ((ref))-((ref)) for both PME, and MG-Johansen's estimators. While moments of Johansen's estimator do not exist, we expect the simulated size and power of MG-Johansen to be good for a sufficiently large $T$ relative to $n$.

Table 3 provides a summary of the results for estimation of ${\Greekmath 010C} _{13,0}$ (\thinspace $=$ $-1)$ in part A, and for ${\Greekmath 010C} _{23,0}$ ($=-1$) in part B. The table report bias, root mean square error (RMSE), size of the tests at 5% nominal level, and power of the tests ($H_{1}:$ ${\Greekmath 010C} _{13}=-0.97$, $ {\Greekmath 010C} _{23,0}=-0.97$). All entries in this table are multiplied by $100$, and averaged across $8$ experiments, defined in terms of error distributions (Gaussian or chi-squared), fit (high or low), and speed of convergence towards the long run (moderate or low).

We first note that the PME estimator works quite well, with relatively small bias and RMSE that decline with $n$ and $T$. The choice of $q=2$ works well for most $n$ and $T$ combinations, and $q=4$ performs better in terms of RMSE only when $T=100$. In terms of size, PME with $q=2$ does better than with $q=4$ for all $n$ and $T$ combinations, and has size close to the nominal $5$ per cent level, together with satisfactory power that rapidly rises to $100$ per cent as $n$ and $T$ are increased. However, even when $ q=2 $ there is some evidence of moderate size distortions, particularly when $T=20$ and $n=3,000$. The size distortion of PME when $q=4$ seems to be largely due to its bias that does not decline with $T$, despite the fall we observe in its RMSE.

The results for the MG-Johansen estimator, also reported in Table 3, show size below 5 percent and a very low power in comparison to PME. The very large RMSE and bias entries could be due to lack of moments for the unit-specific estimators obtained using the Johansen procedure. Clearly, further research is needed for adapting the use of Johansen maximum likelihood approach for use with large panels. We are using MG-Johansen estimators here only for the purpose of comparisons when $r_{0}>1$. Below we do consider other panel cointegration procedures when $r_{0}=1$, and we will no longer report results for MG-Johansen as they are very similar to those summarized in Table 3.

The above summary results do not depend much on which of the two coefficients is considered.

center[center omitted — 3,586 chars of source]
flushleft\scriptsize \singlespacing Notes: The long-run relations are given by $\mathbf{{\Greekmath 010C} }_{1,0}^{\prime } \mathbf{w}_{it}=w_{it,1}-w_{it,3}$ and $\mathbf{{\Greekmath 010C} }_{2,0}^{\prime } \mathbf{w}_{it}=w_{it,2}-w_{it,3}$, and identified using ${\Greekmath 010C} _{11,0}={\Greekmath 010C} _{22,0}=1$, and ${\Greekmath 010C} _{12,0}={\Greekmath 010C} _{21,0}=0$. The 8 experiments differ with respect to distribution (Gaussian or chi-squared), fit (high or low), and speed of convergence toward long run (moderate or low). Reported results are based on $R=2,000$ Monte Carlo replications. Simulated power are computed under $H_{1}:$ ${\Greekmath 010C} _{13}=-0.97$ and ${\Greekmath 010C} _{23}=-0.97$, as alternatives to $-1$ for both coefficients under the null. Results for individual Monte Carlo experiments are available from the authors upon request. A summary of the different MC designs is given in Subsection (ref), with a detailed account of the data generating processes provided in Section (ref) of the supplement. Size and Power are computed at 5 percent nominal level.
center[center omitted — 3,586 chars of source]
flushleft\scriptsize \singlespacing Notes: See the notes to Table 4.
center[center omitted — 7,483 chars of source]
flushleft\scriptsize \singlespacing Notes: The long-run relation is given by $\mathbf{{\Greekmath 010C} }_{1,0}^{\prime } \mathbf{w}_{it}=w_{it,1}-w_{it,2}={\Greekmath 010C} _{11,0}w_{it,1}+{\Greekmath 010C} _{12,0}w_{it,2} $, and identified with ${\Greekmath 010C} _{11,0}=1$. Coefficient ${\Greekmath 010C} _{12,0}=-1$ is estimated. Data generating process used for results reported in this table is taken from \citeN{ChudikPesaranSmith2021PB} . Section 3.1 of \citeN{ChudikPesaranSmith2021PB} provides full account of this design. Reported results are based on $R=2,000$ Monte Carlo replications. Size and Power are computed at 5 percent nominal level. Simulated powers are computed under $H_{1}:{\Greekmath 010C} _{12}=-0.97$, compared to null value of $-1$.

\onehalfspacing

Performance of PME in VARMA(1,1) designs

Results when using VARMA(1,1) designs with $r_{0}=2$ are presented in Tables 4 for models without time effects, and in Table 5 for models with interactive time effects. As before, these results are averages across eight VARMA(1,1) experiments featuring $r_{0}=2$ long-run relations that differ in terms of error distributions (Gaussian or chi-squared), fit (high or low), and speed of convergence toward long run (moderate or low).\ Figures 1-2 give empirical power curves for tests on ${\Greekmath 010C} _{13,0}$ and ${\Greekmath 010C} _{23,0}$ using PME estimators computed with $q=2$ sub-samples, in VARMA(1,1) designs with $r_{0}=2$ long-run relations, and for $n=50,$ $500,$ and $1000.$ Separate panels show $T=20,$ $T=50,$ $T=100.$ This experiment has a slow speed of convergence, $PR_{nT}^{2}=0.2,$ Gaussian errors and without interactive time effects (Figure 1) or with interactive time effects (Figure 2). In both cases, the power increases with $n$ and $T$ as expected.

Qualitatively, the VARMA results in Table 4 are similar to the VAR results summarized in Table 3, with one important exception. Allowing for an MA component in the DGP tends to reduce power of the tests. For example, in the case where $n=500$, $T=20$, and $q=2$, the power of testing ${\Greekmath 010C} _{13,0}=-1 $ against the alternative ${\Greekmath 010C} _{13}=-0.97$ is $67.73$ per cent when using VAR design (Table 3) as compared to $29.62$ per cent when using the VARMA design (Table 4). Otherwise, the results for bias, RMSE and size are comparable across the two designs. Similarly, allowing for interactive time effects (Table 5) does not alter these conclusions, and PME seems to be quite robust to the inclusion of interactive time effects, in line with the theoretical results of Section (ref), so long as the time effects are not trended (deterministic or stochastic).

Comparisons with available estimators in the design with single long-run relation and one-way long-run causality

Last but not least, we investigate how PME compares with existing approaches for the estimation of a single long-run relation, in a design that is favorable to single equation estimators, some of which rely on the direction of long-run causality to be one-way and known, and the lag order of the VAR to be well specified. We report results for panel FMOLS estimator by \citeANP{Pedroni1996} (\citeyearNP{Pedroni1996}, \citeyearNP{Pedroni2001}, \citeyearNP{Pedroni2001ReStat}) , the Pooled Mean Group (PMG) estimator by \citeN{PesaranShinSmith1999} , panel Dynamic OLS (PDOLS) by \citeN{MarkSul2003} , the two-step system estimator of \citeN{Breitung2005} , the system PMG estimator of \citeN{ChudikPesaranSmith2023} , and the pooled Bewley estimator by \citeN{ChudikPesaranSmith2021PB} .\ For large $T$ panels with moderate $n$, we would expect these single equation techniques to perform better than the PME that allows for MA components and does not assume long-run causality. This is confirmed by the results summarized in Table 6. When $T=100$ and $n=50$, PME ($q=2$) has higher RMSE than all other estimators reported in Table 6 except for the FMOLS estimator. In contrast, PME estimator ($q=2$) is much more balanced in terms of bias, RMSE and size of the tests, when $T$ is small ($=20$), and $n$ quite large ($=1000$). None of the alternative estimators to PME in the case of $r_{0}=1$ (and known one-way long-run causality) work well in samples where $T$ is not very large and $n$ much larger than $T$.

In terms of bias and size, the PME estimator with $q=2$ performs much better than all the other single equation estimators under consideration, even if we consider $T=100$ and $n=50$. Overall, Monte Carlo findings show that the PME estimator with $q=2$ can have satisfactory performance for panels where $ n$ is quite large relative to $T$, in particular for panels with $n$ and $T$ combinations similar to the ones we consider in our empirical application discussed below.

Empirical Applications

We provide two empirical applications to illustrate wide applicability of the PME approach to both micro and macro panels.

Estimation of long-run financial relations

The fist application considers micro panels of firms where $n$ is quite large relative to the available time dimension. We use logarithms of six key financial variables from CRSP/Compustat, available from Wharton Research Data Services. The six variables (measured in logarithms) are book value ($ BV_{it}$), market value ($MV_{it}$), short-term debt ($SD_{it}$), long-term debt ($LD_{it}$), total assets ($TA_{it}$) and total debt outstanding ($ DO_{it}$). These are variables among which one would expect some key relations. We consider them in three sets of two or three variables, using two unbalanced panels both begin in 1950, the shorter one ends in 2010 and the longer one in 2021. The maximum $T$ is $71$. We set a minimum $T$ of $20$ , and the average $T$ is around 30. The number of firms $n$ varies from about $1,000$ to $2,500$ depending on the set of variables under consideration. As well as estimating the number of long-run relations and their parameters we test whether the coefficients in the linear combinations of logarithms take the value -1, to relate our results to the ratios used in corporate finance. In corporate finance accounting ratios, constructed from balance sheet data, are commonly used to measure the profitability, liquidity, and solvency of a firm. The rationales for the use of ratios include correcting for size in the cross section dimension and eliminating common trends in the time series dimension to render the ratios stationary. These two objectives are not always compatible. In an early contribution, that remains relevant, \citeN{LevSunder1979} comment \textquotedblleft It appears that the extensive use of financial ratios by both practitioners and researchers is often motivated by tradition and convenience rather than resulting from theoretical considerations or from a careful statistical analysis.\textquotedblright\ \shortciteN{Geelen_etal_2024} provide a more recent study of the use of financial ratios in empirical corporate finance literature.

Given that the theory is often not very specific, it is desirable to have a statistical criteria to judge which are the appropriate long-run relations among the set of variables considered and whether the logarithm of their ratios is stationary. The time series stationarity of finance ratios like the aggregate dividend price ratio have been studied, but the question has not, to our knowledge, been addressed in corporate finance, where the context is somewhat different. Corporate finance studies tend to use unbalanced panels with large $n$ and relatively small $T$ and and the vector of accounting variables, of the sort one gets from Compustat, may include multiple long-run relations. Thus the PME estimator which is appropriate for multiple long-run relations in a large $n$, moderate $T$ panel seems well designed to determine whether there are long-run relations in accounting data.

Using a multiplicative specification, the relation between $y_{it},$ $x_{it}$ and $z_{it}$ can be readily cast in terms of $\mathbf{w}_{it}=(\ln y_{it},\ln x_{it},\ln z_{it})^{\prime }$ and the PME procedure can be used to test (a) if $\mathbf{{\Greekmath 010C} }_{0}^{\prime }\mathbf{w}_{it}$ is stationary; and (b) if ${\Greekmath 010C} _{11,0}=1,$ ${\Greekmath 010C} _{12,0}=-1,$ ${\Greekmath 010C} _{13,0}=0;$ to validate the the use of $\ln \left( y_{it}/x_{it}\right) $ or $y_{it}/x_{it}$ in econometric analysis. If step (a) cannot be validated then $\mathbf{{\Greekmath 010C} }_{0}^{\prime }\mathbf{w}_{it}$ will not be stationary and its use in econometric analysis could lead to spurious results. But if step (a) is validated but not (b), whilst it would not be advisable to use $\ln y_{it}-\ln x_{it}$, one can still consider using $\mathbf{\tilde{{\Greekmath 010C}}} ^{\prime }\mathbf{w}_{it}$, where $\mathbf{\tilde{{\Greekmath 010C}}}$ could be the exactly identified PME estimator of $\mathbf{{\Greekmath 010C} }_{0}$, in second stage panel regressions that allow for short term dynamics as well as other stationary variables.\footnote{ To simplify the notations, we use the symbol tilde in this section to denote PME estimates of the exactly identified long-run relations.} Such a two-step procedure is justified, noting that $\mathbf{\tilde{{\Greekmath 010C}}}$ is super consistent, converging to $\mathbf{{\Greekmath 010C} }_{0}$ at the rate of $T\sqrt{n}$.

Data sources, full definitions of the variables, summary statistics and additional details on construction of the samples for each of the three variable sets are provided in Section (ref) of the supplement. The following data filters are applied sequentially to each variable set and sample period, separately. Firms are omitted if, for a given variable set and sample: they do not have data for all variables in the set (without gaps and covering at least 20 time periods); have nonpositive entries on any of the variables, since we are using logarithms; and average value of key ratios fall below the 1st or above the 99th percentiles estimated after the application of the first two filters. This is similar to the filtering \shortciteN{Geelen_etal_2024} use, except that we require a longer time series dimension.\footnote{ \shortciteN{Geelen_etal_2024} state: \textquotedblleft We winsorize all variables at the 1% and 99% levels to mitigate the impact of outliers. We drop all observations with missing values on one or more variables of interest. We remove observations with a market-to-book ratio larger than 20, negative book equity or negative EBITDA. Our final sample consists of 68,833 firm-year observations with 6,001 unique firms.\textquotedblright\ This gives $\bar{T}=11.5,$ though some regressions use less.}

The variables are grouped into three sets, where we have prior expectations about possible cointegration amongst them. To illustrate the procedure we start with the simplest case where $m=2$ and $\mathbf{w}_{it}$ include the logarithm of total debt outstanding and logarithm of total assets: \{$ DO_{it} $, $TA_{it}$\}. The ratio of total debt to total assets is often used as a measure of leverage, so there is a single hypothesized long-run relation. To showcase the performance of the PME procedure with multiple long-run relations, we consider two other sets of variables with $m=3$. They are: the logarithms of short and long term debt and total assets, \{$SD_{it}$ , $LD_{it}$, $TA_{it}$\}; and the logarithms of total debt outstanding, book value and market value, \{$DO_{it}$, $BV_{it}$, $MV_{it}$\}. We expect two hypothesized long-run relations in both of these sets. Since these variables are often used to construct accounting ratios, whether the long-run relation has a unit coefficient is also of interest. Because of missing firm observations on some variables, the number of firms with minimum of $20$ data points on all the variables under consideration falls as we include more variables in a set. For this reason we do not consider all six variables together.

Table 7\ gives the estimates of the number of long-run relations for two and three variable models for the 1950-2021 and 1950-2010 unbalanced samples using the eigenvalue thresholding\ procedure, given by ((ref)). It also reports the eigenvalues of the correlation matrix, $ \mathbf{R}_{ww}$ defined by ((ref)), together with the associated threshold values, $T_{ave}^{-{\Greekmath 010E} }$, where $T_{ave}=n^{-1} \sum_{i=1}^{n}T_{i}$. Since the panel is unbalanced we base the thresholds on $T_{ave}$ and provide $\tilde{r}$ for ${\Greekmath 010E} =1/2$ and ${\Greekmath 010E} =1/4$, using $q=2$ sub-sample time averages.

The preferred threshold based on ${\Greekmath 010E} =1/4$ gives $\tilde{r}=1$ long-run relation for the panel data models with $m=2,$ and $\tilde{r}=2$ long-run relations for the two cases with $m=3$. The threshold with ${\Greekmath 010E} =1/2$ also yields $\tilde{r}=1$ in the case with $m=2$, but when used in the case of panels with $m=3$ it selects $\tilde{r}=1$ rather than $\tilde{r}=2$. Given the theoretical discussion above and the Monte Carlo results, we proceed using the estimates of the number of long-run relations obtained using the preferred value of ${\Greekmath 010E} =1/4$, namely $one$ long-run relation when $m=2$ and $two$ long-run relations when $m=3$.\footnote{ Table S10 in the supplement reports IPS panel unit root tests by \shortciteN{ImPesaranShin2003} (using a 40-year balanced panel), which do not reject the null of unit root in all cases except for short-term debt (SD). Given the $5$ per cent chance that the IPS test could be in error, we proceed assuming that all the six variables are $I(1)$.}

Tables 8-10 present PME estimates of the coefficients in the long-run relations, their standard errors and $t$-statistics for testing the null hypothesis that the long-run coefficient in question is equal to $-1$. The first set of estimates, in Table 8, are for panels with $\mathbf{w} _{it}=(DO_{it},TA_{it})^{\prime }$. Recall that $DO_{it}$ and $TA_{it}$ are logarithms of debt outstanding and total assets. There is a single hypothesized long-run relation: ${\Greekmath 010C} _{11,0}DO_{it}+{\Greekmath 010C} _{12,0}TA_{it}$. Panel A of Table 8 uses the exact identifying condition ${\Greekmath 010C} _{11,0}=1$ and provides PME estimates of ${\Greekmath 010C} _{12,0}$ and t-statistics for the null value of ${\Greekmath 010C} _{12,0}=-1$. \ The PME estimates of ${\Greekmath 010C} _{12,0}$ at $ -1.142$ and $-1.113$ for the two sample periods ending in 2021 and 2010, respectively, are similar and close to but significantly different from $-1$ . To illustrate that, unlike regression based methods, the PME estimator is invariant to normalization, part B of Table 8 uses the exact identifying condition ${\Greekmath 010C} _{12,0}=1$, and reports the PME estimate of ${\Greekmath 010C} _{11,0}$ . It is confirmed that up to rounding error $\tilde{{\Greekmath 010C}}_{11}=1/\tilde{ {\Greekmath 010C}}_{12}$.

center[center omitted — 3,964 chars of source]
flushleft\scriptsize \singlespacing Notes: This table reports eigenvalues $\tilde{{\Greekmath 0115}}_{j}$ for $ j=1,2...,m\, $\ (in ascending order) of $\mathbf{R}_{_{\bar{w}\bar{w}}}$ given by ((ref)) using $q=2$ sub-sample time averages and the corresponding eigenvalue thresholding estimates of the number of long-run relations given by $\tilde{r}=\sum_{j=1}^{m}\mathcal{I}\left( \tilde{{\Greekmath 0115}} _{j}<T_{ave}^{-{\Greekmath 010E} }\right) $, for ${\Greekmath 010E} =1/4,1/2,$ where $ T_{ave}=n^{-1}\sum_{i=1}^{n}T_{i}$, and $T_{i}\geq 20$ for all $i$. The elements of $\mathbf{w}_{it}$ are logarithms of book value ($BV_{it}$), market value ($MV_{it}$), short-term debt ($SD_{it}$), long-term debt ($ LD_{it}$), total assets ($TA_{it}$) and total debt outstanding ($DO_{it}$). See Section (ref) of the supplement for variable definitions, data sources and availability, and filters applied.

\onehalfspacing

For panels with $m=3$ and $\tilde{r}=2$ we need two exactly identifying conditions on each of the two long-run relations. In the case where $\mathbf{ w}_{it}=(SD_{it},LD_{it},TA_{it})^{\prime }$, we use the conditions ${\Greekmath 010C} _{11,0}=1$ and ${\Greekmath 010C} _{12,0}=0$ to exactly identify the first long-run relation, and conditions ${\Greekmath 010C} _{21,0}=0$ and ${\Greekmath 010C} _{22,0}=1$ to identify the second long-run relation. Hence the two identified long-run relations are $SD_{it}+{\Greekmath 010C} _{13,0}TA_{it}$ and $LD_{it}+{\Greekmath 010C} _{23,0}TA_{it}$. PME estimates for ${\Greekmath 010C} _{13,0}$ in Table 9 are -1.025 and -1.024 for the two unbalanced samples, both very close to -1 and statistically not different from -1. PME estimates for ${\Greekmath 010C} _{23,0}$ are -0.875 and -0.899, both statistically different from -1 at the one percent level. Rotation of the two identified long-run relations reported in Table 9 imply the long-run relations $SD_{it}-1.156LD_{it}$ and $SD_{it}-1.130LD_{it}$ for the samples ending in 2021 and 2010, respectively. Both of these PME estimates, $-1.156$ and $-1.130$, are statistically significantly different from $-1$ at the one per cent level. Hence, in the case of the variable set $\left\{ SD_{it},LD_{it},TA_{it}\right\} $, the unit long-run elasticity hypothesis could be accepted for short term debt to total assets only..

center[center omitted — 1,756 chars of source]
flushleft\scriptsize \singlespacing Notes: $TA_{it}$ and $DO_{it}$ are logarithms of total assets and total debt outstanding. Left panel reports PME estimate of ${\Greekmath 010C} _{12,0}$ using the exact identifying condition ${\Greekmath 010C} _{11,0}=1$. Right panel reports PME estimate of ${\Greekmath 010C} _{11,0}$ using the exact identifying condition ${\Greekmath 010C} _{12,0}=1$. PME estimator, given by ((ref)), is computed using $ q=2$\ sub-sample time averages, subject to $T_{i}\geq 20$, for all $i$. To simplify the notations we use the tilde symbols to denote the exactly identified PME estimates. The row labelled t(${\Greekmath 010C} _{12,0}=-1)$ and t($ {\Greekmath 010C} _{11,0}=-1)$ gives the t statistic for testing $H_{0}:{\Greekmath 010C} _{12,0}=-1$ and $H_{0}:{\Greekmath 010C} _{11,0}=-1$, respectively. See Section (ref) of the supplement for variable definitions, data sources and availability, and filters applied.
center[center omitted — 1,616 chars of source]
flushleft\scriptsize \singlespacing Notes: Long run relations are $\mathbf{{\Greekmath 010C} }_{j,0}^{\prime }\mathbf{w} _{it}={\Greekmath 010C} _{j1,0}SD_{it}+{\Greekmath 010C} _{j2,0}LD_{it}+{\Greekmath 010C} _{j3,0}TA_{it}$, for $ j=1,2$, where $SD_{it}$, $LD_{it}$ and $TA_{it}$ are logarithms of short-term and long-term debts, and total assets. The first long-run relation $(j=1)$ is identified using the exact identifying conditions ${\Greekmath 010C} _{11,0}=1$ and ${\Greekmath 010C} _{12,0}=0$. The second long-run relation ($j=2$) is identified using the exact identifying conditions ${\Greekmath 010C} _{22,0}=1$ and $ {\Greekmath 010C} _{21,0}=0$. This table reports estimates for ${\Greekmath 010C} _{13,0}$ and $ {\Greekmath 010C} _{23,0}$ using PME estimator given by ((ref)) with $q=2$ sub-sample time averages, subject to $T_{i}\geq 20$, for all $i$. To simplify the notations we use tilde symbol to denote the exactly identified PME estimates. See Section (ref) of the supplement for variable definitions, data sources and availability, and filters applied.
center[center omitted — 1,452 chars of source]
flushleft\scriptsize \singlespacing Notes: Long run relations are $\mathbf{{\Greekmath 010C} }_{j,0}^{\prime }\mathbf{w} _{it}={\Greekmath 010C} _{j1,0}DO_{it}+{\Greekmath 010C} _{j2,0}BV_{it}+{\Greekmath 010C} _{j3,0}MV_{it}$, for $ j=1,2$, where $DO_{it}$, $BV_{it}$ and $MV_{it}$ are logarithms of total debt outstanding, book, and market values. The first long-run relation ($j=1$ ) is identified using the exact identifying conditions ${\Greekmath 010C} _{11,0}=1$ and ${\Greekmath 010C} _{12,0}=0$. The second long-run relation ($j=2$) is identified using the exact identifying conditions ${\Greekmath 010C} _{22,0}=1$ and ${\Greekmath 010C} _{21,0}=0$. See also the notes to Table 9.

\onehalfspacing

Table 10 reports similar results for $\mathbf{w}_{it}=(DO_{it},$ $BV_{it},$ $ MV_{it})^{\prime }$ and two exactly identified long-run relations associated with logarithms total debt, market value, and book value. PME estimates are close to -1, but the unit long-run elasticity between pairs of variables in this set is rejected for all cases except the logarithms of total debt outstanding and market value in the sample ending in 2010.\footnote{ Rotation of the two identified long-run relationships reported in Table 10 imply the long-run relations $DO_{it}-0.888BV_{it}$ and $ DO_{it}-0.937BV_{it} $ for the samples ending in 2021 and 2010, respectively, both of these PME estimates are statistically significanlty different from -1 at the 1 percent level.}

Overall, the results support the existence of long-run relations between the logarithms of a number of key financial variables considered in the corporate finance literature. But with the notable exception of the logarithm of short term debt to total asset ratio, the use of other financial ratios as stationary variables in financial analysis is not supported by our empirical findings. Instead our study recommends using estimated long-run relations such as $DO_{i,t-1}-1.14$ $TA_{i,t-1}$, $ LD_{i,t-1}-1.19$ $TA_{i,t-1}$, and $BV_{i,t-1}-0.927MV_{i,t-1}$, as error correction terms in dynamic panel regressors that allow for short term dynamics and other stationary variables. This allows the analysis of short term dynamics of financial variables to be coherently embedded in a long-run equilibrating framework.

International macro applications using Penn World Table

Our second empirical application considers cross country macroeconomic time series data from the Penn World Table\footnote{ We use version 10.01 of PWT database, available at https://www.rug.nl/ggdc/productivity/pwt/, see \shortciteN{FeenstraInklaarTimmer2015} . See also the supplement for a detailed description of data constructions.} (PWT), where $n$ (the number of countries) is smaller and the average time dimension is larger as compared to the corporate finance data. Given the good small sample performance of the PME approach even for smaller values of $n$ in our Monte Carlo experiments, we have confidence in using the PME estimation also in the case of the macro application. We focus on four key macro variables, namely real merchandise exports per capita ($ex_{it}$), real merchandise imports per capita ($im_{it}$), real productivity per hour worked ($prod_{it}$) and real wages per hour worked ($wage_{it}$). The choice of these variables was motivated by two widely maintained hypotheses. Firstly, real wages and productivity should balance for steady state growth to be feasible. Secondly export and imports should balance for international solvency, though the constraint may not be binding for reserve-currency countries such as the US. We first consider the two pairs that we expect to cointegrate separately, then we consider all four variables together. As with the corporate finance application, we are interested both in whether they cointegrate and whether there is a unit coefficient. In both cases, labour market balance and trade balance, there is no clear causal ordering between the variables. Since the constraints that produce balance may be somewhat different for advanced and emerging economies, we report estimates for the country groupings separately as well as together.

The PME results for the analysis of a possible long-run relation between $ ex_{it}$ and $im_{it}$ are summarized in Table 11. This table gives the two eigenvalues of $\mathbf{R}_{_{\bar{w}\bar{w}}}$ for $\mathbf{w}_{it}=\left( ex_{it},im_{it}\right) ^{\prime }$ and $q=2$. (See ((ref))) As can be seen there is a clear separation between the first and second eigenvalues supporting the existence of a long-run relation between $ex_{it}$ and $ im_{it}$. This result holds for both choices of the threshold parameter $ {\Greekmath 010E} =1/2$ and $1/4$, and for all three country groupings. Given this result we then estimated the long-run relation $\mathbf{{\Greekmath 010C} }^{\prime } \mathbf{w}_{it}=$ ${\Greekmath 010C} _{11}ex_{it}+{\Greekmath 010C} _{12}im_{it}$, normalizing on $ {\Greekmath 010C} _{12}=1$, for all three country groupings. There is a marked difference between the estimates of ${\Greekmath 010C} _{11}$ across the advanced and emerging economies. For the advanced economies we have $\hat{{\Greekmath 010C}} _{11}=-0.914$ ($0.029$), which strongly rejects the null hypothesis of $-1$ and suggests systematic differences can persist between merchandise imports and exports for advanced economies. In contrast the estimate for the emerging economies, namely $\hat{{\Greekmath 010C}}_{11}=-0.992$ ( $0.045$), is not significantly different from $-1$. For all countries pooled, $\hat{{\Greekmath 010C}} _{11}=$ $-0.972$ ($0.034$) and does not reject the null of ${\Greekmath 010C} _{11}=-1$. \footnote{ Normalizing on ${\Greekmath 010C} _{11}$ yields the same results for ${\Greekmath 010C} _{12}$ since PME estimate of ${\Greekmath 010C} _{12}$ is exactly equal to $1/\hat{{\Greekmath 010C}}_{11}$ and by construction $Var(\hat{{\Greekmath 010C}}_{12})=Var(1/\hat{{\Greekmath 010C}}_{11})$.} The difference in the estimates obtained for advanced and emerging economies could be due to greater ability of advanced economies in financing their goods trade imbalances by increasing their export of services and having easier access to international capital market.

Similar results are obtained when we consider the relation between wages and labour productivity. For all country groupings and both choices of the thresholds we find $\tilde{r}=1$. See Table 12. The PME estimates of the long-run relation, $\mathbf{{\Greekmath 010C} }^{\prime }\mathbf{w}_{it}=$ ${\Greekmath 010C} _{11}wage_{it}+prod_{it}$ for advanced and emerging economies are $\hat{{\Greekmath 010C} }_{11}=-0.953$ ($0.013$) and $\hat{{\Greekmath 010C}}_{11}=-0.984$ ($0.043$), respectively. Once again estimates of ${\Greekmath 010C} _{11}$ are quite close to $-1$, but as in the case of imports and exports the null of ${\Greekmath 010C} _{11}=-1$ is rejected for advanced economies but not for emerging economies.

center[center omitted — 1,802 chars of source]
flushleft\scriptsize \singlespacing Notes: Standard errors are reported in parentheses. See Section (ref) of the supplement for variable definitions, data sources and availability, and filters applied.
center[center omitted — 2,225 chars of source]
flushleft\scriptsize \singlespacing Notes: See notes to Table 11.

\onehalfspacing

To illustrate how our proposed methods perform in the case of multiple long-run relations we now consider all the four variables together and set $ \mathbf{w}_{it}=\left( ex_{it},im_{it},prod_{it},wage_{it}\right) ^{\prime }$ . Given the above pair-wise results we would expect at least two long-run relations amongst these four variables. But the four eigenvalues of $\mathbf{ R}_{_{\bar{w}\bar{w}}}$ reported in Table 13 clearly suggest that there are three long-run relations among the four variables, irrespective whether we use ${\Greekmath 010E} =1/4$ or $1/2$ as the threshold parameter. This result highlights the advantage of considering the possibility of multiple long-run relations and can help with discovery of hitherto unnoticed or overlooked long relations. For the third long-run relation we consider a possible long-run relation between exports and productivity. There is extensive microeconomic evidence suggesting exporting firms tend to have higher productivity. See, for example, \citeN{BernardJensen2004} . However, there is controversy about whether the relation arises because high productivity firms export more or whether the competitive pressure of exporting boosts productivity. See for example \shortciteN{Aghion_etal2018} .\ Also, while the firm level micro relation has been examined for many countries, the country level macro relation has been less intensively explored. Our application provides cross country evidence on the relation between exports and productivity without making any assumption about the direction of causality between these variables. Accordingly, we consider the following exactly identified long-run relations assuming that $r_{0}=3$ amongst the four variables $\mathbf{w}_{it}$,

equation[equation omitted — 383 chars of source]

PME estimates for ${\Greekmath 010C} _{11}$ and ${\Greekmath 010C} _{23}$ in Table 13 are in line with the estimates in Tables 11 and 12. For the third long-run relation, we obtain $\hat{{\Greekmath 010C}}_{31}=-0.509$ ($0.023$) for advanced economies and $\hat{ {\Greekmath 010C}}_{31}=-0.426$ ( $0.032$) for emerging economies. Hence, we discovered that exports and productivity are related with the expected sign. Unlike the other two long-run relations, we do not have any a priori reason to believe the estimates of ${\Greekmath 010C} _{31}$ should be close to $-1$.

To corroborate the evidence in Table 13 regarding the third long-run relation, Table 14 reports PME findings for $\mathbf{w}_{it}=\left( prod_{it},ex_{it}\right) ^{\prime }$. There is a clear separation of eigenvalues, indicating existence of a long-run relation, in line with findings in Table 14 for the four-variable vector $\mathbf{w}_{it}$. In addition, the estimated coefficients in Table 14 are very similar to the corresponding coefficients reported in Table 13.

center[center omitted — 2,938 chars of source]
flushleft\scriptsize \singlespacing Notes: See Section (ref) of the supplement for variable definitions, data sources and availability, and filters applied.
center[center omitted — 2,199 chars of source]
flushleft\scriptsize \singlespacing Notes: See notes to Table 11.
landscape\begin{center} \singlespacing TABLE 15: Comparison of PME estimates with alternative estimates of single long-run relation, $\mathbf{w}_{it}=\left( w_{1it},w_{2,t}\right) ^{\prime }$ {2pt} \scriptsize \begin{tabular}{rcccrcccrcccrcccrcccrccc} \hline\hline $w_{1it}:$ & \multicolumn{7}{c}{$ex_{it}$} & & \multicolumn{7}{c}{$ prod_{it} $} & & \multicolumn{7}{c}{$ex_{it}$} \\ $w_{2,t}:$ & \multicolumn{7}{c}{$im_{it}$} & & \multicolumn{7}{c}{$ wage_{it} $} & & \multicolumn{7}{c}{$prod_{it}$} \\ \cline{2-8}\cline{10-16}\cline{18-24} Coint. relation: & \multicolumn{3}{c}{${\Greekmath 010C} _{11}w_{1it}+w_{2,t}$} & & \multicolumn{3}{c}{$w_{1it}+{\Greekmath 010C} _{12}w_{2,t}$} & & \multicolumn{3}{c}{$ {\Greekmath 010C} _{11}w_{1it}+w_{2,t}$} & & \multicolumn{3}{c}{$w_{1it}+{\Greekmath 010C} _{12}w_{2,t}$} & & \multicolumn{3}{c}{${\Greekmath 010C} _{11}w_{1it}+w_{2,t}$} & & \multicolumn{3}{c}{$w_{1it}+{\Greekmath 010C} _{12}w_{2,t}$} \\ & \multicolumn{3}{c}{$\hat{{\Greekmath 010C}}_{11}$} & & \multicolumn{3}{c}{$\hat{{\Greekmath 010C}} _{12}$} & & \multicolumn{3}{c}{$\hat{{\Greekmath 010C}}_{11}$} & & \multicolumn{3}{c}{$ \hat{{\Greekmath 010C}}_{12}$} & & \multicolumn{3}{c}{$\hat{{\Greekmath 010C}}_{11}$} & & \multicolumn{3}{c}{$\hat{{\Greekmath 010C}}_{12}$} \\ \cline{2-4}\cline{6-8}\cline{10-12}\cline{14-16}\cline{18-20}\cline{22-24} Sample: & Adv. & Eme. & All & \multicolumn{1}{c} & Adv. & Eme. & All & \multicolumn{1}{c} & Adv. & Eme. & All & \multicolumn{1}{c} & Adv. & Eme. & All & \multicolumn{1}{c} & Adv. & Eme. & All & \multicolumn{1}{c} & Adv. & Eme. & All \\ \hline PME & -0.914 & -0.992 & -0.972 & \multicolumn{1}{c} & -1.094 & -1.008 & -1.029 & \multicolumn{1}{c} & -0.953 & -0.984 & -0.962 & \multicolumn{1}{c} & -1.049 & -1.016 & -1.039 & \multicolumn{1}{c} & -0.510 & -0.357 & -0.432 & \multicolumn{1}{c} & -1.962 & -2.803 & -2.315 \\ & (0.029) & (0.045) & (0.034) & \multicolumn{1}{c} & (0.035) & (0.045) & (0.036) & \multicolumn{1}{c} & (0.013) & (0.043) & (0.016) & \multicolumn{1}{c} & (0.015) & (0.056) & (0.021) & \multicolumn{1}{c} & (0.023) & (0.044) & (0.036) & \multicolumn{1}{c} & (0.075) & (0.251) & (0.119) \\ SPMG & -0.962 & -1.026 & -0.976 & \multicolumn{1}{c} & -1.040 & -0.974 & -1.025 & \multicolumn{1}{c} & -1.072 & -0.987 & -1.043 & \multicolumn{1}{c} & -0.933 & -1.013 & -0.959 & \multicolumn{1}{c} & -0.367 & -0.704 & -0.371 & \multicolumn{1}{c} & -2.729 & -1.421 & -2.697 \\ & (0.004) & (0.006) & (0.004) & \multicolumn{1}{c} & (0.004) & (0.006) & (0.004) & \multicolumn{1}{c} & (0.005) & (0.004) & (0.003) & \multicolumn{1}{c} & (0.005) & (0.004) & (0.003) & \multicolumn{1}{c} & (0.004) & (0.013) & (0.003) & \multicolumn{1}{c} & (0.027) & (0.033) & (0.024) \\ Breitung's 2-step & -0.842 & -0.924 & -0.902 & \multicolumn{1}{c} & -1.048 & -0.882 & -0.922 & \multicolumn{1}{c} & -1.144 & -1.010 & -1.066 & \multicolumn{1}{c} & -0.832 & -0.924 & -0.900 & \multicolumn{1}{c} & -0.522 & -0.371 & -0.455 & \multicolumn{1}{c} & -1.701 & -1.901 & -1.756 \\ & (0.010) & (0.021) & (0.016) & \multicolumn{1}{c} & (0.019) & (0.016) & (0.013) & \multicolumn{1}{c} & (0.008) & (0.020) & (0.005) & \multicolumn{1}{c} & (0.005) & (0.018) & (0.004) & \multicolumn{1}{c} & (0.013) & (0.013) & (0.010) & \multicolumn{1}{c} & (0.042) & (0.095) & (0.040) \\ PMG & -0.985 & -0.996 & -0.989 & \multicolumn{1}{c} & -0.973 & -0.949 & -0.960 & \multicolumn{1}{c} & -1.140 & -0.990 & -1.100 & \multicolumn{1}{c} & -0.837 & -0.962 & -0.886 & \multicolumn{1}{c} & -0.291 & -0.718 & -0.306 & \multicolumn{1}{c} & -1.568 & -1.350 & -1.527 \\ & (0.007) & (0.008) & (0.005) & \multicolumn{1}{c} & (0.009) & (0.009) & (0.006) & \multicolumn{1}{c} & (0.008) & (0.007) & (0.005) & \multicolumn{1}{c} & (0.006) & (0.009) & (0.004) & \multicolumn{1}{c} & (0.007) & (0.020) & (0.006) & \multicolumn{1}{c} & (0.026) & (0.077) & (0.024) \\ PB & -0.868 & -0.813 & -0.828 & \multicolumn{1}{c} & -0.993 & -0.869 & -0.891 & \multicolumn{1}{c} & -1.136 & -0.966 & -1.097 & \multicolumn{1}{c} & -0.844 & -0.961 & -0.888 & \multicolumn{1}{c} & -0.486 & -0.379 & -0.437 & \multicolumn{1}{c} & -1.688 & -1.967 & -1.771 \\ & (0.027) & (0.033) & (0.025) & \multicolumn{1}{c} & (0.039) & (0.037) & (0.031) & \multicolumn{1}{c} & (0.006) & (0.035) & (0.002) & \multicolumn{1}{c} & (0.005) & (0.031) & (0.002) & \multicolumn{1}{c} & (0.028) & (0.056) & (0.038) & \multicolumn{1}{c} & (0.066) & (0.293) & (0.103) \\ PFMOLS & -0.857 & -0.835 & -0.841 & \multicolumn{1}{c} & -1.083 & -0.886 & -0.934 & \multicolumn{1}{c} & -1.241 & -0.978 & -1.070 & \multicolumn{1}{c} & -0.776 & -0.944 & -0.896 & \multicolumn{1}{c} & -0.528 & -0.329 & -0.445 & \multicolumn{1}{c} & -1.723 & -2.051 & -1.822 \\ & (0.023) & (0.035) & (0.026) & \multicolumn{1}{c} & (0.050) & (0.043) & (0.032) & \multicolumn{1}{c} & (0.015) & (0.485) & (0.009) & \multicolumn{1}{c} & (0.010) & (0.163) & (0.008) & \multicolumn{1}{c} & (0.064) & (0.060) & (0.092) & \multicolumn{1}{c} & (0.059) & (0.191) & (0.081) \\ PDOLS & -0.912 & -0.841 & -0.856 & \multicolumn{1}{c} & -1.045 & -0.801 & -0.853 & \multicolumn{1}{c} & -1.093 & -1.029 & -1.089 & \multicolumn{1}{c} & -0.868 & -1.029 & -0.891 & \multicolumn{1}{c} & -0.520 & -0.409 & -0.470 & \multicolumn{1}{c} & -1.853 & -1.957 & -1.901 \\ & (0.004) & (0.006) & (0.004) & \multicolumn{1}{c} & (0.005) & (0.006) & (0.004) & \multicolumn{1}{c} & (0.006) & (0.004) & (0.003) & \multicolumn{1}{c} & (0.004) & (0.004) & (0.003) & \multicolumn{1}{c} & (0.004) & (0.008) & (0.004) & \multicolumn{1}{c} & (0.015) & (0.038) & (0.015) \\ \hline \multicolumn{1}{l}{Sample size} & & & & \multicolumn{1}{c} & & & & \multicolumn{1}{c} & & & & \multicolumn{1}{c} & & & & \multicolumn{1}{c} & & & & \multicolumn{1}{c} & & & \\ \hline $n$ & 38 & 139 & 177 & \multicolumn{1}{c} & 38 & 139 & 177 & \multicolumn{1}{c} & 35 & 24 & 59 & \multicolumn{1}{c} & 35 & 24 & 59 & \multicolumn{1}{c} & 35 & 29 & 64 & \multicolumn{1}{c} & 35 & 29 & 64 \\ $\bar{T}=n^{-1}\sum_{i=1}^{n}T_{i}$ & 61.3 & 56.1 & 57.2 & \multicolumn{1}{c} & 61.3 & 56.1 & 57.2 & \multicolumn{1}{c} & 55.5 & 47.4 & 52.2 & \multicolumn{1}{c} & 55.5 & 47.4 & 52.2 & \multicolumn{1}{c} & 55.5 & 47.0 & 51.7 & \multicolumn{1}{c} & 55.5 & 47.0 & 51.7 \\ \hline\hline \end{tabular} \end{center} \begin{flushleft} \scriptsize \singlespacing Notes: PME and SPMG are the only estimators where $\hat{{\Greekmath 010C}}_{11}=\hat{ {\Greekmath 010C}}_{12}^{-1}$. This is not the case for the remaining estimators. PME is the pooled minimum eigenvalue estimator using $q=2$ subsamples. SPMG\ is the system PMG estimator of \citeN{ChudikPesaranSmith2023} , using $p=2$ lags in levels ($1$ lag in first differences). Breitung's 2-step is the two-step system estimator of \citeN{Breitung2005} , using $p=2$ lags in levels. PMG\ is the Pooled Mean Group (PMG) estimator by \citeN{PesaranShinSmith1999} , using $p=2$ lags in levels. PB is the pooled Bewley estimator by \citeN{ChudikPesaranSmith2021PB} , using $p=2$ lags in levels. PFMOLS is the panel FMOLS estimator by \citeANP{Pedroni1996} (\citeyearNP{Pedroni1996}, \citeyearNP{Pedroni2001}, \citeyearNP{Pedroni2001ReStat}) . PDOLS is the panel Dynamic OLS (PDOLS) by \citeN{MarkSul2003} , using one lead and one lag of first differences. \end{flushleft}

\onehalfspacing

Comparison of PME and their estimates in the case of pair-wise relations

Here we provide comparative estimates for the pair-wise long-run relations for which a number of alternative estimators are proposed in the literature. Specifically, we compare PME estimates reported in Tables 11, 12, and 14 with the SPMG, Breitung's 2-step, PMG, PB, FMOLS, and PDOLS estimators discussed in Sub-Section (ref).\footnote{ The SPMG estimator is proposed by \citeN{ChudikPesaranSmith2023} , the Breitung's two-step system estimator is proposed by \citeN{Breitung2005} , the PMG\ estimator is proposed by \citeN{PesaranShinSmith1999} , the PB (Pooled Bewley) estimator is proposed by \citeN{ChudikPesaranSmith2021PB} , the panel FMOLS\ estimator is proposed by \citeANP{Pedroni1996} (\citeyearNP{Pedroni1996}, \citeyearNP{Pedroni2001}, \citeyearNP{Pedroni2001ReStat}) , and the panel DOLS estimator is proposed by \citeN{MarkSul2003} .}\ The results are summarized in Table 15. PME and SPMG are the only estimators that are invariant to the normalization imposed on $\mathbf{ {\Greekmath 010C} }^{\prime }\mathbf{w}_{it}={\Greekmath 010C} _{11}w_{1,it}+{\Greekmath 010C} _{12}w_{2,it}$. As a result, the equality $\hat{{\Greekmath 010C}}_{11}\hat{{\Greekmath 010C}}_{12}=1$ holds only in the case of PME and SPMG estimators.

The estimates of ${\Greekmath 010C} _{11}$ and ${\Greekmath 010C} _{12}$ are generally similar but there are some large differences. For instance, the PME estimate of the long-run coefficient in ${\Greekmath 010C} _{11}ex_{it}+prod_{it}$ in the case of advanced economies is $-0.510$ ($0.023$) compared with SPMG estimate of $ -0.367$ ($0.004$). Such difference could arise from the possibility that some of the SPMG modeling assumptions regarding short-run dynamics are not met, whereas the PME estimator is more general in that it does not require modeling of short-run dynamics. We also observe that standard errors of the estimates proposed in the literature are in some cases 2 to 10-fold smaller compared to PME. As we have seen in Monte Carlo experiments all of the estimators proposed in the literature severely over-reject the null, whilst this is not the case if we consider the MC results for size reported in Table 6.

Concluding remarks

This paper provides a new, pooled minimum eigenvalue (PME), methodology for the analysis of multiple long-run relations in panel data models where the cross section dimension, $n$, is large relative to the time series dimension, $T$. \ It uses non-overlapping sub-sample time averages as deviations from their full-sample counterpart and estimates the number of long-run relations and their coefficients using eigenvalues and eigenvectors of the pooled covariance matrix of these sub-sample deviations. It applies to unbalanced panels generated from general linear processes with interactive stationary time effects and does not require knowing long-run causal linkages. The PME estimator is consistent and asymptotically normally distributed as $n$ and $T$ $\rightarrow \infty $ jointly, such that $ T\approx n^{d}$, with $d>0$ for consistency and $d>1/2$ for asymptotic normality. Extensive Monte Carlo studies show that the number of long-run relations can be estimated with high precision and the PME estimates of the long-run coefficients show small bias and $RMSE$ and have good size and power properties. The utility of our approach is illustrated with both micro and macro applications. The micro application uncovers long-run relations among key financial variables in an unbalanced panel of US firms from merged CRSP-Compustat data set covering $2,000$ plus firms over the period $ 1950-2021$. The macro application uses cross country macroeconomic time series data covering up to $177$ countries over slightly shorter time period, $1950-2019.$

As well as application in other areas than corporate finance and macroeconomics the procedure opens up a range of theoretical developments. One question that we are considering investigating is whether it is possible to develop a unit root test based on similar principles. Our method is semi-parametric, in that it does not model the short run dynamics. Because the estimates of the long-run relations are super consistent, then having estimated them, the PME estimates of the long-run relations could be used as inputs into second stage models using stationary variables. For instance, they can provide estimates of the disequilibrium terms in error correction models of the short run dynamics, to measure speeds of adjustment. Such models could also be used for medium-term forecasting and counterfactual analysis.