Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
61,224 characters · 12 sections · 0 citation commands
Limit Theorems for Factor Models
\thispagestyle{empty} \setcounter{page}{0} \baselineskip=18.0pt
Data with an underlying factor structure are increasingly used in empirical macroeconomics and finance. Often these data consist of time series of observations for multiple cross-sectional units (assets, portfolios, regions or industries). Quite a few new estimation strategies have appeared in the empirical literature that use both cross-sectional and time series variation in order to estimate global structural parameters. Often the parameter of interest arises from aggregation or estimation using cross-sectional variation of individual parameters for each entity. One example of such a structure is linear factor pricing model in asset pricing (Fama & MacBeth 1973 and Shanken 1992), where for estimation we usually use time series of excess returns for a number of portfolios or assets priced by a small number of risk factors. Each portfolio or stock may have its own (heterogeneous) exposure to risk, often referred to as betas, which can be estimated separately from time series observations for each portfolio. The parameter of interest, a risk premium, is defined as the coefficient of proportionality in the cross-sectional relation between the average excess return on a portfolio and its individual beta.
A vast majority of macroeconomic shocks are only weakly identified via structural VARs that use only time series observations on leading macro variables. A new approach to the estimation of causal effects of a macro shock on the economy is to use cross-sectional variation in data on regions, countries or industries. For example, Serrato & Wingender (2016) use cross-sectional variation in federal spending programs due to a Census shock to identify the causal impact of government spending on the economy. Cross-sectional variation among counties in government spending and in the accuracy of census-based estimates of population provides a better justified treatment effect framework, allows for the estimation of local fiscal multipliers, and finally gives a better global estimate of the fiscal multiplier via aggregation of local multipliers. Hagedorn et al. (2015) estimate the aggregate effect of unemployment-benefit duration on employment and labor force participation using cross-sectional differences across US states. Sarto (2018) discusses how heterogeneous sensitivities of regions to aggregate policy variables, so called micro-global elasticities, can be used to recover macro elasticities of interest such as, for example, a fiscal multiplier.
A shared feature of the above-mentioned examples is the use of time-series observations on multiple entities (stocks, portfolios, counties, states or industries), while data on those entities are not independent and identically distributed. Moreover, variables for different entities often display strong co-movements to the extent that the data have a factor structure, and estimation of these co-movements is the main goal. Indeed, the realization of a risk factor in the economy moves returns on all portfolios simultaneously, while a federal fiscal shock moves spending in all US counties, though in both cases heterogeneously so. A valid estimation procedure must explicitly model and account for the data's factor structure to the extent that the error terms (or residuals) can be considered idiosyncratic; see Kleibergen & Zhan (2015) and Anatolyev & Mikusheva (2018) for how a factor structure that is unaccounted for can lead to misleading results. However, idiosyncrasy of the errors usually implies only that the correlation among errors for different entities is relatively small and does not introduce first-order bias to the estimation procedure. Usually, it is not reasonable to assume that errors for different entities are completely independent; indeed, stocks in the same industry are likely to co-move even after global-economy risks are removed, while errors for neighboring counties are more likely to be correlated even after one accounts for federal shocks. At the same time, we typically want to remain agnostic about the correlation structure of shocks and avoid their structural modeling as long as this does not introduce biases.
The second typical feature of the above-mentioned examples is the two-step nature of the estimation procedure, where in the first step we estimate entity-specific coefficients (risk exposures/betas, local fiscal multipliers, micro-global elasticities) by running a time-series regression separately for each entity. In the second step, we estimate the global coefficient of interest by either aggregating entity-specific coefficients (Serrato & Wingender 2016 and Hagedorn et al. 2015), or by running an OLS regression on the cross-section of entity-specific coefficients (Fama & MacBeth 1973 and Sarto 2018), or by running an IV regression on the cross-section of entity-specific coefficients (Anatolyev & Mikusheva 2018).
The goal of this paper is to establish central limit theorems (CLTs) and to provide a tool for establishing asymptotic normality of estimates obtained in such two-step estimation procedures and for finding ways to do asymptotically correct inference, while being flexible in modeling the cross-sectional dependence of errors. The main difficulty here is that even though the second step cross-sectional regression has nearly uncorrelated errors (which is usually sufficient to obtain consistency of the two-step estimator), this condition is usually insufficient for a CLT, which typically requires that stronger discipline be imposed on the dependence structure (such as independence, or a martingale difference structure, or mixing). Our solution to this problem is to restrict the time series behavior while staying agnostic about the cross-sectional dependence. We assume time-series independence of idiosyncratic errors, which is consistent with market efficiency for factor asset pricing models and the non-predictability of macro shocks in macroeconomic settings. The estimation noise in a two-stage procedure involves aggregation both over time (from the first step) and over entities (from the second step). We show that under certain conditions it is sufficient to have a CLT over just one of these directions, and we use the time-series direction for that.
When the second step uses an OLS or IV estimator, the CLT must adapt to averages of quadratic forms, as both the second-step-dependent variable and the second stage regressor/instrument contain first-stage estimation noise. Our CLT has a linear and a quadratic part. We also note that a need for a CLT for quadratic forms in factor models sometimes arises for the first-step estimators (e.g., Pesaran & Yamagata 2018) or in higher order asymptotic derivations (e.g., Bai & Ng 2010).
There is a growing literature that establishes different CLTs while acknowledging the importance of cross-sectional dependence in the data, which stems from spatial relations and/or from the presence of common factors. Kuersteiner & Prucha (2013) establish a CLT for linear sums in a panel data context with growing cross-sectional dimension $N$ and fixed time-series dimension $T$ allowing for cross-sectional dependence, and Kuersteiner & Prucha (2020) extend these results to quadratic forms as well. Both papers impose conditional moment restrictions, which allows the authors to construct a martingale difference sequence in the cross-sectional direction. The main conditional moment restrictions imply a correct specification of an underlying model, which need not be required by our CLT. However, the mentioned papers allow more flexibility in modeling the time dependence, and do not require large $T$. Another CLT that requires both large $N$ and large $T$ is established in Hahn et al. (2020) for linear terms only.
This paper also contributes to the literature on the CLT for quadratic forms. Various types of CLTs for quadratic forms have been previously established and used in the many instrument literature (see for example, Chao et al. 2012, Hausman et al. 2012, S\o lvsten 2020) and many covariate literature (see Cattaneo et al. 2018), as well as in the literature on semi-parametric estimation (Cattaneo et al. 2014a, 2014b). The CLT used in those papers are established for the cross-sectional dimension only, and rely heavily on the independence assumption. We adapt the ideas used in Chao et al. (2012), specifically the approach of de Jong (1987), to accommodate large cross-sectionally dependent panels; an alternative approach, known as Stein's method, is used in S\o lvsten (2020).
Our second set of results is related to ways of conducting valid statistical inference. Under strengthened conditions on the weakness of the cross-sectional correlation of errors, we show that a conventional variance estimator is consistent, and so the usual asymptotic inference can be applied. When such strengthened conditions do not hold, we propose instead a variant of a wild bootstrap scheme that replicates the original cross-sectional dependence structure. We also conduct a small simulation experiment that provides evidence on the approximation quality of our CLT and on the empirical size and power of wild bootstrap in a moderately sized panel.
The paper proceeds as follows. Section (ref) explains problems with establishing asymptotic Gaussianity for two-step and other estimators and test statistics, and shows how discipline in the time series direction can help. Section (ref) introduces assumptions on idiosyncratic errors, states central limit theorems for two cases, and discusses the relevance of those cases to empirical practice. Section (ref) discusses estimation of asymptotic variances for asymptotic inference and alternative inference tools based on the bootstrap. Section (ref) presents a small simulation experiment that reveals properties of asymptotic and proposed bootstrap inference tools. Section 6 concludes. All proofs appear in the Appendix.
Let the data contain observations on many units indexed by $i=1,...,N,$ and observed for multiple time periods $t=1,...,T.$ We assume that both $N$ and $T$ increase to infinity without restrictions on their rates. The goal of this paper is to find the conditions under which the following statement will hold:
where
and $\Sigma _{\xi }$ is an asymptotic variance matrix. Here, $e_{it}$ are weakly cross-sectionally dependent entity-specific (idiosyncratic)\footnote{ By idiosyncratic error we mean the factor-removed part of entity-specific variables.} errors with $\mathbb{E}\left( e_{it}\right) =0$. Errors $e_{it}$ are uncorrelated with the variables $v_{t}$ and $w_{st}$ that are common to all units $i=1,...,N$ (more exact conditions are to appear in the next Section). We assume $\gamma _{i},$ $i=1,...,N,$ to be non-random entity-specific weights. Further, we want to study the circumstances when one can also consistently estimate the asymptotic covariance -- that is, sufficient conditions for a statement like
As we argue below (see Examples 1--3), statements ((ref)) and ((ref)) are often needed in order to conduct statistical inferences (testing or confidence set construction) about a structural parameter, $ \lambda ,$ which is estimated in two steps. We consider a case when in the first step a researcher estimates a parameter $\beta _{i}$ for each entity/unit/state $i=1,...,N$, typically via running OLS or IV time series regressions. A typical linear estimator can be written as $\widehat{\beta } _{i}=\beta _{i}+\varepsilon _{i},$ where the estimation error has the structure $\varepsilon _{i}=\big( 1+o_{p}(1)\big) \frac{1}{T} \sum_{t=1}^{T}v_{t}e_{it}$, with the $o_{p}(1)$ term uniformly small over the units. In this setting, $v_{t}$ is either a regressor common to all entities, or a common systematic part of entity-specific regressors that have a factor structure.\footnote{ Our setting can accommodate entity-specific regressors, say $v_{it}$, that have a factor structure themselves. Assume that $v_{it}=a_{i}u_{t}+u_{it},$ where $u_{t}$ is a common co-movement in the regressors and $u_{it}$ is idiosyncratic. Then
where $v_{t}=(u_{t},1)^{\prime }$ and $e_{it}^{\ast }=(a_{i}e_{it},u_{it}e_{it})$.}
\paragraph{Example 1.}
There is a variety of estimation approaches that can be used at the second step. The simplest of them is weighted averaging of the first step estimates, viz. $\widehat{\lambda }=\frac{1}{N}\sum_{i=1}^{N}\gamma _{i} \widehat{\beta }_{i}.$ Such an estimator is used in Sarto (2018). In order to justify asymptotic Gaussianity of $\widehat{\lambda }$ and to make statistical inferences about $\lambda ,$ one needs statements on the asymptotic behavior of
Note that the last expression has the structure of normalized averages stated as the first component of $\Xi _{N,T}$ from equation ((ref)). Such `linear' terms, where only the first component of $\Xi _{N,T}$ is involved, are very common in asymptotic derivations in factor models (e.g., Bai & Ng 2006, 2010).
\paragraph{Example 2.}
The second estimation step may invoke a more complex estimator involving a sample covariance between multiple first stage estimators or estimators for multiple first stage parameters. For example, the Fama-MacBeth procedure employs the data on excess returns to a set of portfolios $\{r_{it},\ i=1,...,N,\ t=1,...,T\}$ and time series of a risk factor $\{F_{t},\ t=1,...,T\}.$ Namely, it uses two collections of first stage parameters -- the average return on a portfolio $\beta _{i}^{(1)}=\mathbb{E}r_{it}$ via the sample average return $\widehat{\beta }_{i}^{(1)}=\frac{1}{T} \sum_{t=1}^{T}r_{it},$ and the risk exposure of a portfolio $\beta _{i}^{(2)}=\mathrm{var}(F_{t})^{-1}\mathrm{cov}(r_{it},F_{t})$ via the time series OLS regression of $r_{it}$ on $F_{t}$ resulting in an estimate $ \widehat{\beta }_{i}^{(2)}$. At the second stage of the Fama-MacBeth procedure, one runs the OLS regression of the sample average return $ \widehat{\beta }_{i}^{(1)}$ on the portfolio risk exposure estimated at the first step $\widehat{\beta }_{i}^{(2)}$. In this case,
the second step involves two sample covariances. If one wants to derive the asymptotic distribution of $\widehat{\lambda },$ one needs to establish the asymptotic distribution for a properly normalized sample covariance of the two first step estimators $\frac{1}{N}\sum_{i=1}^{N}\widehat{\beta } _{i}^{(1)}\widehat{\beta }_{i}^{(2)},$ where $\widehat{\beta } _{i}^{(j)}=\beta _{i}^{(j)}+\varepsilon _{i}^{(j)},$ with the estimation error having the structure $\varepsilon _{i}^{(j)}=\big( 1+o_{p}(1)\big) \frac{1}{T}\sum_{t=1}^{T}v_{t}^{(j)}e_{it}$, the term $o_{p}(1)$ being uniform in $i$.
The normalized sample covariance of the two first step estimators contains several terms:
The first term on the right-hand-side of equation ((ref)) is similar to a weighted average of first step estimators and has the form of the first component of $\Xi _{N,T}$ (treating $\beta _{i}^{(j)}$ as constants similar to constants $\gamma _{i}$). The second term in equation ( (ref)) is more complicated and calls for a Central Limit Theorem for quadratic forms:
where $w_{st}=v_{s}^{(1)}v_{t}^{(2)}+v_{t}^{(1)}v_{s}^{(2)}$. Here the first term can be treated as the first component, and the second term as the second component of $\Xi _{N,T}$.
\paragraph{Example 3.}
Anatolyev & Mikusheva (2018) propose a split-sample estimator as an alternative to the Fama-MacBeth procedure for factor asset pricing. There are three sets of parameter estimates produced at the first stage: the sample average return $\widehat{\beta }_{i}^{(1)}=\frac{1}{T} \sum_{t=1}^{T}r_{it}$ and two estimates of the portfolio risk exposure computed as OLS estimates in regressions of $r_{it}$ on $F_{t}$ on different sub-samples, say, $\widehat{\beta }_{i}^{(2)}$ and $\widehat{\beta } _{i}^{(3)}$. In our notation, the usage of a sub-sample is accommodated by setting $v_{t}=0$ for those $t$ not in the currently used sub-sample. The second step IV estimator is constructed as an IV estimator in the regression of $\widehat{\beta }_{i}^{(1)}$ on $\widehat{\beta }_{i}^{(2)}$ using $ \widehat{\beta }_{i}^{(3)}$ as instrument. That is,
In order to make inferences on $\lambda ,$ one needs to obtain the asymptotic distribution of sample covariances between different first step estimates. Statements like ((ref)) and ((ref)) are instrumental to accomplish this.
The configuration in $\Xi _{N,T}$ and a need for statements ((ref)) and ((ref)) occur in other situations as well.
\paragraph{Example 4.}
Pesaran & Yamagata (2018) suggest a new test for factor pricing models that allows many portfolios to be considered simultaneously (with $N$ and $T$ both diverging to infinity). The hypothesis of interest $H_{0}:\alpha _{i}=0$ for all $i=1,...,N$, where $\alpha _{i}$ is a pricing error for the portfolio $i$. To estimate the pricing errors the authors use OLS estimates $ \widehat{\alpha }_{i}$. A large number of portfolios $N$ does not allow one to establish join Gaussianity of all $\widehat{\alpha }_{i}$ or to consistently estimate their covariance. Pesaran & Yamagata (2018) propose to test the hypothesis of interest using statistics based on a weighted sum of squares of $\widehat{\alpha }_{i}$. They create a properly normalized statistic of the form
where $\sigma _{i}^{2}$ are variances of pricing errors. This statistic is directly related to the sample variance of the first step estimator, and a statement of its asymptotic Gaussianity directly follows from ((ref) ) by the same logic as stated above. Pesaran & Yamagata (2018) develop a CLT for quadratic forms that can be applied in this setting. They make an assumption that the idiosyncratic components can be filtered to make them cross-sectionally independent.\footnote{ See Assumptions 2 and 3 in Pesaran & Yamagata (2018).} Here we propose an alternative version of CLT that can be applied under less restrictive assumptions on the cross-sectional dependence of $e_{it}$'s.
\paragraph{Example 5.}
A data-rich IV environment of Bai & Ng (2010) is another example where our linear-quadratic CLTs can be useful. The authors consider an IV setup with many instruments in a panel, where the number of instruments, $N$, is potentially higher than the number of observations, $T$. The instruments are generated by a factor model $z_{it}=\lambda _{i}^{\prime }F_{t}+e_{it},$ with $F_{t}$ and $e_{it}$ independent of the structural error $\varepsilon _{t},$ and time-series and cross-sectional dependence in $e_{it}$ is allowed, though restricted. Bai & Ng (2010) consider the bias-corrected GMM estimator that corrects for inconsistency of the baseline GMM. It is consistent when $N/T=O(1),$ however, its asymptotic Gaussianity is established under a more restrictive assumption when $N/T=o(1).$ The challenge is that when $N/T=O(1),$ the asymptotic expansion for this estimator has, in addition to a linear term, a quadratic form in the idiosyncratic components $e_{it}$ similar to the second component of $\Xi _{N,T}$. Thus, using statement ((ref)), an asymptotic theory could be developed for the bias-corrected GMM estimator without having to impose $ N/T=o(1)$.
In most of these examples, the set of idiosyncratic components $\{e_{it},$ $ i=1,...,N,$ $t=1,...,T\}$ cannot be regarded independent and/or identically distributed. In most realistic applications, one is usually willing to assume that $e_{it}$ do not have a strong (detectable) factor structure, but still allow for some correlation between different units, which would not affect consistency. For example, it is reasonable to think that stocks of firms in the same industry or of the same size may react to some local shocks and be correlated, though when averaged over all stocks (and all industries), this co-movement of returns would have no first-order impact on estimation.
Our attempt to be agnostic with regard to possible cross-sectional correlation among errors and to avoid explicit modeling of its structure whenever possible comes at a cost of more restrictive time series assumptions. In many applications of interest, it is more credible to impose independence assumptions in a time-series direction rather than in a cross-sectional direction. For example, the efficient market hypothesis implies mean non-predictability of excess returns given past history, which is equivalent to a martingale difference property for the errors. The definition of shocks in macroeconomics similarly presumes their time-series independence. In this paper, we assume time-series independence, which in some cases may be weakened to the martingale difference property or stationarity with some proper mixing condition, but we do not pursue this generalization here.
In this paper we consider asymptotics as both cross-sectional and time-series sample sizes, $N$ and $T$, increase to infinity. We allow the data-generating process for all variables to vary with $N$ and $T$. Define $ \mathcal{F}$ to be a $\sigma$-algebra that contains at least the $\sigma$-algebras generated by the full set of variables $\{v_{s},s=1,...,\infty\}$ and $ \{w_{st}, s,t=1,...,\infty\}$ for all $s $ and $t$. It may potentially also contain other events related to common shocks and variables, as long as Assumption (ref) stated below is satisfied. We treat $\gamma _{i} $ as non-random $k_{\gamma }\times 1$ vectors.
In order to simplify the notation, in what follows we will denote $C$ to be a positive generic constant, independent of $N$ and $T,$ which may be different in different equations, but does not depend on or change with $N$ or $T$. We will use the following notation: for a square matrix $A$, we denote by $\mathrm{tr}(A)$ its trace, by $\max \mathrm{ev}(A)$ -- its maximal eigenvalue, and by $\mathrm{dg}(A)$ a diagonal matrix of the same size with the elements from the diagonal of $A$; $\Vert \cdot \Vert $ is the $l_{2}$ norm for a vector or the operator norm for a matrix.
Assumption (ref) imposes very mild restrictions on the time-series behavior of the common (non-entity specific) variables. For example, the part related to $v_{t}$ is trivially satisfied if a time series equal to $v_{t}v_{t}^{\prime }$ is weakly stationary with summable auto-covariances. Assumption (ref) restricts the influence of any one entity in the cross-sectional average and will eventually contribute to asymptotic negligence of the cross-sectional summands needed for the CLT. Assumption (ref)(i) is a restrictive assumption which imposes discipline on the time-series structure, and the restriction $\mathbb{E} (e_{t}|\mathcal{F})=0$ is a form of strict exogeneity in the first step regression. Uniform moment boundedness in Assumption (ref)(ii) is traditional.
Apparently, Assumptions (ref), (ref) and (ref) are insufficient to establish a central limit theorem, and we need to put some restrictions on the cross-sectional dependence and dependence between idiosyncratic errors and common variables. Indeed, we will use a change of summation ordering:
and establish asymptotic convergence in the time-series direction. In order to apply a CLT in the time series direction we need some sort of asymptotic negligibility of summands with different time indexes, in particular, of terms like $\big\{ v_{s}(\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\gamma _{i}e_{is})\big\} _{s}$ and $\big\{ w_{st}(\frac{1}{\sqrt{N}} \sum_{i=1}^{N}e_{it}e_{is})\big\} _{s,t}$. Our goal is to provide low-level assumptions. There is a trade-off in how much dependence of idiosyncratic errors across entities and how much dependence between idiosyncratic errors and common variables can be allowed. Below we consider two particular cases. In the first case, full independence between the $e_{it}$ 's and $\mathcal{F}$ is assumed; as a result, we can be agnostic about the structure of cross-sectional dependence, the corresponding assumptions about it are relatively mild. In the second case, we allow for conditional heteroscedasticity in $e_{it}$ that can be related to some common variables from $\mathcal{F}$ producing dependence in higher-order conditional moments. This flexibility comes at the cost of imposing some structure on the cross-sectional behavior of $e_{it}$.
Numerous papers that establish inferences in factor models commonly assume that the set of factors is independent from the set of idiosyncratic errors, as in Assumption (ref)(i), though cross-sectional dependence of errors is allowed; see, for example, Assumption D in Bai & Ng (2006). We intended for the first part of Assumption (ref) (ii) to impose weak cross-sectional dependence as expressed by the covariance matrix; in particular, it means that no strong factor structure is left in the errors; similar assumptions appear in Onatski (2012) and Bai & Ng (2006). The convergence of the trace in Assumptions (ref)(ii) and (ref)(iii) is needed for the asymptotic covariance matrix to be properly defined.
Assumption (ref)(iv) is another way to restrict pervasive dependence in multiple variables, in particular, precluding outliers to realize in too many error terms simultaneously. For example, imagine that the cross-sectional dependence is induced by several groups with a factor structure, e.g., stock returns are correlated because there are industry-specific shocks and geography-specific shocks. Imagine that there are a finite number, say $G$, groups, indexed by $g=1,...,G$, \ which may be overlapping, with each having independent shocks $f_{g,t}$ at time $t$. Stock $i$ has non-zero loading $\pi _{i,g}$ only if it belongs to group $g$. Let the set of groups, to which $i$ belongs, be denoted by $G(i)$. That is,
where $\eta _{it}$'s are independent both cross-sectionally and across time and have finite fourth cumulants. Then, Assumption (ref) (iv) is essentially equivalent to the following two conditions: $\mathbb{E} (f_{g,t}^{4})<C$ and $\frac{1}{N}\big( \sum_{i=1}^{N}|\pi _{i,g}|\big) ^{2}<C$ for any $g=1,...,G$. Thus, for this example, essentially Assumption (ref)(iv) imposes that the factors $f_{g,t}$ do not produce outliers too often expressed as the moment condition and a statement about pervasiveness.
One of the important steps in the proof of Theorem (ref) verifies asymptotic negligibility of time-series summands by checking boundedness of the fourth moments of the cross-sectional sums $ \frac{1}{\sqrt{N}}\sum_{i=1}^{N}\gamma _{i}e_{is}$ and $\frac{1}{\sqrt{N}} \sum_{i=1}^{N}e_{it}e_{is}$; that imposes the main way we restrict cross-sectional dependence. The fourth cumulant conditions are reminiscent of those in de Jong (1987), which we follow while proving our CLT using Heyde & Brown (1970). There are alternative CLTs for quadratic forms such as Rotar' (1973) that imposes weaker moment conditions on the summands but stricter assumptions on the negligibility of coefficients and eigenvalues of the quadratic form. In our case, following them would require imposing stronger assumptions on the variables $w_{st},$ which we would like to avoid. Another CLT for quadratic forms for time series data can be obtained using Bhansali et al. (2007). The book by Giraitis et al. (2012) has a chapter on this subject and allows for long memory time series as well.
Assumption (ref)(i) of independence is much stronger than Assumption (ref)(i) about exogeneity: it does not allow higher conditional moments of $e_{it}$ to co-move with the common variables; in particular, it imposes conditional homoscedasticity. It may be especially problematic in financial applications where time-varying volatility is of strong empirical relevance, and returns on many stocks display patterns of changing volatility driven by some common variables. The assumptions below allow for conditional heteroscedasticity.
An interesting feature of this example is that it allows the errors to be weakly cross-sectionally dependent to the extent that they may possess a weak (latent) factor structure. The condition $\mathbb{E}(f_{t}f_{t}^{\prime })=I_{k_{f}}$ is a normalization and involves no loss of generality. Assumption (ref)(ii) forces the factors to be weak to such an extent that the factor structure cannot be consistently detected; it implies that the covariance matrix of idiosyncratic errors would satisfy the first half of Assumption (ref)(ii). Moreover, this factor structure may be closely related to the common variables in $\mathcal{F}$, which causes the cross-sectional dependence among the errors $e_{it}$ to change with the common variables and allows a very flexible form of conditional heteroscedasticity. Indeed, the conditional cross-sectional covariance is
Since we do not restrict $\mathbb{E}(f_{t}f_{t}^{\prime }|\mathcal{F})$ beyond proper moment conditions, the strength of any cross-sectional dependence as well as error variances may change stochastically depending on realizations of the common variables.
The moment conditions in Assumption (ref)(i) help to establish asymptotic negligibility of the time-series summands. Assumption (ref)(iii) about $\Gamma _{\omega }$ and Assumption (ref) (iv) allow us to define properly the asymptotic covariance matrix.
In this Section, we first discuss estimation of asymptotic variances for asymptotic inference when this leads to valid inference. Then, we propose alternative tools based on the wild bootstrap to apply in situations when asymptotic inference fails to provide asymptotically correct inference.
Statistical inferences such as confidence set construction and hypotheses testing about the structural parameter typically require consistent estimation of asymptotic variances of all important quantities that are asymptotically Gaussian. The easiest to implement and thus the most appealing from an applied perspective are those that use the same variables and have a structure similar to the original averages, such as the statement in equation ((ref)).
Notice that equation ((ref)) contains the cross-sectional summation outside, and hence it treats the cross-section as nearly uncorrelated observations, or at least it ignores the cross-sectional correlation. A relevant analogue is the difference between the long-run covariance and instantaneous covariance in a classical time series. However, implementing an analogue of long-run covariance estimation here would be a challenge since we do not have any cross-sectional stationarity or a measure of distance between cross-sectional entities. Rather, we explore under which conditions the convergence in ((ref)) holds.
Theorem (ref) below obtains a statement for the case when the common variables are independent from the idiosyncratic errors, while Theorem (ref) establishes a similar statement for the conditionally heteroscedastic case.
The additional assumption ((ref)) in Theorem (ref) strengthens conditions on the weakness of the cross-sectional correlation; in particular, it requires that the covariance matrix converges to a diagonal one. The additional assumption in Theorem (ref) requires that the weights used for averaging the cross-sectional entities are orthogonal to the loadings on the latent factor structure, which precludes the latent factor structure (that represents the cross-sectional dependence) from being amplified. This is a necessary assumption for consistency of the variance estimator. Indeed, let assumptions (ref), (ref), (ref), (ref) hold, and consider the first component of $\xi _{i}$:
where $\tilde{\eta}_{i}=\frac{1}{\sqrt{T}}\sum_{t}v_{t}\gamma _{i}\eta _{it}$ , and $\Upsilon _{T}=\frac{1}{\sqrt{T}}\sum_{t=1}^{T}f_{t}v_{t}$. Note that $ \frac{1}{\sqrt{N}}\sum_{i=1}^{N}\tilde{\eta}_{i}\Rightarrow \mathcal{N} (0,\sigma _{\eta }^{2})$ and $\frac{1}{N}\sum_{i=1}^{N}\tilde{\eta}_{i}^{2} \overset{p}{\rightarrow }\sigma _{\eta }^{2},$ as all conditions of Theorems (ref) and (ref) are satisfied by cross-sectionally and time independent errors $\eta _{is}$. Assumption (ref)(iv) guarantees that $\Upsilon _{T}\Rightarrow \mathcal{N} (0,\Sigma _{fv})$ as $T\rightarrow \infty $, while according to Assumption (ref)(ii), we have $\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\pi _{i}\gamma _{i}^{\prime }\rightarrow \Gamma _{\pi \gamma }$ as $N\rightarrow \infty $. Thus,
while $\frac{1}{N}\sum_{i=1}^{N}\big(\xi _{i}^{(1)}\big)^{2}\overset{p}{\rightarrow } \sigma _{\eta }^{2}$ because $\frac{1}{N}\sum_{i=1}^{N}\pi _{i}\gamma _{i}^{\prime }\gamma _{i}\pi _{i}^{\prime }\rightarrow 0$ by Assumptions (ref) and (ref)(ii).
As one way to conduct valid inferences in settings when $\frac{1}{N} \sum_{i=1}^{N}\xi _{i}\xi _{i}^{\prime }$ is an inconsistent estimator of the variance (in particular, when $\Gamma _{\pi \gamma }\neq 0$ under Assumption (ref)), we propose the following simple wild bootstrap procedure. In each bootstrap repetition,
Then the distribution of $\Xi _{N,T}^{\ast }=\frac{1}{\sqrt{N}} \sum_{i=1}^{N}\xi _{i}^{\ast }$ has the same asymptotic limit as that of $ \Xi _{N,T}$ has, and can be used for inferences. Alternatively, the distribution of $\Xi _{N,T}^{\ast }$ normalized by the bootstrap analogue of the variance estimate $\frac{1}{N}\sum_{i=1}^{N}\xi _{i}^{\ast }\xi _{i}^{\ast \prime }$ can be used to approximate the distribution of $\Xi _{N,T}$ normalized by the variance estimate $\frac{1}{N}\sum_{i=1}^{N}\xi _{i}\xi _{i}^{\prime }.$ We call the two described bootstrap procedures bootstrap and bootstrap-t. In the next Section, we implement both variations of the wild bootstrap for the setup of Assumption (ref).
The wild bootstrap works because it introduces independence in the time direction while preserving the (unknown) cross-sectional dependence. Specifically, we base our proof of Theorem (ref) on the change of order of summations in double/triple summations over $i$ and time index/indices. For example,
We then argue that our assumptions guarantee that the CLT with respect to summation over $t$ is applicable. If the assumptions of the CLT hold, then $ \Xi _{N,T}^{(1)}$ is asymptotically Gaussian with mean zero and variance equal to the limit of $\frac{1}{T}\sum_{t=1}^T\mathbb{E}\big[\big(\frac{1}{\sqrt{N}} \sum_{i}v_{t}\gamma _{i}e_{it}\big)^{2}\big]$. This limit clearly depends on how much there is cross-sectional correlation between $e_{it}$ and $e_{jt}$. Now, in the bootstrapped samples,
Conditional on the original sample, only $\delta _{t}$'s are random, and they are independent and have zero mean and unit variance. When $T$ is large, this bootstrapped normalized sum satisfies the CLT and thus converges to a zero mean Gaussian random variable with the variance equal to the limit of $\frac{1}{T}\sum_{t=1}^{T}\big(\frac{1}{\sqrt{N}}\sum_{i=1}^{N}v_{t}\gamma _{i}e_{it}\big)^{2}$ as $N,T\rightarrow \infty $. This limit coincides with $ \frac{1}{T}\sum_{t=1}^{T}\mathbb{E}\big[\big(\frac{1}{\sqrt{N}}\sum_{i=1}^{N}v_{t} \gamma _{i}e_{it}\big)^{2}\big]$ under relatively weak assumptions, as long as the Law of Large Numbers holds. For example, under Assumptions (ref), we have:
A similar argument can be made about the second component as well. Indeed,
For a bootstrapped statistic, conditional on the original sample, the weights $\frac{1}{\sqrt{N}}\sum_{i=1}^{N}w_{st}e_{it}e_{is}$ are fixed, while $\delta _{t}$'s are independent random variables, and all the conditions of Lemma (ref) are satisfied. Thus, a zero mean Gaussian limit obtains as $T\rightarrow \infty $. It is straightforward to verify that its variance converges to the asymptotic variance of $\Xi _{N,T}^{(2)}$.
In the simulation experiments in the following Section, we check, among other things, that both wild bootstrap variations deliver correctly sized tests even when the asymptotic-t tests fail to do so, and also that the bootstrap-based tests have a non-trivial power.
The goals of this Section are to check finite sample performance of an asymptotic Gaussian approximation for $\Xi _{N,T}$, to explore when the variance estimator $\widehat{\Sigma }_{\xi }=\frac{1}{N} \sum_{i=1}^{N}\xi _{i}\xi _{i}^{\prime }$ allows to construct reliable asymptotic-t inferences, and to evaluate the performance of the wild bootstrap and bootstrap-t procedures in terms of both size and power.
Our setup adheres to Assumption (ref). We generate the errors $ e_{it}$ according to the following weak (unobserved) factor structure:
where $f_{t}\sim iid$ $\mathcal{N}(0,1)$ across $t=1,...,T,$ $\gamma _{i}=1$ for all $i=1,...,N,$ and $\eta _{it}=\omega _{i}\epsilon _{\eta ,it},$ where $\epsilon _{\eta ,it}$ are $iid$ $\mathcal{N}(0,1)$ across $i$ and $t$. The standard deviations are set to $\omega _{i}=c_{\omega }\left( 1+\left\vert \tau _{i}\right\vert \right) ,$ and the factor loadings are $\pi _{i}=\left( c_{\pi }+\tau _{i}\right) /\sqrt{N},$ where $\tau _{i}\sim iid$ $\mathcal{N} (0,1)$ across $i=1,...,N.$ The multiplier $c_{\omega }$ is tuned so that the average cross-sectional variance of $\eta _{it}$ is unity. The parameter $ c_{\pi }$ indexes the degree of cross-sectional dependence as measured by the strength of the factor structure. Specifically, as $\sum_{i=1}^{N}\pi _{i}\pi _{i}^{\prime }\rightarrow 1+c_{\pi }^{2}$, this parameter is assumed to be bounded for the Gaussian approximation to hold. We also have $\Gamma _{\pi \gamma }=c_{\pi }$, hence we can expect consistency of variance estimation only when $c_{\pi }=0$, so we will explore the distortions for different values of $c_{\pi }$. The errors generated this way are cross-sectionally dependent and heteroscedastic, while still satisfying Assumption (ref).
The common variables are generated as follows: $v_{t}=c_{fv}f_{t}+\sqrt{ 1-c_{fv}^{2}}\epsilon _{v,t},$ with $\epsilon _{v,t}\sim iid$\ $\mathcal{N} (0,1)$ across $t=1,...,T,$ and $w_{st}=v_{s}v_{t},$ $s,t=1,...,T.$ All the disturbances $f_{t},$ $\tau _{i},$ $\epsilon _{\eta ,it}$ and $\epsilon _{v,t}$ are mutually independent. The parameter $c_{fv}$ indexes the dependence between common variables and $e_{it}$. The mean zero Assumption (ref)(i) requires $c_{fv}=0$; the non-zero values of $c_{fv}$ index deviations from the null hypothesis $\mathbb{E}(\xi _{i})=0,$ and will be used to study the power properties of the proposed wild bootstrap. In wild bootstrap samples, we generate the bootstrap errors by $e_{it}^{\ast }=\delta _{t}e_{it},$ where $\delta _{t}=2\zeta _{t}-1,$ and $\zeta _{t}\sim iid\ \mathcal{B}(\frac{1}{2})$ across $t,...,T$. The bootstrap analogues of $ \gamma _{i},$ $v_{t}$ and $w_{st}$ are set equal to their original sample values.
The distribution characteristics are computed from 10,000 simulations, while the rejection rates are based on 5,000 simulation runs. In all simulations, we set $N=T=500.$ The number of bootstrap repetitions is 600.
Table (ref) contains distributional characteristics of $\Xi _{N,T}$. We report averages, coefficients of skewness, coefficients of kurtosis, and right 5% quantiles of normalized marginal distributions of both elements of $\Xi _{N,T}.$ For the exactly normal distributions, these values are $0$, $0$, $3$ and $1.645,$ respectively.
The actual distribution of the linear component of $\Xi _{N,T}$ is very close to Gaussian, in all respects: all the moments and the right tail are almost equal to their theoretical counterparts. The quadratic component of $ \Xi _{N,T},$ however, albeit mean unbiased, is somewhat positively skewed and a bit leptokurtic. The shifted right quantile confirms slight over-dispersion. The distortions, however, do not seem to increase with the strength of the error factor structure.
In Table (ref) we document the empirical rejection rates for tests with $10\%,$ $5\%$ and $1\%$ declared size based on the asymptotic-t, wild bootstrap and wild bootstrap-t approaches. In the asymptotic-t approach we create a $t$-statistic using $\widehat{\Sigma }_{\xi }$ as a variance estimator and compare it with the symmetric standard Gaussian critical values. We also explore the performance of two wild bootstrap procedures -- one that bootstraps $\Xi _{N,T}$ and another that bootstraps the $t$ -statistics (referred to in Table (ref) as bootstrap $\xi $ and bootstrap $t$). In both bootstrap procedures, we compare the absolute value of the statistic from the sample to the right quantile of the absolute value of the bootstrapped statistic. We verify the empirical size by setting $ c_{fv}=0$ and empirical power by setting $c_{fv}=0.1$ for relatively small deviations from the null and $c_{fv}=0.2$ for relatively large deviations from the null.
As expected, the size of the test based on asymptotic approximations sometimes deviates from nominal rates by a wide margin, especially for the linear component of $\xi _{N,T},$ the gap quickly increasing with the strength of the error factor structure. This happens due to inconsistency of the variance estimator $\widehat{\Sigma }_{\xi }$ and becomes more pronounced with stronger cross-sectional dependence. What is surprising is that for the quadratic component of $\xi _{N,T},$ the distortions are relatively minor and not very sensitive to the strength of the factor structure. In contrast, both bootstrap procedures exhibit excellent size control and stability thereof across the strength of the error factor structure for both components of $\xi _{N,T}.$ In terms of power, however, the two bootstrap statistics are approximately equally powerful for the linear component of $\xi _{N,T},$ while there is a gap, sometimes sizable, between power figures for its quadratic component. It seems that bootstrapping the statistic itself is preferable.
Possible directions for future research may be relaxing the error time-series independence to martingale difference structures and inventing ways to consistently estimate the asymptotic variance matrix when it is not diagonal in the limit. Other interesting areas involve establishing formal properties of the proposed wild bootstrap schemes, exploring the possibility of asymptotic refinements, and examining the superiority of one bootstrap scheme over the other.