The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
61,207 characters
Limit Theorems for Factor Models
\title{Limit Theorems for Factor Models\thanks{
We thank the Editor Peter C.B. Phillips, Co-Editor Guido Kuersteiner, and three anonymous referees for guidance and helpful suggestions that have greatly improved the
paper.\smallskip }}
\author{Stanislav Anatolyev\thanks{
Address: Stanislav Anatolyev, CERGE-EI, Politickych v\v{e}z\v{n}\r{u} 7,
11121 Prague 1, Czech Republic; e-mail \texttt{
[email removed]}. Czech Science Foundation support under
grants 17-26535S and 20-28055S is gratefully acknowledged.\smallskip } \\
CERGE-EI and NES \and Anna Mikusheva\thanks{
Address: Department of Economics, M.I.T., 50 Memorial Drive, Building E52,
Cambridge, MA, 02142, USA; e-mail \texttt{[email removed]}. National
Science Foundation support under grant 1757199 is gratefully
acknowledged. } \\
MIT\bigskip }
\date{September 2020 \\
}
\maketitle
\begin{abstract}
\quad \newline
\noindent The paper establishes central limit theorems and proposes how to perform valid inference in factor models. We consider a setting where many counties/regions/assets are observed for many time periods, and when estimation of a global parameter includes aggregation of a cross-section of heterogeneous micro-parameters estimated separately for each entity. The central limit theorem applies for quantities involving both cross-sectional and time series aggregation, as well as for quadratic forms in time-aggregated errors. The paper studies the conditions when one can consistently estimate the asymptotic variance, and proposes a bootstrap scheme for cases when one cannot. A small simulation study illustrates performance of the asymptotic and bootstrap procedures. The results are useful for making inferences in two-step estimation procedures related to factor models, as well as in other related contexts. Our treatment avoids structural modeling of cross-sectional dependence but imposes time-series independence.
\bigskip \bigskip
\noindent \textbf{Keywords:} factor models, two-step procedure, dimension
asymptotics, central limit theorem.\bigskip
\noindent \textbf{JEL classification codes:} C13, C33, C38, C55.\bigskip
\bigskip \bigskip
\end{abstract}
\thispagestyle{empty} \setcounter{page}{0} \baselineskip=18.0pt
\section{INTRODUCTION}
Data with an underlying factor structure are increasingly used in empirical
macroeconomics and finance. Often these data consist of time series of
observations for multiple cross-sectional units (assets, portfolios, regions
or industries). Quite a few new estimation strategies have appeared in the
empirical literature that use both cross-sectional and time series variation
in order to estimate global structural parameters. Often the parameter of
interest arises from aggregation or estimation using cross-sectional
variation of individual parameters for each entity. One example of such a
structure is linear factor pricing model in asset pricing (Fama \& MacBeth
1973 and Shanken 1992), where for estimation we usually use time series of
excess returns for a number of portfolios or assets priced by a small number
of risk factors. Each portfolio or stock may have its own (heterogeneous)
exposure to risk, often referred to as betas, which can be estimated
separately from time series observations for each portfolio. The parameter
of interest, a risk premium, is defined as the coefficient of
proportionality in the cross-sectional relation between the average excess
return on a portfolio and its individual beta.
A vast majority of macroeconomic shocks are only weakly identified via
structural VARs that use only time series observations on leading macro
variables. A new approach to the estimation of causal effects of a macro
shock on the economy is to use cross-sectional variation in data on regions,
countries or industries. For example, Serrato \& Wingender (2016) use
cross-sectional variation in federal spending programs due to a Census shock
to identify the causal impact of government spending on the economy.
Cross-sectional variation among counties in government spending and in the
accuracy of census-based estimates of population provides a better justified
treatment effect framework, allows for the estimation of local fiscal
multipliers, and finally gives a better global estimate of the fiscal
multiplier via aggregation of local multipliers. Hagedorn et al. (2015)
estimate the aggregate effect of unemployment-benefit duration
on employment and labor force participation using cross-sectional
differences across US states. Sarto (2018) discusses how heterogeneous
sensitivities of regions to aggregate policy variables, so called
micro-global elasticities, can be used to recover macro elasticities of
interest such as, for example, a fiscal multiplier.
A shared feature of the above-mentioned examples is the use of time-series
observations on multiple entities (stocks, portfolios, counties, states or
industries), while data on those entities are not independent and
identically distributed. Moreover, variables for different entities often
display strong co-movements to the extent that the data have a factor
structure, and estimation of these co-movements is the main goal. Indeed,
the realization of a risk factor in the economy moves returns on all
portfolios simultaneously, while a federal fiscal shock moves spending in
all US counties, though in both cases heterogeneously so. A valid estimation
procedure must explicitly model and account for the data's factor structure
to the extent that the error terms (or residuals) can be considered
idiosyncratic; see Kleibergen \& Zhan (2015) and Anatolyev \& Mikusheva
(2018) for how a factor structure that is unaccounted for can lead to
misleading results. However, idiosyncrasy of the errors usually implies
only that the correlation among errors for different entities is relatively
small and does not introduce first-order bias to the estimation procedure.
Usually, it is not reasonable to assume that errors for different entities
are completely independent; indeed, stocks in the same industry are likely
to co-move even after global-economy risks are removed, while errors for
neighboring counties are more likely to be correlated even after one
accounts for federal shocks. At the same time, we typically want to remain
agnostic about the correlation structure of shocks and avoid their
structural modeling as long as this does not introduce biases.
The second typical feature of the above-mentioned examples is the two-step
nature of the estimation procedure, where in the first step we estimate
entity-specific coefficients (risk exposures/betas, local fiscal
multipliers, micro-global elasticities) by running a time-series regression
separately for each entity. In the second step, we estimate the global
coefficient of interest by either aggregating entity-specific coefficients
(Serrato \& Wingender 2016 and Hagedorn et al. 2015), or by
running an OLS regression on the cross-section of entity-specific
coefficients (Fama \& MacBeth 1973 and Sarto 2018), or by running an IV
regression on the cross-section of entity-specific coefficients (Anatolyev
\& Mikusheva 2018).
The goal of this paper is to establish central limit theorems (CLTs) and to
provide a tool for establishing asymptotic normality of estimates obtained
in such two-step estimation procedures and for finding ways to do
asymptotically correct inference, while being flexible in modeling the
cross-sectional dependence of errors. The main difficulty here is that even
though the second step cross-sectional regression has nearly uncorrelated
errors (which is usually sufficient to obtain consistency of the two-step
estimator), this condition is usually insufficient for a CLT, which
typically requires that stronger discipline be imposed on the dependence
structure (such as independence, or a martingale difference structure, or
mixing). Our solution to this problem is to restrict the time series
behavior while staying agnostic about the cross-sectional dependence. We
assume time-series independence of idiosyncratic errors, which is consistent
with market efficiency for factor asset pricing models and the
non-predictability of macro shocks in macroeconomic settings. The estimation
noise in a two-stage procedure involves aggregation both over time (from the
first step) and over entities (from the second step). We show that under
certain conditions it is sufficient to have a CLT over just one of these
directions, and we use the time-series direction for that.
When the second step uses an OLS or IV estimator, the CLT must adapt to
averages of \textit{quadratic forms}, as both the second-step-dependent
variable and the second stage regressor/instrument contain first-stage
estimation noise. Our CLT has a linear and a quadratic part. We also note
that a need for a CLT for quadratic forms in factor models sometimes arises
for the first-step estimators (e.g., Pesaran \& Yamagata 2018) or in
higher order asymptotic derivations (e.g., Bai \& Ng 2010).
There is a growing literature that establishes different CLTs while
acknowledging the importance of cross-sectional dependence in the data,
which stems from spatial relations and/or from the presence of common
factors. Kuersteiner \& Prucha (2013) establish a CLT for linear sums in a
panel data context with growing cross-sectional dimension $N$ and fixed
time-series dimension $T$ allowing for cross-sectional dependence, and
Kuersteiner \& Prucha (2020) extend these results to quadratic forms as
well. Both papers impose conditional moment restrictions, which allows the
authors to construct a martingale difference sequence in the cross-sectional
direction. The main conditional moment restrictions imply a correct
specification of an underlying model, which need not be required by our CLT.
However, the mentioned papers allow more flexibility in modeling the time
dependence, and do not require large $T$. Another CLT that requires both
large $N$ and large $T$ is established in Hahn et al. (2020) for linear
terms only.
This paper also contributes to the literature on the CLT for quadratic
forms. Various types of CLTs for quadratic forms have been previously
established and used in the many instrument literature (see for example,
Chao et al. 2012, Hausman et al. 2012, S\o lvsten 2020) and many covariate
literature (see Cattaneo et al. 2018), as well as in the literature on
semi-parametric estimation (Cattaneo et al. 2014a, 2014b). The CLT used in
those papers are established for the cross-sectional dimension only, and
rely heavily on the independence assumption. We adapt the ideas used in Chao
et al. (2012), specifically the approach of de Jong (1987), to accommodate
large cross-sectionally dependent panels; an alternative approach, known as
Stein's method, is used in S\o lvsten (2020).
Our second set of results is related to ways of conducting valid statistical
inference. Under strengthened conditions on the weakness of the
cross-sectional correlation of errors, we show that a conventional variance
estimator is consistent, and so the usual asymptotic inference can be
applied. When such strengthened conditions do not hold, we propose instead a
variant of a wild bootstrap scheme that replicates the original
cross-sectional dependence structure. We also conduct a small simulation
experiment that provides evidence on the approximation quality of our CLT
and on the empirical size and power of wild bootstrap in a moderately sized
panel.
The paper proceeds as follows. Section \ref{section-notation} explains
problems with establishing asymptotic Gaussianity for two-step and other
estimators and test statistics, and shows how discipline in the time series
direction can help. Section \ref{section-CLT} introduces assumptions on
idiosyncratic errors, states central limit theorems for two cases, and
discusses the relevance of those cases to empirical practice. Section \ref
{section-variance} discusses estimation of asymptotic variances for
asymptotic inference and alternative inference tools based on the
bootstrap. Section \ref{section-simulations} presents a small simulation
experiment that reveals properties of asymptotic and proposed bootstrap
inference tools. Section 6 concludes. All proofs appear in the Appendix.
\section{GOALS\ AND\ EXAMPLES}
\label{section-notation}Let the data contain observations on many units
indexed by $i=1,...,N,$ and observed for multiple time periods $t=1,...,T.$
We assume that both $N$ and $T$ increase to infinity without restrictions on
their rates. The goal of this paper is to find the conditions under which
the following statement will hold:
\begin{equation}
\Xi _{N,T}\equiv \frac{1}{\sqrt{N}}\sum_{i=1}^{N}\xi _{i}\Rightarrow
\mathcal{N}(0,\Sigma _{\xi }), \label{eq: CLT}
\end{equation}
where
\begin{equation*}
\xi _{i}=\left(
\begin{array}{c}
\frac{1}{\sqrt{T}}\sum_{s=1}^{T}v_{s}\gamma _{i}e_{is} \\
\frac{1}{T}\sum_{s=1}^{T}\sum_{t<s}w_{st}e_{it}e_{is}
\end{array}
\right) ,
\end{equation*}
and $\Sigma _{\xi }$ is an asymptotic variance matrix. Here, $e_{it}$ are weakly cross-sectionally dependent entity-specific (idiosyncratic)\footnote{
By idiosyncratic error we mean the factor-removed part of entity-specific
variables.} errors with $\mathbb{E}\left( e_{it}\right) =0$. Errors $e_{it}$
are uncorrelated with the variables $v_{t}$ and $w_{st}$ that are common to
all units $i=1,...,N$ (more exact conditions are to appear in the next
Section). We assume $\gamma _{i},$ $i=1,...,N,$ to be non-random
entity-specific weights. Further, we want to study the circumstances when
one can also consistently estimate the asymptotic covariance -- that is,
sufficient conditions for a statement like
\begin{equation}
\frac{1}{N}\sum_{i=1}^{N}\xi _{i}\xi _{i}^{\prime }\overset{p}{\rightarrow }
\Sigma _{\xi }. \label{eq: variance}
\end{equation}
As we argue below (see Examples 1--3), statements (\ref{eq: CLT}) and (\ref
{eq: variance}) are often needed in order to conduct statistical inferences
(testing or confidence set construction) about a structural parameter, $
\lambda ,$ which is estimated in two steps. We consider a case when in the
first step a researcher estimates a parameter $\beta _{i}$ for each
entity/unit/state $i=1,...,N$, typically via running OLS or IV time series
regressions. A typical linear estimator can be written as $\widehat{\beta }
_{i}=\beta _{i}+\varepsilon _{i},$ where the estimation error has the
structure $\varepsilon _{i}=\big( 1+o_{p}(1)\big) \frac{1}{T}
\sum_{t=1}^{T}v_{t}e_{it}$, with the $o_{p}(1)$ term uniformly small over
the units. In this setting, $v_{t}$ is either a regressor common to all
entities, or a common systematic part of entity-specific regressors that
have a factor structure.\footnote{
Our setting can accommodate entity-specific regressors, say $v_{it}$, that
have a factor structure themselves. Assume that $v_{it}=a_{i}u_{t}+u_{it},$
where $u_{t}$ is a common co-movement in the regressors and $u_{it}$ is
idiosyncratic. Then
\begin{equation*}
\varepsilon _{i}=\frac{1}{T}\sum_{t=1}^{T}v_{it}e_{it}=\frac{1}{T}
\sum_{t=1}^{T}u_{t}(a_{i}e_{it})+\frac{1}{T}\sum_{t=1}^{T}u_{it}e_{it}=\frac{
1}{T}\sum_{t=1}^{T}v_{t}e_{it}^{\ast },
\end{equation*}
where $v_{t}=(u_{t},1)^{\prime }$ and $e_{it}^{\ast
}=(a_{i}e_{it},u_{it}e_{it})$.}
\paragraph{Example 1.}
There is a variety of estimation approaches that can be used at the second
step. The simplest of them is weighted averaging of the first step
estimates, viz. $\widehat{\lambda }=\frac{1}{N}\sum_{i=1}^{N}\gamma _{i}
\widehat{\beta }_{i}.$ Such an estimator is used in Sarto (2018). In order to
justify asymptotic Gaussianity of $\widehat{\lambda }$ and to make
statistical inferences about $\lambda ,$ one needs statements on the
asymptotic behavior of
\begin{equation*}
\sqrt{\frac{T}{N}}\sum_{i=1}^{N}\gamma _{i}\varepsilon _{i}=\frac{1}{\sqrt{N}
}\sum_{i=1}^{N}\frac{1}{\sqrt{T}}\sum_{t=1}^{T}v_{t}\gamma _{i}e_{it}.
\end{equation*}
Note that the last expression has the structure of normalized averages stated
as the first component of $\Xi _{N,T}$ from equation (\ref{eq: CLT}). Such `linear' terms, where only the
first component of $\Xi _{N,T}$ is involved, are very common in asymptotic
derivations in factor models (e.g., Bai \& Ng 2006, 2010).
\paragraph{Example 2.}
The second estimation step may invoke a more complex estimator involving a
sample covariance between multiple first stage estimators or estimators for
multiple first stage parameters. For example, the Fama-MacBeth procedure
employs the data on excess returns to a set of portfolios $\{r_{it},\
i=1,...,N,\ t=1,...,T\}$ and time series of a risk factor $\{F_{t},\
t=1,...,T\}.$ Namely, it uses two collections of first stage parameters --
the average return on a portfolio $\beta _{i}^{(1)}=\mathbb{E}r_{it}$ via
the sample average return $\widehat{\beta }_{i}^{(1)}=\frac{1}{T}
\sum_{t=1}^{T}r_{it},$ and the risk exposure of a portfolio $\beta
_{i}^{(2)}=\mathrm{var}(F_{t})^{-1}\mathrm{cov}(r_{it},F_{t})$ via the time
series OLS regression of $r_{it}$ on $F_{t}$ resulting in an estimate $
\widehat{\beta }_{i}^{(2)}$. At the second stage of the Fama-MacBeth
procedure, one runs the OLS regression of the sample average return $
\widehat{\beta }_{i}^{(1)}$ on the portfolio risk exposure estimated at the
first step $\widehat{\beta }_{i}^{(2)}$. In this case,
\begin{equation*}
\widehat{\lambda }=\big(\sum_{i=1}^{N}\widehat{\beta }_{i}^{(2)}\widehat{\beta }
_{i}^{(2)}\big)^{-1}\sum_{i=1}^{N}\widehat{\beta }_{i}^{(1)}\widehat{\beta }
_{i}^{(2)},
\end{equation*}
the second step involves two sample covariances. If one wants to derive the
asymptotic distribution of $\widehat{\lambda },$ one needs to establish the
asymptotic distribution for a properly normalized sample covariance of the
two first step estimators $\frac{1}{N}\sum_{i=1}^{N}\widehat{\beta }
_{i}^{(1)}\widehat{\beta }_{i}^{(2)},$ where $\widehat{\beta }
_{i}^{(j)}=\beta _{i}^{(j)}+\varepsilon _{i}^{(j)},$ with the estimation
error having the structure $\varepsilon _{i}^{(j)}=\big( 1+o_{p}(1)\big)
\frac{1}{T}\sum_{t=1}^{T}v_{t}^{(j)}e_{it}$, the term $o_{p}(1)$ being uniform in $i$.
The normalized sample covariance of the two first step estimators contains
several terms:
\begin{equation*}
\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\left( \widehat{\beta }_{i}^{(1)}\widehat{
\beta }_{i}^{(2)}-\mathbb{E}\big[\widehat{\beta }_{i}^{(1)}\widehat{\beta }
_{i}^{(2)}\big]\right) \hspace{3in}
\end{equation*}
\begin{equation}
=\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\left( \beta _{i}^{(1)}\varepsilon
_{i}^{(2)}+\beta _{i}^{(2)}\varepsilon _{i}^{(1)}\right) +\frac{1}{\sqrt{N}}
\sum_{i=1}^{N}\left( \varepsilon _{i}^{(1)}\varepsilon _{i}^{(2)}-\mathbb{E}
\big[\varepsilon _{i}^{(1)}\varepsilon _{i}^{(2)}\big]\right) . \label{eq: some eq}
\end{equation}
The first term on the right-hand-side of equation (\ref{eq: some eq}) is
similar to a weighted average of first step estimators and has the form of
the first component of $\Xi _{N,T}$ (treating $\beta _{i}^{(j)}$ as
constants similar to constants $\gamma _{i}$). The second term in equation (
\ref{eq: some eq}) is more complicated and calls for a Central Limit Theorem
for quadratic forms:
\begin{equation*}
\frac{T}{\sqrt{N}}\sum_{i=1}^{N}\left( \varepsilon _{i}^{(1)}\varepsilon
_{i}^{(2)}-\mathbb{E}\big[\varepsilon _{i}^{(1)}\varepsilon _{i}^{(2)}\big]\right)
\hspace{3in}
\end{equation*}
\begin{equation*}
=\frac{1}{T\sqrt{N}}\sum_{i=1}^{N}
\sum_{t=1}^{T}v_{t}^{(1)}v_{t}^{(2)}\big(e_{it}^{2}-\mathbb{E}[e_{it}^{2}]\big)+\frac{1
}{\sqrt{N}}\sum_{i=1}^{N}\frac{1}{T}\sum_{s=1}^{T}
\sum_{t<s}w_{st}e_{it}e_{is},
\end{equation*}
where $w_{st}=v_{s}^{(1)}v_{t}^{(2)}+v_{t}^{(1)}v_{s}^{(2)}$. Here the first
term can be treated as the first component, and the second term as the
second component of $\Xi _{N,T}$.
\paragraph{Example 3.}
Anatolyev \& Mikusheva (2018) propose a split-sample estimator as an
alternative to the Fama-MacBeth procedure for factor asset pricing. There
are three sets of parameter estimates produced at the first stage: the
sample average return $\widehat{\beta }_{i}^{(1)}=\frac{1}{T}
\sum_{t=1}^{T}r_{it}$ and two estimates of the portfolio risk exposure
computed as OLS estimates in regressions of $r_{it}$ on $F_{t}$ on different
sub-samples, say, $\widehat{\beta }_{i}^{(2)}$ and $\widehat{\beta }
_{i}^{(3)}$. In our notation, the usage of a sub-sample is accommodated by
setting $v_{t}=0$ for those $t$ not in the currently used sub-sample. The
second step IV estimator is constructed as an IV estimator in the regression
of $\widehat{\beta }_{i}^{(1)}$ on $\widehat{\beta }_{i}^{(2)}$ using $
\widehat{\beta }_{i}^{(3)}$ as instrument. That is,
\begin{equation*}
\widehat{\lambda }=\big(\sum_{i=1}^{N}\widehat{\beta }_{i}^{(3)}\widehat{\beta }
_{i}^{(1)}\big)^{-1}\sum_{i=1}^{N}\widehat{\beta }_{i}^{(3)}\widehat{\beta }
_{i}^{(2)}.
\end{equation*}
In order to make inferences on $\lambda ,$ one needs to obtain the
asymptotic distribution of sample covariances between different first step
estimates. Statements like (\ref{eq: CLT}) and (\ref{eq: variance}) are
instrumental to accomplish this.\bigskip
The configuration in $\Xi _{N,T}$ and a need for statements (\ref{eq: CLT})
and (\ref{eq: variance}) occur in other situations as well.
\paragraph{Example 4.}
Pesaran \& Yamagata (2018) suggest a new test for factor pricing models that
allows many portfolios to be considered simultaneously (with $N$ and $T$
both diverging to infinity). The hypothesis of interest $H_{0}:\alpha _{i}=0$
for all $i=1,...,N$, where $\alpha _{i}$ is a pricing error for the
portfolio $i$. To estimate the pricing errors the authors use OLS estimates $
\widehat{\alpha }_{i}$. A large number of portfolios $N$ does not allow one
to establish join Gaussianity of all $\widehat{\alpha }_{i}$ or to
consistently estimate their covariance. Pesaran \& Yamagata (2018) propose
to test the hypothesis of interest using statistics based on a weighted sum
of squares of $\widehat{\alpha }_{i}$. They create a properly normalized
statistic of the form
\begin{equation*}
\sum_{i=1}^{N}\left( \frac{\widehat{\alpha }_{i}^{2}}{\sigma _{i}^{2}}
-1\right) ,
\end{equation*}
where $\sigma _{i}^{2}$ are variances of pricing errors. This statistic is
directly related to the sample variance of the first step estimator, and a
statement of its asymptotic Gaussianity directly follows from (\ref{eq: CLT}
) by the same logic as stated above. Pesaran \& Yamagata (2018) develop a
CLT for quadratic forms that can be applied in this setting. They make an
assumption that the idiosyncratic components can be filtered to make them
cross-sectionally independent.\footnote{
See Assumptions 2 and 3 in Pesaran \& Yamagata (2018).} Here we propose an
alternative version of CLT that can be applied under less restrictive
assumptions on the cross-sectional dependence of $e_{it}$'s.
\paragraph{Example 5.}
A data-rich IV environment of Bai \& Ng (2010) is another example where our
linear-quadratic CLTs can be useful. The authors consider an IV setup with
many instruments in a panel, where the number of instruments, $N$, is
potentially higher than the number of observations, $T$. The instruments are
generated by a factor model $z_{it}=\lambda _{i}^{\prime }F_{t}+e_{it},$
with $F_{t}$ and $e_{it}$ independent of the structural error $\varepsilon
_{t},$ and time-series and cross-sectional dependence in $e_{it}$ is
allowed, though restricted. Bai \& Ng (2010) consider the bias-corrected GMM
estimator that corrects for inconsistency of the baseline GMM. It is
consistent when $N/T=O(1),$ however, its asymptotic Gaussianity is
established under a more restrictive assumption when $N/T=o(1).$ The
challenge is that when $N/T=O(1),$ the asymptotic expansion for this
estimator has, in addition to a linear term, a quadratic form in the
idiosyncratic components $e_{it}$ similar to the second component of $\Xi
_{N,T}$. Thus, using statement (\ref{eq: CLT}), an asymptotic theory could
be developed for the bias-corrected GMM estimator without having to impose $
N/T=o(1)$.\bigskip
In most of these examples, the set of idiosyncratic components $\{e_{it},$ $
i=1,...,N,$ $t=1,...,T\}$ cannot be regarded independent and/or identically
distributed. In most realistic applications, one is usually willing to
assume that $e_{it}$ do not have a strong (detectable) factor structure, but
still allow for some correlation between different units, which would not
affect consistency. For example, it is reasonable to think that stocks
of firms in the same industry or of the same size may react to some local
shocks and be correlated, though when averaged over all stocks (and all
industries), this co-movement of returns would have no first-order impact on
estimation.
Our attempt to be agnostic with regard to possible
cross-sectional correlation among errors and to avoid explicit modeling of
its structure whenever possible comes at a cost of more restrictive time series assumptions. In many applications of interest, it is
more credible to impose independence assumptions in a time-series
direction rather than in a cross-sectional direction. For example, the
efficient market hypothesis implies mean non-predictability of excess
returns given past history, which is equivalent to a martingale difference
property for the errors. The definition of shocks in macroeconomics
similarly presumes their time-series independence. In this paper, we assume
time-series independence, which in some cases may be weakened to the
martingale difference property or stationarity with some proper mixing
condition, but we do not pursue this generalization here.
\section{CENTRAL\ LIMIT\ THEOREM}
\label{section-CLT}
In this paper we consider asymptotics as both cross-sectional and
time-series sample sizes, $N$ and $T$, increase to infinity. We allow the
data-generating process for all variables to vary with $N$ and $T$. Define $
\mathcal{F}$ to be a $\sigma$-algebra that contains at least the $\sigma$-algebras
generated by the full set of variables $\{v_{s},s=1,...,\infty\}$ and $
\{w_{st}, s,t=1,...,\infty\}$ for all $s $ and $t$. It may potentially also
contain other events related to common shocks and variables, as long as
Assumption \ref{ass: errors} stated below is satisfied. We treat $\gamma
_{i} $ as non-random $k_{\gamma }\times 1$ vectors.
In order to simplify the notation, in what follows we will denote $C$ to be
a positive generic constant, independent of $N$ and $T,$ which may be
different in different equations, but does not depend on or change with $N$
or $T$. We will use the following notation: for a square matrix $A$, we
denote by $\mathrm{tr}(A)$ its trace, by $\max \mathrm{ev}(A)$ -- its
maximal eigenvalue, and by $\mathrm{dg}(A)$ a diagonal matrix of the same
size with the elements from the diagonal of $A$; $\Vert \cdot \Vert $ is the
$l_{2}$ norm for a vector or the operator norm for a matrix.
\begin{ass}
\label{ass: common vars} The random $k_{v}$-vector $v_{s}$ and $k_{w}$-vector $w_{st}$ are
measurable with respect to the $\sigma$-algebra $\mathcal{F}$ for all $s,t$, and
\begin{itemize}
\item[(i)] $\frac{1}{T}\sum_{s=1}^{T}\mathbb{E}\left( v_{s}v_{s}^{\prime
}\right) \rightarrow \Omega _{v}$ and $\frac{1}{T^2}\sum_{s=1}^{T}\sum_{t<s}
\mathbb{E}\left( w_{st}w_{st}^{\prime }\right) \rightarrow \Omega _{w}$,
where $\Omega _{v}$ and $\Omega _{w}$ are full rank matrices;
\item[(ii)] $\max_{1\leq s\leq T}\mathbb{E}\big[ \Vert v_{s}\Vert ^{4}
\big] <C$ and $\max_{1\leq t,s\leq T}\mathbb{E}\big[ \Vert w_{st}\Vert
^{4}\big] <C$;
\item[(iii)] $\mathbb{E}\left[ \left\Vert \frac{1}{T^2}\sum_{s=1}^{T}
\sum_{t<s}\big(w_{st}w_{st}^{\prime }-\mathbb{E}[w_{st}w_{st}^{\prime
}]\big)\right\Vert ^{2}\right] \rightarrow 0;$
\item[(iv)] $\mathbb{E}\left[ \left\Vert \frac{1}{T}
\sum_{s=1}^{T}\big(v_{s}v_{s}^{\prime }-\mathbb{E}[v_{s}v_{s}^{\prime
}]\big)\right\Vert ^{2}\right] \rightarrow 0.$
\end{itemize}
\end{ass}
\begin{ass}
\label{ass: loadings} $\max_{1\leq i\leq N}\Vert \gamma _{i}\Vert <C$.
\end{ass}
\begin{ass}
\label{ass: errors}
\begin{itemize}
\item[(i)] Conditional on $\mathcal{F}$, the random $N$-vectors $
e_{t}=(e_{1t},...,e_{Nt})^{\prime }$ are serially independent, and $\mathbb{E
}(e_{t}|\mathcal{F})=0$ for all $t$;
\item[(ii)] $\max_{1\leq i\leq N,1\leq t\leq T}\mathbb{E}\left(
e_{it}^{4}\right) <C$.
\end{itemize}
\end{ass}
Assumption \ref{ass: common vars} imposes very mild restrictions on the
time-series behavior of the common (non-entity specific) variables. For
example, the part related to $v_{t}$ is trivially satisfied if a time series
equal to $v_{t}v_{t}^{\prime }$ is weakly stationary with summable
auto-covariances. Assumption \ref{ass: loadings} restricts the influence of
any one entity in the cross-sectional average and will eventually contribute
to asymptotic negligence of the cross-sectional summands needed for the CLT.
Assumption \ref{ass: errors}(i) is a restrictive assumption which imposes
discipline on the time-series structure, and the restriction $\mathbb{E}
(e_{t}|\mathcal{F})=0$ is a form of strict exogeneity in the first step
regression. Uniform moment boundedness in Assumption \ref{ass: errors}(ii)
is traditional.
Apparently, Assumptions \ref{ass: common vars}, \ref{ass: loadings} and \ref
{ass: errors} are insufficient to establish a central limit theorem, and we
need to put some restrictions on the cross-sectional dependence and
dependence between idiosyncratic errors and common variables. Indeed, we
will use a change of summation ordering:
\begin{equation*}
\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\xi _{i}=\left(
\begin{array}{c}
\frac{1}{\sqrt{T}}\sum_{s=1}^{T}v_{s}\left( \frac{1}{\sqrt{N}}
\sum_{i=1}^{N}\gamma _{i}e_{is}\right) \\
\frac{1}{T}\sum_{s=1}^{T}\sum_{t<s}w_{st}\left( \frac{1}{\sqrt{N}}
\sum_{i=1}^{N}e_{it}e_{is}\right)
\end{array}
\right) ,
\end{equation*}
and establish asymptotic convergence in the time-series direction. In
order to apply a CLT in the time series direction we need some sort of asymptotic
negligibility of summands with different time indexes, in particular, of
terms like $\big\{ v_{s}(\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\gamma
_{i}e_{is})\big\} _{s}$ and $\big\{ w_{st}(\frac{1}{\sqrt{N}}
\sum_{i=1}^{N}e_{it}e_{is})\big\} _{s,t}$. Our goal is to provide
low-level assumptions. There is a trade-off in how much dependence of
idiosyncratic errors across entities and how much dependence between
idiosyncratic errors and common variables can be allowed. Below we consider two
particular cases. In the first case, full independence between the $e_{it}$
's and $\mathcal{F}$ is assumed; as a result, we can be agnostic about the
structure of cross-sectional dependence, the corresponding assumptions about
it are relatively mild. In the second case, we allow for conditional
heteroscedasticity in $e_{it}$ that can be related to some common variables
from $\mathcal{F}$ producing dependence in higher-order conditional moments.
This flexibility comes at the cost of imposing some structure on the
cross-sectional behavior of~$e_{it}$.
\subsection{Independence from common variables}
\begin{ass}
\label{ass: independence}
\begin{itemize}
\item[(i)] The errors $e_{t}=(e_{1t},...,e_{Nt})^{\prime },$ $t=1,\ldots ,T$
are independent from the $\sigma$-algebra $\mathcal{F}$ and identically
distributed across $t$;
\item[(ii)] For the $N\times N$ covariance matrix $\mathcal{E}_{N,T}=\mathbb{
E}\left( e_{t}e_{t}^{\prime }\right) $, $\limsup_{N,T\rightarrow \infty
}\max \mathrm{ev}\left( \mathcal{E}_{N,T}\right) <\infty ,$ and $\frac{1}{N}
\mathrm{tr}\big(\mathcal{E}_{N,T}^{2}\big)\rightarrow a<\infty$;
\item[(iii)] $\frac{1}{N}\gamma ^{\prime }\mathcal{E}_{N,T}\gamma
\rightarrow \Gamma _{\sigma }$, where $\Gamma _{\sigma }$ is full rank;
\item[(iv)] $\frac{1}{N^{2}}\sum_{i_{1}=1}^{N}\sum_{i_{2}=1}^{N}
\sum_{i_{3}=1}^{N}\sum_{i_{4}=1}^{N}\left\vert \mathbb{E}\left(
e_{i_{1}t}e_{i_{2}t}e_{i_{3}t}e_{i_{4}t}\right) \right\vert <C.$
\end{itemize}
\end{ass}
\begin{theorem}
\label{the: gaussianity for independence} Under Assumptions \ref{ass: common
vars}, \ref{ass: loadings}, \ref{ass: errors} and \ref{ass: independence},
the central limit theorem stated in equation (\ref{eq: CLT}) holds with $\Sigma_{\xi}=\left(
\begin{array}{cc}
\Sigma_V & 0 \\
0 & \Sigma_W \\
\end{array}
\right)
$, where $\Sigma_V=\Gamma_\sigma\otimes\Omega_v$ and $\Sigma_W=a\Omega_w.$
\end{theorem}
Numerous papers that establish inferences in factor models commonly assume
that the set of factors is independent from the set of idiosyncratic errors,
as in Assumption \ref{ass: independence}(i), though cross-sectional
dependence of errors is allowed; see, for example, Assumption D in Bai \& Ng
(2006). We intended for the first part of Assumption \ref{ass: independence}
(ii) to impose weak cross-sectional dependence as expressed by the
covariance matrix; in particular, it means that no strong factor structure
is left in the errors; similar assumptions appear in Onatski (2012) and Bai
\& Ng (2006). The convergence of the trace in Assumptions \ref{ass:
independence}(ii) and \ref{ass: independence}(iii) is needed for the
asymptotic covariance matrix to be properly defined.
Assumption \ref{ass: independence}(iv) is another way to restrict pervasive
dependence in multiple variables, in particular, precluding outliers to
realize in too many error terms simultaneously. For example, imagine that
the cross-sectional dependence is induced by several groups with a factor
structure, e.g., stock returns are correlated because there are
industry-specific shocks and geography-specific shocks. Imagine that there
are a finite number, say $G$, groups, indexed by $g=1,...,G$, \ which may be
overlapping, with each having independent shocks $f_{g,t}$ at time $t$.
Stock $i$ has non-zero loading $\pi _{i,g}$ only if it belongs to group $g$.
Let the set of groups, to which $i$ belongs, be denoted by $G(i)$. That is,
\begin{equation*}
e_{it}=\sum_{g\in G(i)}\pi _{i,g}f_{g,t}+\eta _{it},
\end{equation*}
where $\eta _{it}$'s are independent both cross-sectionally and across time
and have finite fourth cumulants. Then, Assumption \ref{ass: independence}
(iv) is essentially equivalent to the following two conditions: $\mathbb{E}
(f_{g,t}^{4})<C$ and $\frac{1}{N}\big( \sum_{i=1}^{N}|\pi _{i,g}|\big)
^{2}<C$ for any $g=1,...,G$. Thus,
for this example, essentially Assumption \ref{ass: independence}(iv)
imposes that the factors $f_{g,t}$ do not produce outliers too often
expressed as the moment condition and a statement about pervasiveness.
One of the important steps in the proof of Theorem \ref{the: gaussianity for
independence} verifies asymptotic negligibility of time-series summands by
checking boundedness of the fourth moments of the cross-sectional sums $
\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\gamma _{i}e_{is}$ and $\frac{1}{\sqrt{N}}
\sum_{i=1}^{N}e_{it}e_{is}$; that imposes the main way we restrict
cross-sectional dependence. The fourth cumulant conditions are reminiscent
of those in de Jong (1987), which we follow while proving our CLT using Heyde
\& Brown (1970). There are alternative CLTs for quadratic forms such as
Rotar' (1973) that imposes weaker moment conditions on the summands but
stricter assumptions on the negligibility of coefficients and eigenvalues of
the quadratic form. In our case, following them would require imposing
stronger assumptions on the variables $w_{st},$ which we would like to avoid. Another CLT for quadratic forms for time series data can be obtained using Bhansali et al. (2007).
The book by Giraitis et al. (2012) has a chapter on this subject and allows for long memory time series as well.
\subsection{Conditional heteroscedasticity}
Assumption \ref{ass: independence}(i) of independence is much stronger than
Assumption \ref{ass: errors}(i) about exogeneity: it does not allow higher
conditional moments of $e_{it}$ to co-move with the common variables; in
particular, it imposes conditional homoscedasticity. It may be especially
problematic in financial applications where time-varying volatility is of
strong empirical relevance, and returns on many stocks display patterns of
changing volatility driven by some common variables. The assumptions below
allow for conditional heteroscedasticity.
\begin{ass}
\label{ass: heterosk} The errors $e_{it}$ have the following
weak (unobserved) factor structure:
\begin{equation*}
e_{it}=\pi _{i}^{\prime }f_{t}+\eta _{it},
\end{equation*}
where the following assumptions hold:
\begin{itemize}
\item[(i)] The $k_{f}\times 1$ process $f_{t}$, where $k_{f}$ is fixed, is serially independent,
conditionally on $\mathcal{F}$, with $\mathbb{E}(f_{t}|\mathcal{F})=0$, $
\mathbb{E}(f_{t}f_{t}^{\prime })=I_{k_{f}},$ $\max_{1\leq t,s\leq T}\mathbb{E
}\big[ (\Vert v_{s}\Vert ^{4}+1)\Vert f_{t}\Vert ^{4}\big] <C,$ and $
\max_{1\leq s,t,t^{\ast }\leq T}\mathbb{E}\big[ \Vert w_{st^{\ast }}\Vert
^{4}\Vert f_{t}\Vert ^{8}\big] <C$;
\item[(ii)] $\max \mathrm{ev}\big( \sum_{i=1}^{N}\pi _{i}\pi _{i}^{\prime
}\big) <C$ and $\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\pi _{i}\gamma
_{i}^{\prime }\rightarrow \Gamma _{\pi \gamma }$;
\item[(iii)] The random variables $\eta _{it}$ are independent both
cross-sectionally and across time, independent from both $f_{s}$'s and $
\mathcal{F}$, have mean zero and variances $\mathrm{var}(\eta _{it})=\omega
_{i}^{2}$ that are bounded from above and such that $\frac{1}{N}
\sum_{i=1}^{N}\omega _{i}^{4}\rightarrow \omega ^{4}<\infty$, $\frac{1}{N}
\sum_{i=1}^{N}\omega _{i}^{2}\gamma _{i}\gamma _{i}^{\prime }\rightarrow
\Gamma _{\omega }$, where $\Gamma _{\omega }$ is finite and has full rank, and $
\max_{1\leq i\leq N,1\leq t\leq T}\mathbb{E}\left( \eta _{it}^{4}\right) <C$;
\item[(iv)] Additionally, if $\Gamma _{\pi \gamma }\neq 0$, then there
exists a matrix $\Sigma _{fv}$ such that
\begin{equation*}
\mathbb{E}\left[ \left\Vert \frac{1}{T}\sum_{s=1}^{T}(f_{s}f_{s}^{\prime
})\otimes (v_{s}v_{s}^{\prime })-\Sigma _{fv}\right\Vert ^{2}\right]
\rightarrow 0.
\end{equation*}
\end{itemize}
\end{ass}
\begin{theorem}
\label{th: CLT for HC} Under Assumptions \ref{ass: common vars}, \ref{ass:
loadings}, \ref{ass: errors} and \ref{ass: heterosk}, the statement of the
central limit theorem stated in equation (\ref{eq: CLT}) holds with $\Sigma_{\xi}=\left(
\begin{array}{cc}
\Sigma_V & 0 \\
0 & \Sigma_W \\
\end{array}
\right)
$, where $\Sigma_W=\omega^4\Omega_w$ and $\Sigma_V=\left(\Gamma_{\pi\gamma}^\prime \otimes I_{k_v}\right)\Sigma_{fv}\left(\Gamma_{\pi\gamma} \otimes I_{k_v}\right)+\Gamma_\omega\otimes\Omega_v$.
\end{theorem}
An interesting feature of this example is that it allows the errors to be
weakly cross-sectionally dependent to the extent that they may possess a
weak (latent) factor structure. The condition $\mathbb{E}(f_{t}f_{t}^{\prime
})=I_{k_{f}}$ is a normalization and involves no loss of generality.
Assumption \ref{ass: heterosk}(ii) forces the factors to be weak to such an
extent that the factor structure cannot be consistently detected; it implies
that the covariance matrix of idiosyncratic errors would satisfy the first
half of Assumption \ref{ass: independence}(ii). Moreover, this factor
structure may be closely related to the common variables in $\mathcal{F}$,
which causes the cross-sectional dependence among the errors $e_{it}$ to
change with the common variables and allows a very flexible form of
conditional heteroscedasticity. Indeed, the conditional cross-sectional
covariance is
\begin{equation*}
\mathbb{E}(e_{it}e_{jt}|\mathcal{F})=\pi _{i}^{\prime }\mathbb{E}
(f_{t}f_{t}^{\prime }|\mathcal{F})\pi _{j}+\mathbb{I}_{\{i=j\}}\omega
_{i}^{2}.
\end{equation*}
Since we do not restrict $\mathbb{E}(f_{t}f_{t}^{\prime }|\mathcal{F})$
beyond proper moment conditions, the strength of any cross-sectional
dependence as well as error variances may change stochastically depending on
realizations of the common variables.
The moment conditions in Assumption \ref{ass: heterosk}(i) help to establish
asymptotic negligibility of the time-series summands. Assumption \ref{ass:
heterosk}(iii) about $\Gamma _{\omega }$ and Assumption \ref{ass: heterosk}
(iv) allow us to define properly the asymptotic covariance matrix.
\section{VALID\ INFERENCE}
\label{section-variance}In this Section, we first discuss estimation of
asymptotic variances for asymptotic inference when this leads to valid
inference. Then, we propose alternative tools based on the wild bootstrap to
apply in situations when asymptotic inference fails to provide
asymptotically correct inference.
\subsection{Asymptotic inference}
Statistical inferences such as confidence set construction and hypotheses
testing about the structural parameter typically require consistent
estimation of asymptotic variances of all important quantities that are
asymptotically Gaussian. The easiest to implement and thus the most
appealing from an applied perspective are those that use the same variables
and have a structure similar to the original averages, such as the statement
in equation (\ref{eq: variance}).
Notice that equation (\ref{eq: variance}) contains the cross-sectional
summation outside, and hence it treats the cross-section as nearly
uncorrelated observations, or at least it ignores the cross-sectional
correlation. A relevant analogue is the difference between the long-run
covariance and instantaneous covariance in a classical time series. However,
implementing an analogue of long-run covariance estimation here would be a
challenge since we do not have any cross-sectional stationarity or a measure
of distance between cross-sectional entities. Rather, we explore under which
conditions the convergence in (\ref{eq: variance}) holds.
Theorem \ref{th: covariance for independence} below obtains a statement for
the case when the common variables are independent from the idiosyncratic
errors, while Theorem \ref{th: covariance for HC} establishes a similar
statement for the conditionally heteroscedastic case.
\begin{theorem}
\label{th: covariance for independence} If in addition to Assumptions \ref
{ass: common vars}, \ref{ass: loadings}, \ref{ass: errors}, \ref{ass:
independence} we also have that
\begin{equation} \label{eq: almost diagonal}
\Vert \mathcal{E}_{N,T}-\mathrm{dg}(\mathcal{E}_{N,T})\Vert \rightarrow
0\quad \text{as }N,T\rightarrow \infty ,
\end{equation}
then consistency statement (\ref{eq: variance}) holds.
\end{theorem}
\begin{theorem}
\label{th: covariance for HC} If in addition to Assumptions \ref{ass: common
vars}, \ref{ass: loadings}, \ref{ass: errors}, \ref{ass: heterosk} we also
have that $\Gamma _{\pi \gamma }=0$, then consistency statement (\ref{eq:
variance}) holds.
\end{theorem}
The additional assumption (\ref{eq: almost diagonal}) in Theorem \ref{th:
covariance for independence} strengthens conditions on the weakness of the
cross-sectional correlation; in particular, it requires that the covariance
matrix converges to a diagonal one. The additional assumption in Theorem \ref
{th: covariance for HC} requires that the weights used for averaging the
cross-sectional entities are orthogonal to the loadings on the latent factor
structure, which precludes the latent factor structure (that represents the
cross-sectional dependence) from being amplified. This is a necessary
assumption for consistency of the variance estimator. Indeed, let
assumptions \ref{ass: common vars}, \ref{ass: loadings}, \ref{ass: errors},
\ref{ass: heterosk} hold, and consider the first component of $\xi _{i}$:
\begin{equation*}
\xi _{i}^{(1)}=\frac{1}{\sqrt{T}}\sum_{t}v_{t}\gamma _{i}e_{it}=\pi
_{i}\gamma _{i}^{\prime }\Upsilon _{T}+\tilde{\eta}_{i},
\end{equation*}
where $\tilde{\eta}_{i}=\frac{1}{\sqrt{T}}\sum_{t}v_{t}\gamma _{i}\eta _{it}$
, and $\Upsilon _{T}=\frac{1}{\sqrt{T}}\sum_{t=1}^{T}f_{t}v_{t}$. Note that $
\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\tilde{\eta}_{i}\Rightarrow \mathcal{N}
(0,\sigma _{\eta }^{2})$ and $\frac{1}{N}\sum_{i=1}^{N}\tilde{\eta}_{i}^{2}
\overset{p}{\rightarrow }\sigma _{\eta }^{2},$ as all conditions of Theorems
\ref{th: CLT for HC} and \ref{th: covariance for HC} are satisfied by
cross-sectionally and time independent errors $\eta _{is}$. Assumption \ref
{ass: heterosk}(iv) guarantees that $\Upsilon _{T}\Rightarrow \mathcal{N}
(0,\Sigma _{fv})$ as $T\rightarrow \infty $, while according to Assumption
\ref{ass: heterosk}(ii), we have $\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\pi
_{i}\gamma _{i}^{\prime }\rightarrow \Gamma _{\pi \gamma }$ as $N\rightarrow
\infty $. Thus,
\begin{equation*}
\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\xi _{i}^{(1)}\Rightarrow \mathcal{N}
\big(0,\Gamma _{\pi \gamma }\Sigma _{fv}\Gamma _{\pi \gamma }^{\prime }+\sigma
_{\eta }^{2}\big),
\end{equation*}
while $\frac{1}{N}\sum_{i=1}^{N}\big(\xi _{i}^{(1)}\big)^{2}\overset{p}{\rightarrow }
\sigma _{\eta }^{2}$ because $\frac{1}{N}\sum_{i=1}^{N}\pi _{i}\gamma
_{i}^{\prime }\gamma _{i}\pi _{i}^{\prime }\rightarrow 0$ by Assumptions \ref
{ass: loadings} and \ref{ass: heterosk}(ii).
\subsection{Bootstrap inference}
As one way to conduct valid inferences in settings when $\frac{1}{N}
\sum_{i=1}^{N}\xi _{i}\xi _{i}^{\prime }$ is an inconsistent estimator of
the variance (in particular, when $\Gamma _{\pi \gamma }\neq 0$ under
Assumption \ref{ass: heterosk}), we propose the following simple wild
bootstrap procedure. In each bootstrap repetition,
\begin{enumerate}
\item[(i)] simulate independent draws of random variables $\delta _{t}\sim
\{+1,-1\},$ $t=1,...,T,$ with success probability 1/2;
\item[(ii)] compute the bootstrap analogues of errors $e_{it}^{\ast }=\delta
_{t}e_{it},$ $i=1,...,N,$ $t=1,...,T$;
\item[(iii)] compute the bootstrap analogues $\xi _{i}^{\ast }$ using the
definition of $\xi _{i},$ with $e_{it}^{\ast }$ in place of $e_{it}$ for all
$i=1,...,N,$ $t=1,...,T.$
\end{enumerate}
\noindent Then the distribution of $\Xi _{N,T}^{\ast }=\frac{1}{\sqrt{N}}
\sum_{i=1}^{N}\xi _{i}^{\ast }$ has the same asymptotic limit as that of $
\Xi _{N,T}$ has, and can be used for inferences. Alternatively, the
distribution of $\Xi _{N,T}^{\ast }$ normalized by the bootstrap analogue of
the variance estimate $\frac{1}{N}\sum_{i=1}^{N}\xi _{i}^{\ast }\xi
_{i}^{\ast \prime }$ can be used to approximate the distribution of $\Xi
_{N,T}$ normalized by the variance estimate $\frac{1}{N}\sum_{i=1}^{N}\xi
_{i}\xi _{i}^{\prime }.$ We call the two described bootstrap procedures
\textit{bootstrap} and \textit{bootstrap-t}. In the next Section, we
implement both variations of the wild bootstrap for the setup of Assumption
\ref{ass: heterosk}.
The wild bootstrap works because it introduces independence in the time
direction while preserving the (unknown) cross-sectional dependence.
Specifically, we base our proof of Theorem \ref{th: CLT for HC} on the
change of order of summations in double/triple summations over $i$ and time index/indices.
For example,
\begin{equation*}
\Xi _{N,T}^{(1)}=\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\xi _{i}^{(1)}=\frac{1}{
\sqrt{T}}\sum_{t=1}^{T}\left( \frac{1}{\sqrt{N}}\sum_{i=1}^{N}v_{t}\gamma
_{i}e_{it}\right) .
\end{equation*}
We then argue that our assumptions guarantee that the CLT with respect to
summation over $t$ is applicable. If the assumptions of the CLT hold, then $
\Xi _{N,T}^{(1)}$ is asymptotically Gaussian with mean zero and variance
equal to the limit of $\frac{1}{T}\sum_{t=1}^T\mathbb{E}\big[\big(\frac{1}{\sqrt{N}}
\sum_{i}v_{t}\gamma _{i}e_{it}\big)^{2}\big]$. This limit clearly depends on how much
there is cross-sectional correlation between $e_{it}$ and $e_{jt}$. Now, in
the bootstrapped samples,
\begin{equation*}
\big(\Xi _{N,T}^{(1)}\big)^{\ast }=\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\big(\xi
_{i}^{(1)}\big)^{\ast }=\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\left( \frac{1}{\sqrt{N}}
\sum_{i=1}^{N}v_{t}\gamma _{i}e_{it}\right) \delta _{t}.
\end{equation*}
Conditional on the original sample, only $\delta _{t}$'s are random, and
they are independent and have zero mean and unit variance. When $T$ is
large, this bootstrapped normalized sum satisfies the CLT and thus converges
to a zero mean Gaussian random variable with the variance equal to the limit
of $\frac{1}{T}\sum_{t=1}^{T}\big(\frac{1}{\sqrt{N}}\sum_{i=1}^{N}v_{t}\gamma
_{i}e_{it}\big)^{2}$ as $N,T\rightarrow \infty $. This limit coincides with $
\frac{1}{T}\sum_{t=1}^{T}\mathbb{E}\big[\big(\frac{1}{\sqrt{N}}\sum_{i=1}^{N}v_{t}
\gamma _{i}e_{it}\big)^{2}\big]$ under relatively weak assumptions, as long as the
Law of Large Numbers holds. For example, under Assumptions \ref{ass:
heterosk}, we have:
\begin{eqnarray*}
\lim \frac{1}{T}\sum_{t=1}^{T}\left( \frac{1}{\sqrt{N}}\sum_{i=1}^{N}v_{t}
\gamma _{i}e_{it}\right) ^{2} &=&\lim \frac{1}{T}\sum_{t=1}^{T}\left( \Gamma
_{\pi \gamma }v_{t}f_{t}+\frac{1}{\sqrt{N}}\sum_{i=1}^{N}v_{t}\gamma
_{i}\eta _{it}\right) ^{2} \\
&&\overset{p}{\rightarrow }\Gamma _{\pi \gamma }\Sigma _{fv}\Gamma _{\pi
\gamma }^{\prime }+\sigma _{\eta }^{2}.
\end{eqnarray*}
A similar argument can be made about the second component as well. Indeed,
\begin{equation*}
\big(\Xi _{N,T}^{(2)}\big)^{\ast }=\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\big(\xi
_{i}^{(2)}\big)^{\ast }=\frac{1}{T}\sum_{s=1}^{T}\sum_{t<s}\left( \frac{1}{\sqrt{
N}}\sum_{i=1}^{N}w_{st}e_{it}e_{is}\right) \delta _{t}\delta _{s}.
\end{equation*}
For a bootstrapped statistic, conditional on the original sample, the
weights $\frac{1}{\sqrt{N}}\sum_{i=1}^{N}w_{st}e_{it}e_{is}$ are fixed,
while $\delta _{t}$'s are independent random variables, and all the
conditions of Lemma \ref{th: Heyde and brown} are satisfied. Thus, a zero
mean Gaussian limit obtains as $T\rightarrow \infty $. It is straightforward
to verify that its variance converges to the asymptotic variance of $\Xi
_{N,T}^{(2)}$.
In the simulation experiments in the following Section, we check, among
other things, that both wild bootstrap variations deliver correctly sized
tests even when the asymptotic-t tests fail to do so, and also that the
bootstrap-based tests have a non-trivial power.
\section{MONTE CARLO SIMULATIONS}
\label{section-simulations}The goals of this Section are to check finite
sample performance of an asymptotic Gaussian approximation for $\Xi _{N,T}$,
to explore when the variance estimator $\widehat{\Sigma }_{\xi }=\frac{1}{N}
\sum_{i=1}^{N}\xi _{i}\xi _{i}^{\prime }$ allows to construct reliable
asymptotic-t inferences, and to evaluate the performance of the wild
bootstrap and bootstrap-t procedures in terms of both size and power.
\subsection{Setup}
Our setup adheres to Assumption \ref{ass: heterosk}. We generate the errors $
e_{it}$ according to the following weak (unobserved) factor structure:
\begin{equation*}
e_{it}=\pi _{i}^{\prime }f_{t}+\eta _{it},
\end{equation*}
where $f_{t}\sim iid$ $\mathcal{N}(0,1)$ across $t=1,...,T,$ $\gamma _{i}=1$
for all $i=1,...,N,$ and $\eta _{it}=\omega _{i}\epsilon _{\eta ,it},$ where
$\epsilon _{\eta ,it}$ are $iid$ $\mathcal{N}(0,1)$ across $i$ and $t$. The
standard deviations are set to $\omega _{i}=c_{\omega }\left( 1+\left\vert
\tau _{i}\right\vert \right) ,$ and the factor loadings are $\pi _{i}=\left(
c_{\pi }+\tau _{i}\right) /\sqrt{N},$ where $\tau _{i}\sim iid$ $\mathcal{N}
(0,1)$ across $i=1,...,N.$ The multiplier $c_{\omega }$ is tuned so that the
average cross-sectional variance of $\eta _{it}$ is unity. The parameter $
c_{\pi }$ indexes the degree of cross-sectional dependence as measured by
the strength of the factor structure. Specifically, as $\sum_{i=1}^{N}\pi
_{i}\pi _{i}^{\prime }\rightarrow 1+c_{\pi }^{2}$, this parameter is assumed
to be bounded for the Gaussian approximation to hold. We also have $\Gamma
_{\pi \gamma }=c_{\pi }$, hence we can expect consistency of variance
estimation only when $c_{\pi }=0$, so we will explore the distortions for
different values of $c_{\pi }$. The errors generated this way are
cross-sectionally dependent and heteroscedastic, while still satisfying
Assumption \ref{ass: heterosk}.
The common variables are generated as follows: $v_{t}=c_{fv}f_{t}+\sqrt{
1-c_{fv}^{2}}\epsilon _{v,t},$ with $\epsilon _{v,t}\sim iid$\ $\mathcal{N}
(0,1)$ across $t=1,...,T,$ and $w_{st}=v_{s}v_{t},$ $s,t=1,...,T.$ All the
disturbances $f_{t},$ $\tau _{i},$ $\epsilon _{\eta ,it}$ and $\epsilon
_{v,t}$ are mutually independent. The parameter $c_{fv}$ indexes the
dependence between common variables and $e_{it}$. The mean zero Assumption
\ref{ass: errors}(i) requires $c_{fv}=0$; the non-zero values of $c_{fv}$
index deviations from the null hypothesis $\mathbb{E}(\xi _{i})=0,$ and will
be used to study the power properties of the proposed wild bootstrap. In
wild bootstrap samples, we generate the bootstrap errors by $e_{it}^{\ast
}=\delta _{t}e_{it},$ where $\delta _{t}=2\zeta _{t}-1,$ and $\zeta _{t}\sim
iid\ \mathcal{B}(\frac{1}{2})$ across $t,...,T$. The bootstrap analogues of $
\gamma _{i},$ $v_{t}$ and $w_{st}$ are set equal to their original sample
values.
The distribution characteristics are computed from 10,000 simulations, while
the rejection rates are based on 5,000 simulation runs. In all simulations,
we set $N=T=500.$ The number of bootstrap repetitions is 600.
\subsection{Results}
Table \ref{tab: table 1} contains distributional characteristics of $\Xi
_{N,T}$. We report averages, coefficients of skewness, coefficients of
kurtosis, and right 5\% quantiles of normalized marginal distributions of
both elements of $\Xi _{N,T}.$ For the exactly normal distributions, these
values are $0$, $0$, $3$ and $1.645,$ respectively.
\begin{table}[h]
\caption{\label{tab: table 1} Characteristics of simulated finite-sample distribution of $\Xi_{N,T}$ for $N=T=500$}\medskip
\begin{centering}
\begin{tabular}{ccccccccc}
\toprule
element of $\Xi_{N,T}$$\,\rightarrow $\quad & \multicolumn{4}{c}{linear} & \multicolumn{4}{c}{
quadratic} \\
\midrule
$c_{\pi }$ $\downarrow $ & mean & skew & kurt & quant & mean & skew & kurt &
quant \\
\midrule
0.5 & $0.01$ & $-0.05$ & $3.01$ & $1.63$ & $-0.01$ & $0.25$ & $3.20$ & $1.71$
\\
1 & $0.01$ & $-0.04$ & $2.97$ & $1.64$ & $-0.01$ & $0.25$ & $3.20$ & $1.71$
\\
2 & $0.00$ & $-0.00$ & $2.95$ & $1.65$ & $0.02$ & $0.20$ & $3.11$ & $1.71$
\\
\bottomrule
\end{tabular}
\par\end{centering}
\medskip
{\footnotesize{\emph{Notes: }Based on 10,000 simulations. The first component of $\Xi_{N,T}$ is
labeled `linear' and second component is labeled `quadratic'. `Mean' stands for average, `skew' for skewness coefficient, `kurt' for kurtosis coefficient, and `quant' for right 5\% quantile of
simulated marginal distribution of components of $\Xi _{N,T}$ normalized by
corresponding standard deviations. Rows with $c_{\pi }=0.5,$ $1$ and $2$
correspond to a very weak, weak and moderately strong error factor structure. }}
\end{table}
The actual distribution of the linear component of $\Xi _{N,T}$ is very
close to Gaussian, in all respects: all the moments and the right tail are
almost equal to their theoretical counterparts. The quadratic component of $
\Xi _{N,T},$ however, albeit mean unbiased, is somewhat positively skewed
and a bit leptokurtic. The shifted right quantile confirms slight
over-dispersion. The distortions, however, do not seem to increase with the
strength of the error factor structure.
In Table \ref{tab: table 2} we document the empirical rejection rates for
tests with $10\%,$ $5\%$ and $1\%$ declared size based on the asymptotic-t,
wild bootstrap and wild bootstrap-t approaches. In the asymptotic-t approach
we create a $t$-statistic using $\widehat{\Sigma }_{\xi }$ as a variance
estimator and compare it with the symmetric standard Gaussian critical
values. We also explore the performance of two wild bootstrap procedures --
one that bootstraps $\Xi _{N,T}$ and another that bootstraps the $t$
-statistics (referred to in Table \ref{tab: table 2} as bootstrap $\xi $ and
bootstrap $t$). In both bootstrap procedures, we compare the absolute value
of the statistic from the sample to the right quantile of the absolute value
of the bootstrapped statistic. We verify the empirical size by setting $
c_{fv}=0$ and empirical power by setting $c_{fv}=0.1$ for relatively small
deviations from the null and $c_{fv}=0.2$ for relatively large deviations
from the null.
As expected, the size of the test based on asymptotic approximations
sometimes deviates from nominal rates by a wide margin, especially for the
linear component of $\xi _{N,T},$ the gap quickly increasing with the
strength of the error factor structure. This happens due to inconsistency of
the variance estimator $\widehat{\Sigma }_{\xi }$ and becomes more pronounced
with stronger cross-sectional dependence. What is surprising is that for the
quadratic component of $\xi _{N,T},$ the distortions are relatively minor and
not very sensitive to the strength of the factor structure. In contrast,
both bootstrap procedures exhibit excellent
size control and stability thereof across the strength of the error factor
structure for both components of $\xi _{N,T}.$ In terms of power, however,
the two bootstrap statistics are approximately equally powerful for the
linear component of $\xi _{N,T},$ while there is a gap, sometimes sizable,
between power figures for its quadratic component. It seems that
bootstrapping the statistic itself is preferable.
\begin{landscape}
\begin{table}
\caption{\label{tab: table 2} Simulated rejection rates for asymptotic and wild bootstrap tests}\medskip
\begin{centering}
\begin{tabular}{ccccccccccccccccccc}
\toprule
element $\rightarrow $\quad & \multicolumn{9}{c}{linear} & \multicolumn{9}{c}{
quadratic} \\
\midrule
& \multicolumn{3}{c}{1\%} & \multicolumn{3}{c}{5\%} & \multicolumn{3}{c}{10\%
} & \multicolumn{3}{c}{1\%} & \multicolumn{3}{c}{5\%} & \multicolumn{3}{c}{
10\%} \\
\midrule
$c_{\pi }$ $\downarrow $& asy & \multicolumn{2}{c}{bootstrap} & asy & \multicolumn{2}{c}{bootstrap} & asy & \multicolumn{2}{c}{bootstrap} & asy & \multicolumn{2}{c}{bootstrap} &asy & \multicolumn{2}{c}{bootstrap} & asy & \multicolumn{2}{c}{bootstrap} \\
\midrule
& t & $\xi$ & t & t & $\xi$ & t& t & $\xi$ & t& t & $\xi$ & t & t & $\xi$ & t & t & $\xi$ & t\\ \midrule
& \multicolumn{18}{c}{Size: $c_{fv}=0$} \\
\midrule
0.5 & 2 & 1 & 1 & 8 & 5 & 5 & 14 & 10 & 10 & 0 & 1 & 1 & 3 & 6 & 5 & 6 & 10
& 10 \\
1 & 7 & 1 & 1 & 16 & 5 & 5 & 24 & 10 & 10 & 0 & 1 & 1 & 3 & 6 & 5 & 6 & 10 & 10
\\
2 & 25 & 1 & 1 & 38 & 5 & 5 & 46 & 10 & 10 & 0 & 1 & 1 & 3 & 6 & 6 & 7 & 11
& 12 \\
\midrule
& \multicolumn{18}{c}{Power: $c_{fv}=0.1$} \\
\midrule
0.5 & 9 & 6 & 6 & 22 & 17 & 17 & 32 & 27 & 26 & 0 & 2 & 1 & 3 & 7 & 6 & 7 &
12 & 12 \\
1 & 39 & 16 & 16 & 56 & 35 & 35 & 67 & 47 & 46 & 1 & 2 & 2 & 3 & 7 & 6 & 6 &
13 & 11 \\
2 & 79 & 29 & 29 & 87 & 51 & 51 & 90 & 63 & 63 & 1 & 6 & 3 & 5 & 14 & 9 & 11
& 20 & 17 \\
\midrule
& \multicolumn{18}{c}{Power: $c_{fv}=0.2$} \\
\midrule
0.5 & 35 & 28 & 26 & 58 & 51 & 49 & 68 & 63 & 62 & 1 & 5 & 3 & 5 & 13 & 9 & 10
& 19 & 16 \\
1 & 90 & 73 & 71 & 96 & 88 & 87 & 97 & 93 & 93 & 2 & 10 & 5 & 9 & 21 & 15 &
17 & 30 & 25 \\
2 & 100 & 93 & 92 & 100 & 98 & 98 & 100 & 99 & 99 & 18 & 45 & 29 & 40 & 61 &
50 & 53 & 68 & 61 \\
\bottomrule
\end{tabular}
\end{centering}
\medskip
{\footnotesize{\emph{Notes: } The table contains actual rates, computed from 5,000
simulations for $10\%,$ $5\%$ and $1\%$ declared size
asymptotic-t (`asy t'), bootstrap (`bootstrap $\xi$') and bootstrap-t (`bootstrap t') two-sided tests for deviations of
each component of $\Xi _{N,T}$ from the zero value.
The first component of $\Xi _{N,T}$ is labeled `linear' and second component is
labeled `quadratic'. The size figures are in
panel with $c_{fv}=0,$ and power figures are in panels with $c_{fv}\neq 0.$
Rows with $c_{\pi }=0.5,$ $1$ and $2$ correspond to a very weak, weak and
moderately strong error factor structure. }}
\end{table}
\end{landscape}
\section{CONCLUDING REMARKS}
Possible directions for future research may be relaxing the error time-series independence to martingale difference structures and inventing ways to consistently estimate the asymptotic variance matrix when it is not diagonal in the limit.
Other interesting areas involve establishing formal properties of the proposed wild bootstrap schemes, exploring the possibility of asymptotic refinements, and examining the superiority of one bootstrap scheme over the other.