Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
79,376 characters · 15 sections · 95 citation commands
Common Correlated Effects Estimation of Nonlinear Panel Data Models
\affil[1]{HSBC Business School, Peking Unviersity} \affil[2]{School of Economics, Shanghai University of Finance and Economics}
Nonlinear panel data models have posed a long-standing and challenging problem for identification and estimation. This difficulty is especially pronounced when the number of cross-sectional observations (denoted by $N$) is large, but the number of time-series observations (denoted by $T$) is small. The central challenge in this setting is to identify the coefficients of the observed regressors in the presence of fixed effects that have an unrestricted joint distribution with the regressors. Over the years, a considerable amount of work has been devoted to this problem, including classical results on binary choice models by manski1987semiparametric, honore2000panel, honore2002semiparametric, and chamberlain2010binary, as well as more recent developments by davezies2020fixed, honore2020moment, khan2020identification and zhu2022sufficient.
In long panels, where the number of time-series observations is comparable to the number of cross-sectional observations, the identification problem can be resolved by treating the fixed effects as parameters that are jointly estimated with the coefficients of the regressors, and the primary focus of this literature is to correct the asymptotic biases caused by the estimation errors of the fixed effects, also known as the incidental parameter biases --- see neyman1948consistent. This line of research was pioneered by hahn2004jackknife, and further developed by dhaene2015split.
Recent research on long panels has focused on models with interactive fixed effects that use a factor model structure to characterize the unobserved errors. These models nest the conventional two-way fixed effects models as special cases, providing a more general way to capture cross-sectional dependence in panels (see chudik2015large). For linear models with interactive fixed effects, fundamental contributions have been made by pesaran2006estimation, bai2009panel and moon2015linear. In recent years, chen2016estimation, boneva2017discrete, chen2020nonlinear, ando2022bayesian and gao2023binary have investigated the estimation of nonlinear models with interactive fixed effects.\footnote{wang2022maximum studies the estimation of nonliner factor models, which can be viewed as a special case of nonlinear panel data models with interactive fixed effects. However, in that paper, the main focus is the estimation of the factors and factor loadings. }
Panel data models with interactive fixed effects involve three sets of parameters: a finite-dimensional vector of coefficients (denoted by $\bm{\beta}$) for the observed regressors, a $T\times r$ matrix of latent factors (denoted by $\bm{F}$) representing the global shocks to all individuals, and an $N\times r$ matrix of factors loadings (denoted by $\bm{\Lambda}$) measuring the individual-specific responses to the factors.\footnote{Throughout the paper, $r$ is used to denote the number of factors.} While the main object of interest is $\bm{\beta}$, the latter two sets of parameters are introduced to account for individual heterogeneity and cross-sectional dependence. To estimate these parameters in nonlinear models, the papers mentioned above usually take two different approaches. The first approach, utilized by chen2016estimation, chen2020nonlinear, ando2022bayesian and gao2023binary, involves estimating $(\bm{\beta},\bm{F},\bm{\Lambda})$ jointly using some iterative algorithms. The second approach, introduced by boneva2017discrete, extends the CCE estimation method of pesaran2006estimation to nonlinear models.
This paper fills a gap in the literature by studying the estimation of nonlinear panel data models with interactive fixed effects using the CCE framework, where the observed regressors are assumed to be driven by the same latent factors $\bm{F}$ and the coefficients of the regressors are homogeneous across individuals (see Table 1 below). As the main contribution of this paper, a two-step estimator for this type of models is proposed. In the first step, an estimator for the latent factors is constructed using the observed regressors; in the second step, given the estimated factors, $\bm{\beta}$ and $\bm{\Lambda}$ are estimated jointly to maximize the objective function (e.g., the log likelihood function). Notably, the proposed method for estimating the factors in the first step is different from the standard CCE approach, which can suffer from the problem of degenerated regressors (see Karabiyik201760). Asymptotic properties, particularly asymptotic biases, of the proposed estimators are derived under the framework of long panels and other general conditions, with a Bahadur representation established for the estimator of $\bm{\beta}$ where the leading biases are shown to be of order $1/T+1/N$. To the best of our knowledge, this is the first result of this kind for CCE estimators of nonlinear panels. In addition, both analytical and split-panel jackknife methods are introduced to correct the asymptotic biases, providing the basis for valid inference in large samples. Through Monte Carlo simulations, we find that the proposed bias-correction procedures significantly reduce the biases of the estimators and improve the empirical coverage rates of the confidence intervals in finite samples.
Compared with the approach that estimate $(\bm{\beta},\bm{F},\bm{\Lambda})$ jointly, the main advantage of the CCE estimation method is its much lower computational cost, because the estimation problem is decomposed into a simpler first step and a second step that can be easily implemented in standard software such as Stata and R. Additionally, for the leading examples of binary choice models, the log likelihood function is convex in $(\bm{\beta},\bm{\Lambda})$ given $\bm{F}$, simplifying the search for the global maximum of the objective function. In contrast, joint estimation of these parameters can be computationally challenging even for linear models, as the objective functions are generally not convex in $(\bm{\beta},\bm{F},\bm{\Lambda})$ (see bai2009panel). However, the success of the CCE approach relies on the assumption that it is possible to consistently estimate the (space of) latent factors using the observed regressors, which contradicts the usual assumption of cross-sectional independence. Nonetheless, as pointed out in andrews2005cross: “...it seems apparent that common shocks (macroeconomic, technological, legal/institutional, political, environmental, health, and sociological shocks) are a likely feature of cross-section economic data. This is true whether the population units in the cross-section regression are individuals, households, firms, industries, plants, cities, states, countries, or products.” Thus, the seemly stronger CCE assumption that the cross-sectional dependence of observations are driven by the same latent factors could be potentially advantageous in certain applications.
Finally, in a closely related paper, boneva2017discrete also considered the CCE estimation of nonlinear panels, but their approach assumes that the coefficients of the regressors are heterogeneous across individuals. Consequently, the estimators of these coefficients converge at the rate of $\sqrt{T}$, and their asymptotic distributions are free of asymptotic biases. In this paper, the CCE estimator of the homogeneous coefficients converges at the rate of $\sqrt{NT}$, making it far more challenging to establish its asymptotic distribution, because many higher order terms in the stochastic expansion of the estimator now become asymptotic biases whose analytical forms need to be carefully derived.
The rest of the paper is structured as follows: Section 2 introduces the model along with the CCE estimators for the coefficients and the average partial effect of the regressors. Moving on to Section 3, we establish the asymptotic properties of the estimators and demonstrate how to correct their asymptotic biases and consistently estimate their asymptotic variances. Simulation results are presented in Section 4 to evaluate the performance of the estimators in finite samples. Section 5 describes an empirical application where we use the proposed method to study the arbitrage behavior of U.S. nonfinancial firms across different security markets. Finally, Section 6 concludes. The appendix provides the proof of Theorem 1, whereas the proofs of the other theorems are available in an online appendix to save space.
Let $y_{it}\in\mathbb{R}$ and $\boldsymbol{x}_{it}\in\mathbb{R}^k$ be the observed outcome and covariates respectively for individual $i$ at period $t$, and let $\boldsymbol{\lambda}_i, \boldsymbol{f}_t \in \mathbb{R}^r$ be the unobserved factor loadings (or individual effects) and factors (or time effects) respectively. Suppose that we have a random sample $\{ y_{it}, \boldsymbol{x}_{it}\}$ for $i=1,\ldots,N, t=1,\ldots,T$, where the realized value of $\boldsymbol{\lambda}_i, \boldsymbol{f}_t$ are denoted as $\boldsymbol{\lambda}_{0i}, \boldsymbol{f}_{0t}$. Moreover, for some $\boldsymbol{\beta}_0 \in \mathbb{R}^k$, the likelihood function of $y_{it}$ given $(\boldsymbol{x}_{it},\boldsymbol{\lambda}_{0i}, \boldsymbol{f}_{0t} )$ can be written as \[ L(y_{it}, \boldsymbol{\beta}_0'\boldsymbol{x}_{it} + \boldsymbol{\lambda}_{0i}'\boldsymbol{f}_{0t}).\] Through out the paper, the following examples are used to illustrate the applicability of our general theoretical results.
For any $\boldsymbol{\beta}\in\mathbb{R}^k$ and $\boldsymbol{\lambda}_i,\boldsymbol{f}_t\in \mathbb{R}^r$, write $ l_{it}( \boldsymbol{\beta}, \boldsymbol{\lambda}_i'\boldsymbol{f}_t) = \log \left[ L(y_{it}, \boldsymbol{\beta}'\boldsymbol{x}_{it} + \boldsymbol{\lambda}_i'\boldsymbol{f}_t) \right]$, then the objective function can be written as \[ \mathcal{L}_{NT}( \boldsymbol{\beta}, \boldsymbol{\Lambda}, \boldsymbol{F} ) = \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} l_{it}( \boldsymbol{\beta}, \boldsymbol{\lambda}_i'\boldsymbol{f}_t) , \] where $\boldsymbol{\Lambda}=(\bm{\lambda}_1,\ldots,\bm{\lambda}_N)'$ and $\boldsymbol{F}=(\bm{f}_1,\ldots,\boldsymbol{f}_t)'$. In the spirit of bai2009panel, chen2020nonlinear proposes to estimate $(\boldsymbol{\beta}, \boldsymbol{\Lambda}, \boldsymbol{F})$ jointly to maximize the above objective function. This estimation procedure leaves the relationship between the regressors, the factors and factor loadings unspecified, thus it is usually termed as the fixed-effects estimator. The practical implementation of this estimation method usually involves iterations between $\bm{\beta}$ and $(\boldsymbol{\Lambda}, \boldsymbol{F})$. In particular, chen2016estimation proposed an EM-type iterative algorithm that converges to local maximums of the objective function. Moreover, it is usually assumed that the number of factor $r$ is known in the asymptotic theory and in the practical implementation of the fixed-effects estimator.\footnote{chen2020nonlinear proposed an method adapted from the eigen-ratio estimator of ahn2013eigenvalue to estimate $r$, but the consistency of their method was not established.} The effect of overestimating $r$ in linear panel data models is analyzed by moon2015linear, but extending their analysis to nonlinear models is much more challenging.
In this paper, we try to overcome the problems of the fixed-effects estimator mentioned above by taking a CCE approach pioneered by pesaran2006estimation. The CCE approach starts by assuming a linear relationship between the regressors and the common factors as follows:
for $i=1,\ldots,N, t=1\ldots,T$, where $\boldsymbol{\Gamma}_i$ is a $k\times r$ matrix of non-random constants, and $\boldsymbol{e}_{it} \in \mathbb{R}^k$ is a vector of idiosyncratic components. The key of the CCE estimator is to approximate the common factors by the cross-sectional averages of the regressors: $\bar{ \boldsymbol{x} }_t = N^{-1}\sum_{i=1}^{N} \boldsymbol{x}_{it}.$\footnote{In linear panel data models, (ref) implies that the dependent variables have a factor structure, thus the cross-sectional averages of the dependent variables are also used to approximate the common factors. However, in nonlinear models, the dependent variables generally do not have a factor structure. Thus, only the cross-sectional averages of the regressors are used.} To ensure that the space of the common factors can be spanned by $\bar{ \boldsymbol{x} }_t$, the following assumption is usually imposed:
Under Assumption 1 and some other restrictions on the cross-sectional dependence of $\boldsymbol{e}_{it}$, it is easy to show that \[ \bar{\boldsymbol{x}}_t = \boldsymbol{\Gamma}_0 \boldsymbol{f}_{0t} + o_P(1) \text{ for all }t=1,\ldots,T.\] When $k=r$, the above equation implies that $\bar{\boldsymbol{x}}_t$ can be used as approximations of $\boldsymbol{f}_{0t}$ since they span the same space asymptotically. However, when $k>r$, using $\bar{\boldsymbol{x}}_t$ as estimators of $\boldsymbol{f}_{0t}$ amounts to overestimating the number of factors. In particular, the fact that $T^{-1}\sum_{t=1}^T \bar{\boldsymbol{x}}_t\bar{\boldsymbol{x}}_t'$ converges in probability to a singular matrix when $k>r$ implies asymptotic multicolinearity of the estimated factors, leading to the problem of degenerated regressors as pointed out in Karabiyik201760. For linear models, Karabiyik201760 showed that the standard CCE estimator for $\bm{\beta}_0$ is still consistent but it suffers from extra asymptotic biases due to this problem.
To overcome the problem of degenerated regressors in the standard CCE method, in this paper we use an alternative approach to estimate the common factors. Let $\hat{\boldsymbol{\Sigma}}_{\bar{\boldsymbol{x}}}=T^{-1}\sum_{t=1}^T \bar{\boldsymbol{x}}_t\bar{\boldsymbol{x}}_t'$ and $\hat{\boldsymbol{\Psi}}$ be the $k\times r$ matrix of eigenvectors associated with the first $r$ eigenvalues of $\hat{\boldsymbol{\Sigma}}_{\bar{\boldsymbol{x}}}$, then the estimated factors are defined as $\hat{\boldsymbol{f}}_t = \hat{\boldsymbol{\Psi}}' \bar{\boldsymbol{x}}_t$. To establish the properties of $\hat{\boldsymbol{f}}_t$, we need to impose the following assumptions:
Moreover, define $\hat{\bm{H}} =\hat{\boldsymbol{\Psi}}' \bar{\boldsymbol{\Gamma}} $. Let $\bm{D}$ be a $r \times r$ diagonal matrix with the non-zeros eigenvalues of $ \boldsymbol{\Gamma}_0 \boldsymbol{\Sigma}_{f_0} \boldsymbol{\Gamma}_0'$ in decreasing order, and let $\boldsymbol{\Psi}_{0}$ be the matrix of corresponding eigenvectors such that $\boldsymbol{\Gamma}_0 \boldsymbol{\Sigma}_{f_0} \boldsymbol{\Gamma}_0' \boldsymbol{\Psi}_0 = \boldsymbol{\Psi}_0 \bm{D}$. Then it can be shown that:
The above results implies that $\hat{\bm{f}}_t$ is a consistent estimator of $\bm{f}_t$ up to a non-singular normalization matrix. Given $\hat{\boldsymbol{f}}_t$, the CCE estimator is defined as
where $ l_{it}( \boldsymbol{\beta} , \boldsymbol{\lambda}_i' \hat{\boldsymbol{f}}_{t}) = \log L(y_{it}, \boldsymbol{\beta}'\boldsymbol{x}_{it} + \boldsymbol{\lambda}_i' \hat{\boldsymbol{f}}_{t})$, and $\mathcal{B}\subset\mathbb{R}^k, \mathcal{A}\subset\mathbb{R}^r$. Note that in practice, the CCE estimator can be obtained using standard packages in Matlab or R by treating $y_{it }$ as the dependent variable and $( \boldsymbol{x}_{it}, \boldsymbol{1}\{ i=1\} \hat{\boldsymbol{f}}_t, \ldots, \boldsymbol{1}\{ i=N\} \hat{\boldsymbol{f}}_t )$ as the regressors. It should also be mentioned that the computational cost of the CCE estimator is much lower than the fixed-effects estimator, because in the maximization problem (ref) there are $k+Nr$ parameters while for the fixed-effects estimator there are $k+(N+T)r$ parameters. More importantly, since the objective function $\mathcal{L}_{NT}( \boldsymbol{\beta}, \boldsymbol{\Lambda}, \boldsymbol{F} )$ is generally not convex in $( \boldsymbol{\beta}, \boldsymbol{\Lambda}, \boldsymbol{F} ) $, the joint estimation of these estimators normally involves iterative procedures that do not necessary find the global maximum of the objective function. However, it is well known that given $\bm{F}$, the objective function becomes convex in $(\boldsymbol{\beta}, \boldsymbol{\Lambda}) $ in our leading examples (e.g., probit and logit models). Thus, the estimator defined in (ref) can be easily obtained using standard optimization methods such as the gradient descent algorithm, without the need of a good initial estimator.
Finally, the analysis above assumes that $r$ is known. In practice, $r$ needs to be estimated before implementing the CCE method. Observe that if $\hat{\boldsymbol{\Sigma}}_{f_0}=T^{-1}\sum_{t=1}^{T}\boldsymbol{f}_{0t}\boldsymbol{f}_{0t}'$ converges to a positive definite matrix, it is easy to show that $\hat{\boldsymbol{\Sigma}}_{\bar{\boldsymbol{x}}}$ converges in probability to a matrix with rank $r$. Thus, to estimate $r$, we can just estimate the rank of $\hat{\boldsymbol{\Sigma}}_{\bar{\boldsymbol{x}}}$. In particular, let $\hat{\rho}_1,\ldots, \hat{\rho}_k$ be the eigenvalues of $\hat{\boldsymbol{\Sigma}}_{\bar{\boldsymbol{x}}}$ in decreasing order, and let $P_{NT}$ be a sequence of constants converging to 0 as $N,T\rightarrow \infty$, the estimator of $r$ can be simply defined as: \[ \hat{r} = \sum_{j=1}^{k} \boldsymbol{1} \{ \hat{\rho}_j \geq P_{NT}\}.\] The following result is established in chen2022twostep.\footnote{The proofs of Proposition 1 and Proposition 2 are identical to the proofs of Proposition 1 and Lemma 1 in chen2022twostep. Thus, they are omitted to save space.}
Given the above result, the true number of factors $r$ can be treated as known in the rest of the paper.\footnote{See footnote 5 of bai2003inferential.}
For models with limited dependent variables, the coefficient $\boldsymbol{\beta}_0$ usually cannot capture the partial effects of the regressors, which are the main object of interests for most practitioners. Consider the binary choice models, and let $\boldsymbol{x}^{(0)},\boldsymbol{x}^{(1)}\in\mathbb{R}^k$ be the values of the regressors before and after some policy intervention, the effect of the policy on the probability of success is given by \[ \delta( \boldsymbol{x}^{(0)},\boldsymbol{x}^{(1)}; \boldsymbol{\beta}_0, \boldsymbol{\lambda}_{0i}' \boldsymbol{f}_{0t}) = G(\boldsymbol{\beta}_0' \boldsymbol{x}^{(1)}+\boldsymbol{\lambda}_{0i}'\boldsymbol{f}_{0t} ) -G(\boldsymbol{\beta}_0' \boldsymbol{x}^{(0)}+\boldsymbol{\lambda}_{0i}'\boldsymbol{f}_{0t} ) . \] A typical example is that the first element of $x$ is a binary indicator for some treatment or policy change, then $ \boldsymbol{x}^{(0)} =(0, x_2,\ldots, x_k)$ and $ \boldsymbol{x}^{(1)} =(1, x_2,\ldots, x_k)$. In this case, $ \delta( \boldsymbol{x}^{(0)},\boldsymbol{x}^{(1)}; \boldsymbol{\beta}_0, \boldsymbol{\lambda}_{0i}' \boldsymbol{f}_{0t})$ denotes the partial effect of the treatment for individual $i$ at time $t$, while his/her other characteristics are fixed at $(x_2,\ldots, x_k)$. Similar partial effects can be defined for other nonlinear models, depending on the applications at hand. However, it should be noted that in nonlinear models such partial effects generally depends on $\boldsymbol{x}^{(0)}, \boldsymbol{x}^{(1)}, \boldsymbol{\lambda}_{0i},\boldsymbol{f}_{0t}$.
To summarize the partial effects for all the individuals in the dataset, we consider the following average partial effect (APE): \[ \bar{\delta}_{0} ( \boldsymbol{x}^{(0)},\boldsymbol{x}^{(1)} ) = \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}\delta( \boldsymbol{x}^{(0)},\boldsymbol{x}^{(1)}; \boldsymbol{\beta}_0, \boldsymbol{\lambda}_{0i}'\boldsymbol{f}_{0t}), \] which is different from the definitions of chen2020nonlinear and boneva2017discrete, who take expectations of the partial effects with respect to the distribution of the regressors. Note that our definition of APE is closer in spirit to the definition of hahn2004jackknife, in the sense that it only averages out the individual and time effects while fixing the values of the regressors. Given $(\hat{\boldsymbol{\beta}}, \hat{\boldsymbol{\Lambda}}, \hat{\boldsymbol{F}} )$, the estimator of APE is simply given by \[ \hat{\delta} ( \boldsymbol{x}^{(0)},\boldsymbol{x}^{(1)} ) = \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}\delta( \boldsymbol{x}^{(0)},\boldsymbol{x}^{(1)}; \hat{\boldsymbol{\beta}}, \hat{\boldsymbol{\lambda}}_{i}' \hat{\boldsymbol{f}}_{t}). \]
This section presents the main theoretical results of this paper. Sections 3.1 and 3.2 provide the asymptotic distributions of the CCE estimators for $\bm{\beta}_0$ and the APE. Section 3.3 discusses how to correct the asymptotic biases of the estimators. Finally, estimators for the asymptotic variances are proposed in Section 3.4. It should be noted that, as in chen2020nonlinear and chen2022twostep, the assumptions and the asymptotic results below are all conditional on $(\bm{\Lambda},\bm{F})= (\bm{\Lambda}_0,\bm{F}_0)$.
Let $c_{0,it}= \boldsymbol{\lambda}_{0i}'\boldsymbol{f}_{0t}$. Write $\tilde{\boldsymbol{f}}_{0t} = \boldsymbol{H}_0 \boldsymbol{f}_{0t} $ and $\tilde{\boldsymbol{\lambda}}_{0i} = (\boldsymbol{H}_0^{-1})' \boldsymbol{\lambda}_{0i}$. Note that $ \tilde{\boldsymbol{\lambda}}_{0i}' \tilde{\boldsymbol{f}}_{0t} = \boldsymbol{\lambda}_{0i}'\boldsymbol{f}_{0t}=c_{0,it}$. Let $\mathcal{F}$ be a compact subset of $\mathbb{R}^r$ such that $\tilde{\boldsymbol{f}}_{0t} \in \mathcal{F}$ for all $t$. Let $\mathcal{C}$ be a compact subset of $\mathbb{R}$ such that $\boldsymbol{\lambda}'\boldsymbol{f} \in \mathcal{C}$ for all $\boldsymbol{\lambda} \in \mathcal {A}$ and all $\boldsymbol{f} \in \mathcal{F}$. Moreover, define \[ l_{it}^{(j)}(\boldsymbol{\beta}, \boldsymbol{\lambda}_i'\boldsymbol{f}_t) = \frac{ \partial^j \log [ L(y,z)]}{\partial z^j }|_{y=y_{it},z=\boldsymbol{\beta}'\boldsymbol{x}_{it}+\boldsymbol{\lambda}_i'\boldsymbol{f}_t } \text{ for }j=1,\ldots,4, \] and \[ \underbrace{\boldsymbol{A}_i}_{r\times r}= \frac{1}{T}\sum_{t=1}^{T} \mathbb{E}[ l_{it}^{(2)}]\boldsymbol{f}_{0t} \boldsymbol{f}_{0t}' , \quad \underbrace{\boldsymbol{B}_i}_{k\times r} = \frac{1}{T}\sum_{t=1}^{T} \mathbb{E}[ l_{it}^{(2)}\boldsymbol{x}_{it} ] \boldsymbol{f}_{0t}', \quad \dot{\boldsymbol{x}}_{it} =\boldsymbol{x}_{it} - \boldsymbol{B}_i\boldsymbol{A}^{-1}_i \boldsymbol{f}_{0t} ,\] where we suppress the arguments of $l_{it}^{(j)}$ when they are evaluated at $(\boldsymbol{\beta}_0,c_{0,it})$ to simplify the notations. Assume the following conditions hold:
Most of the conditions in Assumption 4 are standard in the literature, with a few exceptions. First, Assumption 4(i) excludes cross-sectional dependence in the data,\footnote{As stressed in the beginning of this section, these assumptions are made conditional on the factors and factor loadings. Thus, cross-sectional dependence due to the common factors are still allowed.} but it allows for general time-series dependence for each $i$ that is excluded by Assumption 1(i) of chen2020nonlinear. Second, unlike Assumption 1(v) of chen2020nonlinear, Assumption B of ando2022bayesian and Assumption 2.2.C of gao2023binary, we don't need $N^{-1}\boldsymbol{\Lambda}_0' \boldsymbol{\Lambda}_0$ to converge to some positive definite matrix. Thus, our assumption allow some columns of $\boldsymbol{\Lambda}_0$ to be 0, meaning that some factors may affect the dependent variables $y_{it}$ only indirectly through the regressors $\bm{x}_{it}$. Third, our moment restrictions on $\bm{x}_{it}$ are generally much weaker than those imposed in existing studies. For example, both chen2020nonlinear and ando2022bayesian require $\bm{x}_{it}$ to have bounded support, while gao2023binary assumes that $\max_{i\leq N,t\leq T}\|\bm{x}_{it}\| = O_p(\log NT)$, which essentially require all the moments of $\bm{x}_{it}$ to exist.
In order to establish the asymptotic distribution of the CCE estimator, the following definition are needed: \[\underbrace{\boldsymbol{C}_t }_{k\times r} = \frac{1}{N}\sum_{i=1}^{N} \mathbb{E}[ l_{it}^{(2)} \dot{\boldsymbol{x}}_{it} ] \boldsymbol{\lambda}_{0i}', \quad \underbrace{\boldsymbol{D}_{t,j}}_{r\times r}= \frac{1}{N} \sum_{i=1}^{N}\boldsymbol{\lambda}_{0i}\boldsymbol{B}_{i,j} \boldsymbol{A}^{-1}_i \bar{l}_{it}^{(2)}, \] \[ \underbrace{\boldsymbol{G}_{t,j}}_{r\times r} =\frac{1}{N}\sum_{i=1}^{N} \mathbb{E}\left[ l_{it}^{(3)} \dot{\boldsymbol{x}}_{it,j} \right]\boldsymbol{\lambda}_{0i} \boldsymbol{\lambda}_{0i}', \quad \underbrace{\boldsymbol{Q}_{i}}_{r\times r}= \frac{1}{T}\sum_{t=1}^{T}\sum_{s=1}^{T} \mathbb{E}\left[ l_{it}^{(1)}l_{is}^{(1)}\right] \boldsymbol{f}_{0t}\boldsymbol{f}_{0s}' , \] where $\boldsymbol{B}_{i,j}$ is the $j$th row of $\boldsymbol{B}_{i}$.
Then we can show that:
The proof of Theorem 1 is based on the following Bahadur representation for $\hat{\bm{\beta}}$: \[ \boldsymbol{\Delta} (\hat{\boldsymbol{\beta}} -\boldsymbol{\beta}_0)+o_P(\| \hat{\boldsymbol{\beta}} -\boldsymbol{\beta}_0\|) =- \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} \bm{w}_{it} \\ +\frac{\bm{b}}{T}+ \frac{\bm{d}}{N}+ o_P(T^{-1}), \] where the bias term $\bm{b}/T$ is due to the estimation error of $\bm{\hat{\Lambda}}$, and the bias term $\bm{d}/N$ is caused by the estimation error of $\bm{\hat{F}}$. Similar Bahadur representations were established by fernandez2016individual and chen2020nonlinear for nonlinear panel data models, and by chen2022twostep for quantile panel data models. This representation provides the theoretical basis for the analytical and split-panel jackknife bias corrections that will be discussed in Section 3.3 below.
In the next two remarks, we compare Theorem 1 above with the main results of hahn2004jackknife and chen2020nonlinear. Following these two papers, we assume that there is no time series dependence to facilitate the comparison. Note that in this case, the expressions of $\bm{Q}_i$, $\bm{\Omega}$ and $\bm{b}^2$ reduce to \[\boldsymbol{Q}_{i}= \frac{1}{T}\sum_{t=1}^{T} \mathbb{E}\left[ l_{it}^{(1)}\right]^2 \boldsymbol{f}_{0t}\boldsymbol{f}_{0t}' ,\quad \boldsymbol{\Omega} = \lim_{N,T\rightarrow \infty}\frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} \mathbb{E} \left[ \boldsymbol{w}_{it}\boldsymbol{w}_{it}'\right],\] \[\bm{b}^2=\lim_{N,T\rightarrow\infty} \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}\left[ l_{it}^{(2)}l_{it}^{(1)} \dot{\boldsymbol{x}}_{it}\right] \cdot \boldsymbol{f}_{0t} ' \boldsymbol{A}^{-1}_i \boldsymbol{f}_{0t}. \]
To derive the asymptotic distribution of the APE estimator, we first define the partial derivatives of $\delta$: $ \boldsymbol{\delta}^{\boldsymbol{\beta}}(\boldsymbol{\beta},c ) = \partial \delta (\boldsymbol{x}^{(0)},\boldsymbol{x}^{(1)}; \boldsymbol{\beta},c ) / \partial \boldsymbol{\beta}, \delta^{c}( \boldsymbol{\beta},c ) = \partial \delta ( \boldsymbol{x}^{(0)},\boldsymbol{x}^{(1)};\boldsymbol{\beta},c ) / \partial c$. Moreover, $\boldsymbol{\delta}^{\boldsymbol{\beta} c}( \boldsymbol{\beta},c )$, $\boldsymbol{\delta}^{\boldsymbol{\beta} \boldsymbol{\beta}}( \boldsymbol{\beta},c )$, $\delta^{c c}(\boldsymbol{\beta},c )$ and $\delta^{c cc}( \boldsymbol{\beta},c ) $ can be defined in a similar fasion. For simplicity, write $\boldsymbol{\delta}^{\boldsymbol{\beta}}_{0,it} = \boldsymbol{\delta}^{\boldsymbol{\beta}}(\boldsymbol{\beta}_0, c_{0,it} ) $, $\delta^{c}_{0,it} = \delta^{c}( \boldsymbol{\beta}_0, c_{0,it} ) $ and $\delta^{cc}_{0,it} = \delta^{cc}( \boldsymbol{\beta}_0, c_{0,it} ). $ In addition, define \[\boldsymbol{\gamma} = \lim_{N,T\rightarrow\infty} \left( \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} \boldsymbol{\delta}^{\boldsymbol{\beta}}_{0,it} -\frac{1}{N}\sum_{i=1}^{N} \boldsymbol{B}_i \boldsymbol{A}_i^{-1}\boldsymbol{\gamma}_i \right),\quad \boldsymbol{\gamma}_i = \frac{1}{T}\sum_{t=1}^{T} \delta^{c}_{0,it} \boldsymbol{f}_{0t} , \quad \boldsymbol{\gamma}_t= \frac{1}{N} \sum_{i=1}^{N} \delta^{c}_{0,it}\boldsymbol{\lambda}_{0i},\] \[ \boldsymbol{R}_t =\frac{1}{N}\sum_{i=1}^{N} \boldsymbol{\lambda}_{0i} \boldsymbol{\gamma}_i' \boldsymbol{A}_i^{-1} \bar{l}_{it}^{(2)}, \quad \boldsymbol{W}_t = \frac{1}{N}\sum_{i=1}^{N}\left( \delta^{cc}_{0,it} -0.5 \bar{l}_{it}^{(3)} \cdot \boldsymbol{\gamma}_i' \boldsymbol{A}^{-1}_i \boldsymbol{f}_{0t} \right) \boldsymbol{\lambda}_{0i}\boldsymbol{\lambda}_{0i}' , \] and assume that:
Then, it can be shown that:
In order to make valid inference, we need to eliminate the asymptotic biases of the CCE estimator and the APE estimator. This can be done by either analytical bias correction or by split-panel jackknife method.
For analytical bias correction, we need to construct consistent estimators of $\boldsymbol{\Delta}, \bm{b}^1,\bm{b}^2,\bm{d}^1,\bm{d}^2$ and $\boldsymbol{\gamma},b^3,b^4,d^3,d^4$. First, consider the bias correction of $\bm{\hat{\beta}}$ and define: \[\hat{l}_{it}^{(j)}=l_{it}^{(j)}(\hat{\boldsymbol{\beta}},\hat{c}_{it}) , \quad \hat{\boldsymbol{A}}_i= \frac{1}{T}\sum_{t=1}^{T} \hat{l}_{it}^{(2)} \hat{\boldsymbol{f}}_{t} \hat{\boldsymbol{f}}_{t}' , \quad \hat{\boldsymbol{B}}_i = \frac{1}{T}\sum_{t=1}^{T} \hat{l}_{it}^{(2)}\boldsymbol{x}_{it} \hat{\boldsymbol{f}}_{t}', \quad \hat{\dot{\boldsymbol{x}}}_{it} =\boldsymbol{x}_{it} - \hat{\boldsymbol{B}}_i \hat{\boldsymbol{A}}^{-1}_i \hat{\boldsymbol{f}}_{t} ,\] \[ \hat{\boldsymbol{\Delta}} = \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} \hat{l}_{it}^{(2)} \hat{\dot{\boldsymbol{x}}}_{it} \hat{\dot{\boldsymbol{x}}}_{it}', \quad \hat{\boldsymbol{D}}_{t,j} = \frac{1}{N} \sum_{i=1}^{N} \hat{\boldsymbol{\lambda}}_{i} \hat{\boldsymbol{B}}_{i,j} \hat{\boldsymbol{A}}^{-1}_i \hat{l}_{it}^{(2)}, \quad \hat{\boldsymbol{G}}_{t,j} =\frac{1}{N}\sum_{i=1}^{N} \hat{l}_{it}^{(3)} \hat{\dot{\boldsymbol{x}}}_{it,j} \hat{\boldsymbol{\lambda}}_{i} \hat{\boldsymbol{\lambda}}_{i}',\quad \hat{\boldsymbol{\Upsilon}}= \hat{\boldsymbol{\Psi}}' , \] \[\hat{\boldsymbol{Q}}_i=\frac{1}{T}\sum_{t=1}^{T}\sum_{s=1}^{T}\hat{l}_{it}^{(1)}\hat{l}_{is}^{(1)}\hat{\boldsymbol{f}}_{t}\hat{\boldsymbol{f}}_{s}'k\left(\frac{t-s}{L}\right), \quad \hat{\boldsymbol{e}}_{it} = \boldsymbol{x}_{it} - \hat{\boldsymbol{\Gamma}}_i \hat{\boldsymbol{f}}_t, \quad \hat{\boldsymbol{\Gamma}}_i' = \left( \sum_{t=1}^{T} \hat{\boldsymbol{f}}_t \hat{\boldsymbol{f}}_t '\right)^{-1}\left( \sum_{t=1}^{T} \hat{\boldsymbol{f}}_t \boldsymbol{x}_{it} '\right) ,\] where $L\rightarrow\infty$ as $N,T\rightarrow\infty$ and $k(x)=(1-|x|)\boldsymbol{1}\{|x|\leq1\}$ is the Bartlett kernel function, corresponding to the HAC estimators of newey1987. Then the estimators of the biases of $\bm{\hat{\beta}}$ can be constructed as follows: \[ \hat{\bm{b}}^{1}= -0.5\frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} \hat{l}_{it}^{(3)} \hat{\dot{\boldsymbol{x}}}_{it} \hat{\boldsymbol{f}}_{t}' \hat{\boldsymbol{A}}^{-1}_i \hat{\boldsymbol{Q}}_i \hat{\boldsymbol{A}}^{-1}_i \hat{\boldsymbol{f}}_{t} ,\quad \hat{\bm{b}}^{2} = \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}\sum_{s=1}^{T} \hat{l}_{it}^{(2)} \hat{l}_{is}^{(1)} \hat{\dot{\boldsymbol{x}}}_{it} \hat{\boldsymbol{f}}_{t}' \hat{\boldsymbol{A}}^{-1}_i \hat{\boldsymbol{f}}_{s}k\left(\frac{t-s}{L}\right) ,\] \[ \hat{\bm{d}}^{1}= - \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} \hat{l}_{it}^{(2)} \hat{\dot{\boldsymbol{x}}}_{it} \hat{\boldsymbol{e}}_{it} ' \hat{\boldsymbol{\Upsilon}}' \hat{\boldsymbol{\lambda}}_{i}, \quad \hat{d}^{2}_j = \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} \operatorname*{Tr} \left( \hat{\boldsymbol{e}}_{it} \hat{\boldsymbol{e}}_{it}' \cdot \hat{\boldsymbol{\Upsilon}}' ( \hat{\boldsymbol{D}} _{t,j} -0.5 \hat{\boldsymbol{G}}_{t,j }) \hat{\boldsymbol{\Upsilon}} \right), \] and $\hat{\bm{d}}^2=(d_1^2,\dots,d_k^2)'$, $\hat{\bm{b}}=\hat{\bm{b}}^1+\hat{\bm{b}}^2$, $\hat{\bm{d}}=\hat{\bm{d}}^1+\hat{\bm{d}}^2$. The bias-corrected CCE estimator is then defined as \[ \hat{ \boldsymbol{\beta}}_{ABC} = \hat{\boldsymbol{\beta}} - \hat{\boldsymbol{\Delta}}^{-1}\left(\frac{\hat{\bm{b}}}{T} + \frac{\hat{\bm{d}}}{N}\right) .\]
Next, consider the bias correction of the APE estimator and define $\hat{\delta}^{c}_{it} = \delta^c(\hat{\boldsymbol{\beta}},\hat{c}_{it})$, $\hat{\boldsymbol{\delta}}^{\boldsymbol{\beta}}_{it} = \boldsymbol{\delta}^{\boldsymbol{\beta}}(\hat{\boldsymbol{\beta}},\hat{c}_{it})$, $\hat{\delta}^{cc}_{it} = \delta^{cc}(\hat{\boldsymbol{\beta}},\hat{c}_{it})$, \[ \hat{\boldsymbol{\gamma}}_i = \frac{1}{T}\sum_{t=1}^{T} \hat{\delta}^{c}_{it} \hat{\boldsymbol{f}}_{t} , \quad \hat{\boldsymbol{\gamma}}_t= \frac{1}{N} \sum_{i=1}^{N} \hat{\delta}^{c}_{it}\hat{\boldsymbol{\lambda}}_{i}, \quad \hat{\boldsymbol{\gamma}} = \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} \hat{\boldsymbol{\delta}}^{\boldsymbol{\beta}}_{it} -\frac{1}{N}\sum_{i=1}^{N} \hat{\boldsymbol{B}}_i \hat{\boldsymbol{A}}_i^{-1} \hat{\boldsymbol{\gamma}}_i, \] \[ \hat{\boldsymbol{R}}_t =\frac{1}{N}\sum_{i=1}^{N} \hat{\boldsymbol{\lambda}}_{i} \hat{\boldsymbol{\gamma}}_i' \hat{\boldsymbol{A}}_i^{-1} \hat{l}_{it}^{(2)}, \quad \hat{\boldsymbol{W}}_t = \frac{1}{N}\sum_{i=1}^{N}\left( \hat{\delta}^{cc}_{it} -0.5 \hat{l}_{it}^{(3)} \cdot \hat{\boldsymbol{\gamma}}_i' \hat{\boldsymbol{A}}^{-1}_i \hat{\boldsymbol{f}}_{t} \right) \hat{\boldsymbol{\lambda}}_{i} \hat{\boldsymbol{\lambda}}_{i}' , \] \[\hat{b}^3 = \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} \left( \hat{\delta}^{cc}_{it} -0.5 \hat{l}_{it}^{(3)} \cdot \hat{\boldsymbol{\gamma}}_i' \hat{\boldsymbol{A}}^{-1}_i \hat{\boldsymbol{f}}_{t} \right) \cdot \hat{\boldsymbol{f}}_{t}' \hat{\boldsymbol{A}}^{-1}_i \hat{\boldsymbol{Q}}_{i} \hat{\boldsymbol{A}}^{-1}_i \hat{\boldsymbol{f}}_{t},\] \[\hat{b}^4 = \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}\sum_{s=1}^{T} \hat{ l}_{it}^{(2)} \hat{l}_{is}^{(1)} \cdot \hat{\boldsymbol{\gamma}}_i' \hat{\boldsymbol{A}}^{-1}_i \hat{\boldsymbol{f}}_{t} \cdot \hat{ \boldsymbol{f}}_{t} ' \hat{\boldsymbol{A}}^{-1}_i \hat{\boldsymbol{f}}_{s}k\left(\frac{t-s}{L}\right),\] \[\hat{d}^3 = - \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} \hat{\boldsymbol{\gamma}}_{i}' \hat{\boldsymbol{A}}^{-1}_i \hat{\boldsymbol{f}}_{t} \cdot \hat{\boldsymbol{\lambda}}_{i} ' \hat{\boldsymbol{\Upsilon}} \cdot \hat{l}_{it}^{(2)} \hat{\boldsymbol{e}}_{it} , \quad \hat{d}^4 = \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} \operatorname*{Tr}\left[ \hat{\boldsymbol{\Upsilon}}' ( \hat{\boldsymbol{W}}_t - \hat{\boldsymbol{R}}_t) \hat{\boldsymbol{\Upsilon}} \cdot \hat{\boldsymbol{e}}_{it} \hat{\boldsymbol{e}}_{it}' \right].\] The bias-corrected APE estimator is then defined as \[ \hat{\delta}_{ABC} ( \boldsymbol{x}^{(0)},\boldsymbol{x}^{(1)} ) =\hat{\delta} ( \boldsymbol{x}^{(0)},\boldsymbol{x}^{(1)} ) - \frac{ \hat{\boldsymbol{\gamma}}' \hat{\boldsymbol{\Delta}}^{-1} \hat{ \bm{b}} + \hat{b}^3 +\hat{b}^4}{T} -\frac{ \hat{\boldsymbol{\gamma}}' \hat{\boldsymbol{\Delta}}^{-1} \hat{\bm{d}} + \hat{d}^3 +\hat{d}^4}{N}. \] It can be shown that the above bias-corrected estimators are free of asymptotic biases.
Following dhaene2015split, fernandez2016individual and chen2020nonlinear, bias correction can also achieved by a simple procedure called split-panel jackknife (SPJ, hereafter). Let $\hat{\boldsymbol{\beta}}_{N/2,T}^1$ and $\hat{\boldsymbol{\beta}}_{N/2,T}^2$ be the CCE estimators using subsamples $\{ (i,t):i=1,\ldots, N/2; t=1,\ldots,T\}$ and $\{ (i,t):i=N/2+1,\ldots, N; t=1,\ldots,T\}$ respectively. Similarly, let $\hat{\boldsymbol{\beta}}_{N,T/2}^1$ and $\hat{\boldsymbol{\beta}}_{N,T/2}^2$ be the CCE estimators using subsamples $\{ (i,t):i=1,\ldots, N; t=1,\ldots,T/2\}$ and $\{ (i,t):i=1,\ldots, N; t=T/2+1,\ldots,T\}$ respectively. Define \[ \hat{ \boldsymbol{\beta}}_{SPJ} = 3\hat{\boldsymbol{\beta}} -\frac{1}{2}\left(\hat{\boldsymbol{\beta}}_{N/2,T}^1+\hat{\boldsymbol{\beta}}_{N/2,T}^2 \right) -\frac{1}{2}\left(\hat{\boldsymbol{\beta}}_{N,T/2}^1+\hat{\boldsymbol{\beta}}_{N,T/2}^2\right) . \] Then under some type of homogeneity conditions to guarantee that the asymptotic biases of $ \hat{\boldsymbol{\beta}}_{N/2,T}^1,\hat{\boldsymbol{\beta}}_{N/2,T}^2, \hat{\boldsymbol{\beta}}_{N,T/2}^1,\hat{\boldsymbol{\beta}}_{N,T/2}^2$ and $\hat{\boldsymbol{\beta}}$ all converge to the same limit, it can be shown that \[ \sqrt{NT} (\hat{\boldsymbol{\beta}}_{SPJ} -\boldsymbol{\beta}_0) \overset{d}{\rightarrow} \mathcal{N} \left( 0,\boldsymbol{\Delta}^{-1}\boldsymbol{\Omega} \boldsymbol{\Delta}^{-1} \right) \text{ as }N,T\rightarrow\infty. \] A bias-corrected estimator using the SPJ for the APE can be defined in a similar fashion.
For the CCE estimator, define \[ \hat{\boldsymbol{C}}_t = \frac{1}{N}\sum_{i=1}^{N} \hat{l}_{it}^{(2)} \hat{\dot{\boldsymbol{x}}}_{it} \hat{\boldsymbol{\lambda}}_{i}', \quad \hat{\boldsymbol{w}} _{it} = \hat{l}_{it}^{(1)} \hat{\dot{\boldsymbol{x}}}_{it} + \hat{\boldsymbol{C}}_t \hat{\boldsymbol{\Upsilon}} \hat{\boldsymbol{e}}_{it}, \quad \hat{\boldsymbol{\Omega}} =\frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}\sum_{s=1}^{T} \hat{\boldsymbol{w}} _{it}\hat{\boldsymbol{w}} _{is}'k\left(\frac{t-s}{L}\right) .\] For the APE estimator, define \[\hat{v}_{it} = \hat{\boldsymbol{\gamma}}' \hat{\boldsymbol{\Delta}}^{-1}\hat{\boldsymbol{w}}_{it} + (\hat{\boldsymbol{R}}_t \hat{\boldsymbol{f}}_{t} -\hat{\boldsymbol{\gamma}}_t)' \hat{\boldsymbol{\Upsilon}} \hat{\boldsymbol{e}}_{it} + \hat{l}_{it}^{(1)}\hat{\boldsymbol{\gamma}}_i' \hat{\boldsymbol{A}}_i^{-1}\hat{\boldsymbol{f}}_{t}, \quad \hat{\sigma}^2 = \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}\sum_{s=1}^{T} \hat{v}_{it}\hat{v}_{is}k\left(\frac{t-s}{L}\right). \] It can be show that:
The above result ensures that we can make asymptotically valid inference based on the estimated variances and bias-corrected estimators.
In this section, the finite sample performance of the proposed estimation method is evaluated using Monte Carlo simulations. The data generating process (DGP) of $y_{it}$ is given by: \[ y_{it} = \boldsymbol{1} \{ x_{it,1} +x_{it,2}+x_{it,3}+x_{it,4}+\lambda_{i,1}f_{t,1} + \lambda_{i,2}f_{t,2} - \epsilon_{it} \geq 0 \} , \] where $\epsilon_{it}$ are i.i.d with the standard logistic distributions, $f_{t,1} = 0.3+0.7 f_{t-1,1} + u_{1t} , f_{t,2} = 0.6+0.4 f_{t-1,2} + u_{2t},$ $u_{1t},u_{2t}\sim \text{i.i.d } \mathcal{N}(0,1)$ and $\lambda_{i,1},\lambda_{i,2}\sim \text{i.i.d } \mathcal{N}(1,1)$. In addition, the covariates are generated by \[ x_{it,1} = \theta_{1i}f_{t,1}+f_{t,2}+e_{it,1} , \quad x_{it,2} = \theta_{2i} f_{t,2} +e_{it,2}, \quad x_{it,3} = 1.5e_{it,3}, \quad x_{it,4} = e_{it,4},\] where $\theta_{1i},\theta_{2i}\sim \text{i.i.d } \mathcal{N}(1,1)$. As for $e_{it,j}$, $j=1,2,3,4$, two cases are considered : (i) $e_{it,j}\sim \text{i.i.d } \mathcal{N}(0,1)$; (ii) $e_{it,j} = 0.6e_{i,t-1,j}+h_{it,j}$ where $h_{it,j}\sim \text{i.i.d } \mathcal{N}(0,1)$. For the first case, there is no serial dependence in $(y_{it},\bm{x}_{it})$ conditional on $\{\bm{\lambda}_i\}$ and $\{\bm{f}_t\}$. For the second case, we need to take into account the time-series dependence when constructing the bias-corrected estimators and estimating the asymptotic variances.
Our focus is whether the bias correction methods proposed in Section 3.3 can effectively reduce the bias of the CCE estimator, and therefore improve the empirical coverage rate of the confidence interval. To this end, we compare three estimators: the CCE estimator without bias correction $\hat{\bm{\beta}}$, the CCE estimator with analytical bias correction $\hat{\bm{\beta}}_{ABC}$, and the CCE estimator with SPJ bias correction $\hat{\bm{\beta}}_{SPJ}$. Table 2 below reports the biases and standard errors of these three estimators for the above model, along with the empirical coverage rates of their confidence intervals from 500 replications for $N,T\in\{50,100,200\}$. For the DGP without serial dependence, we set $L=0$ and the results are reported in the upper panel. For the DGP with serial dependence, the results with $L=1,2,3$ are reported in the lower three panels. Moreover, to save space, we only show results for the coefficient of $x_{it,1}$, and the results for the other three coefficients (which are very similar) are available upon request.
Based on the results presented in Table 2, several key observations can be made. Firstly, both bias correction methods are found to be effective in significantly reducing the biases associated with CCE estimators. The analytical bias correction method is observed to perform better in models without serial dependence, whilst the SPJ method is more effective in models with serial dependence. Secondly, the standard errors of the CCE estimators are found to not be inflated in most cases, and in some instances the standard errors of the bias-corrected estimators are even lower. This suggests that bias correction does not come at the expense of increased uncertainty. Thirdly, the smaller biases of the bias-corrected estimators result in empirical coverage rates of their confidence intervals that are closer to their nominal rates of 95$\%$. Lastly, it is found that increasing $L$ from 1 to 3 does not result in improved performance of the bias-corrected estimators in models with serial dependence, thus supporting the recommendation of hahn2011bias and galvao2016smoothed that using $L=1$ in practice is a reasonably good choice.
In a recent study, ma2019nonfinancial documented that a sizable fraction of financial activities comes from firms that simultaneously issue in one financial market and repurchase in another.\footnote{Most studies discussed the capital market-driven firm financing activities in a single market, see Baker2000, BAKER2003261, HONG2008119 and DHT2012, among others.} For example, using the U.S. data from 1985 to 2015, the author found that about 45$\%$ of equity repurchases in value come from firms that concurrently net issue debt, and about 35$\%$ of net debt issuance comes from the firms that net repurchase equity. The central question of that paper is how such cross-market arbitrage of the nonfinancial firms is affected by relative valuations across debt and equity markets. In this section, we revisit this important question by employing the proposed estimation method in this paper. Our main objective is to compare the estimation results of panel logistic models with only individual effects, as assumed in ma2019nonfinancial, with models featuring interactive effects while also demonstrating the efficacy of the bias correction methods in both type of models.
Following ma2019nonfinancial, let the dependent variable $y_{it}$ be an indicator that identifies instances in which a firm $i$ both issues debt and repurchases equity at time period $t$: $y_{it} = \mathbf{1}\{ s_{it}>0, d_{it}>0 \} $, where $s_{it}$ is the net equity repurchases, defined as the net purchase of Common and Preferred Stock (i.e., PRSTKC-SSTK) of firm $i$ in quarter $t$, and $d_{it}$ is the net debt issuance, defined as long-term debt issuance (DLTIS) minus long-term debt reduction (DLTR).
The explanatory variables in this study consist of three measures of valuations in both the debt and equity markets. To measure the debt market valuation, firm-level spreads are constructed since most firms have more than one outstanding bond. A bond's credit spread is defined as the yield difference between its yield and the contemporaneous yield on the nearest-maturity Treasury, while the term spread is defined as the yield difference between the nearest-maturity Treasury and the three-month Treasury bill. The firm-level credit and term spreads are then calculated as the equal-weighted averages of its outstanding bonds' spreads. Meanwhile, the book-to-market ratio (B/P) is used as a measure of a firm's valuation condition in the equity market.\footnote{As discussed in ma2019nonfinancial and DHT2012, the value-to-price ratio is more appropriate if we want to study the effect of equity market valuation on $y_{it}$. We use the book-to-market ratio here mainly because we don't have access to some of the datasets needed to construct the value-to-price ratio.} Other firm-level variables that may impact financing decisions, such as net income, cash holding, financing flows driven by investment plans (CAPX), deviations from target capital structure (measured by ex ante distance to target leverage), asset growth (which captures a firm's expansion tendency), and firm size, will also be controlled for in the analysis.\footnote{See ma2019nonfinancial for detailed construction of these variables.}
Our analysis draws on three primary data sources: Compustat for firm-level balance sheet and cash flow variables, CRSP for equity market valuation data, and the Trade Reporting and Compliance Engine (TRACE) database for bond pricing data. As per established literature, all flow variables, such as net issuance, in a given quarter are normalized by lagged assets at the end of the previous quarter, while all stock variables, such as cash holdings, are normalized by assets in the same quarter. Our final sample consists of quarterly data for 145 nonfinancial firms over the period 2013 to 2022, corresponding to a balanced panel with $T=40$ and $5800$ total observations.
We consider the following models and estimators.
Model A: (only individual effects) $\mathbb{E}[ y_{it} | \boldsymbol{x}_{it},\alpha_i] = 1/(1+e^{-\bm{\beta}_0'\bm{x}_{it}-\alpha_i } )$ \\ (A1) Conditional MLE (same as ma2019nonfinancial);\\ (A2) fixed-effects estimator, analytical bias correction with $L = 1$ (hahn2004jackknife);\\ (A3) fixed-effects estimator, SPJ bias correction (dhaene2015split).
Model B: (interactive effects) $\mathbb{E}[ y_{it} | \boldsymbol{x}_{it},\boldsymbol{\lambda}_{i},\boldsymbol{f}_{t}] = 1/(1+e^{-\bm{\beta}_0'\bm{x}_{it}-\boldsymbol{\lambda}_{i}'\boldsymbol{f}_{t} } )$ \\ (B1) CCE estimator, no bias correction;\\ (B2) CCE estimator, analytical bias correction with $L= 1$;\\ (B3) CCE estimator, SPJ bias correction.
Model A is the traditional panel logistic model with only individual effects, used as the benchmark model for comparison. For this model, we consider three estimation methods: (A1) the classical conditional maximum likelihood estimator of andersen1970asymptotic, which is also the estimator used by ma2019nonfinancial; (A2) the fixed-effects estimator proposed by hahn2004jackknife, where the asymptotic bias is corrected using analytical bias correction; (A3) the same fixed-effects estimator with SPJ bias correction proposed by dhaene2015split.
Compared with Model A, Model B introduces interactive fixed effects and the CCE framework. As in the simulations, three different estimators proposed in this paper are considered: (B1) the CCE estimator without bias correction; (B2) the CCE estimator with analytical bias correction where the bandwidth parameter is chosen as $L=1$; (B3) the CCE estimator with SPJ bias correction. Using the method proposed in Section 2.2 with $P_{NT}= \min\{N,T\}^{-1/3}$, the estimated number of factors is equal to 4.
Table 3 below displays our estimation results. Overall, the results are consistent with those of ma2019nonfinancial, despite our use of a different sample period. Specifically, coefficients for the credit spread and term spread are negative, while the coefficient for B/P ratio is positive. These findings suggest that concurrently issuing debt and repurchasing equity is more likely when the cost of debt is low and the cost of equity is high.
However, in Model B, the coefficients of the credit spread and B/P ratio are significantly larger in absolute value, while the coefficient for the term spread is less robust across different estimation methods. Notably, in most coefficients of Model B, significant differences are evident between the CCE estimators and their bias-corrected counterparts. Overall, our empirical application underscores the importance of interactive fixed effects in nonlinear panel data models and highlights the efficacy of the proposed bias correction methods.
In this paper, we introduced a novel CCE estimator for nonlinear panel data models with interactive fixed effects and homogeneous coefficients, filling an important gap in the literature. We established the theoretical properties of the estimator and demonstrated the effectiveness of both analytical and split-panel jackknife methods for eliminating its asymptotic bias. Monte Carlo simulations confirmed that the proposed estimators provide reliable results in finite samples. We also presented an empirical application that shows the applicability of our method for analyzing the cross-market arbitrage behavior of nonfinancial firms. Although our proposed method shows promising results, further research is needed to investigate the optimal choice of bandwidth parameter for the HAC estimators in the presence of serial dependence. Overall, our findings suggest that our proposed estimator is a useful tool for empirical researchers interested in analyzing nonlinear panel data models with interactive fixed effects.
{
}