Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,590 characters · 20 sections · 85 citation commands
Low-rank Panel Quantile Regression: Estimation and Inference
\ \ \ \
Panel quantile regressions are widely used to estimate the conditional quantiles, which can capture the heterogeneous effects that may vary across the distribution of the outcomes. Such effects are usually assumed to be homogeneous across individuals and over time periods. However, in empirical analyses, it is usually unknown whether the slope coefficients are homogeneous across individuals and/or time. Mistakenly forcing slopes to be homogeneous across time and individuals may lead to inconsistent estimation and misleading inferences. This prompts two questions to be answered: how can we estimate the true model at different quantiles when we allow for heterogeneous slopes across individuals and time at the same time? How to conduct specification tests for homogeneous effects over individuals or time and tests for the additive structure of the slope coefficients?
To answer the first question, we propose an estimation procedure for heterogeneous panel quantile regression models where we allow the fixed effects to be either additive or interactive, and the slope coefficients to be heterogeneous over both individuals and time. We impose a low-rank structure for both the intercept and slope coefficient matrices and estimate them via nuclear norm regularization (NNR) followed by the sample splitting, row- and column-wise quantile regressions and debiasing steps. The estimation algorithm is inspired by chernozhukov2019inference, where the main difference is that we split the full sample into three subsamples rather than two because we need certain uniform results which require independence of regressors and regressand used in the debiasing step, and we do not have the closed form for the quantile regression estimates. At last, we derive the asymptotic distributions for the estimators of the factors and factor loadings associated with slope coefficient matrices.
To answer the second question, under the case when the rank of slope coefficient matrix equals one, we conduct sup-type specification tests for homogeneous effects over individuals or time following the lead of castagnetti2015inference and lu2021uniform. We show that our sup-test statistics follow the Gumbel distribution under the null, and the tests have non-trivial power against certain classes of local alternatives. Under the case when the rank of slope matrix equals two, our sup-type test statistic is also shown to follow the Gumbel distribution under the null that the slope coefficient exhibits an additive structure.
This paper relates to three bunches of literature. First, we contribute to the large literature on panel quantile regressions (PQRs). Since koenker2004quantile studied the PQRs with individual fixed effects, there has been an increasing number of papers on PQRs. galvao2010penalized, kato2012asymptotics, galvao2015efficient, galvao2016smoothed, machado2019quantiles, and galvao2020unbiased study the asymptotics for PQRs with individual fixed effects. chen2021quantile study quantile factor models and chen2019two considers PQRs with interactive fixed effects (IFEs). We complement the literature by allowing for unobserved heterogeneity in the slope coefficients of PQRs.
Second, our paper also pertains to slope heterogeneity in panel data models. Latent group structures across individuals and structural changes over time are two common types of slope heterogeneity that have received vast attention in the literature. To recover the unobserved group structures, various methods have been proposed. For example, lin2012estimation, bonhomme2015grouped and ando2016panel use the K-means algorithm; su2016identifying propose the C-lasso algorithm which is further studied and extended by su2018identifying, su2019sieve and wang2019heterogeneous; wang2018homogeneity propose an clustering algorithm in regression via data-driven segmentation called CARDS; wang2021identifying propose a sequential binary segmentation algorithm to identify the latent group structures in nonlinear panels. Recent literature on the estimation with structural changes in panel data models includes, but is not limited to, chen2015estimating, cheng2016shrinkage, ma2018estimation, baltagi2021estimating. In addition, galvao2018testing and zhang2019quantile consider individual heterogeneity in PQRs while they assume homogeneity across time. To allow for both latent groups and structural breaks, okui2021heterogeneous study a linear panel data model with individual fixed effects where each latent group has common breaks and the breaking points can be different across different groups, and they propose a grouped adaptive group fused lasso (GAGFL) approach to estimate slope coefficients. lumsdaine2021estimation consider a linear panel data model with a grouped pattern of heterogeneity where the latent group membership structure and/or the values of slope coefficients can change at a breaking point, and they propose a K-means-type estimation algorithm and establish the asymptotic properties of the resulting estimators. Compared with the models studied above, our model combines both individual and time heterogeneity and only requires certain low-rank structure in the slope coefficient matrix. So the unobserved heterogeneity takes a more flexible form in our model than those in the literature such as okui2021heterogeneous and lumsdaine2021estimation.
Last, our paper also connects with the burgeoning literature on nuclear norm regularization. Such a method has been widely adopted to study panel and network models. See, alidaee2020recovering, athey2021matrix, bai2019rank, belloni2019high, chen2020noisy, chernozhukov2019inference, feng2019regularized, Hong_Su_Jiang2022, Miao_Phillips_Su2022, among others. In the least squares panel framework, moon2018nuclear consider a homogeneous panel with IFEs by using NNR-based estimator as an initial estimator to construct iterative estimators that are asymptotically equivalent to the least squares estimators; chernozhukov2019inference study a heterogenous panel where both the intercept and slope coefficient matrices exhibit a low-rank structure and establish the asymptotic distribution theory based on NNR. In the presence of endogeneity, Hong_Su_Jiang2022 proposes a profile GMM method to estimate panel data models with IFEs. In the panel quantile regression setting, feng2019regularized develops error bounds for the low-rank estimates in terms of Frobenius norms under independence assumption; belloni2019high relaxes the independence assumption to the $\beta $-mixing condition along the time dimension. Our paper extends chernozhukov2019inference from the least squares framework to the PQR framework, derives the asymptotic distribution theory and develops various specification tests under some strong mixing conditions along the time dimension that is weaker than the $\beta $-mixing condition. We also rely on the sequential symmetrization technique developed by rakhlin2015sequential to obtain the convergence rates of the nuclear norm regularized estimators.
The rest of the paper is organized as follows. We first introduce the low-rank structure PQR model and the estimation algorithm in Section 2. We study the asymptotic properties of our estimators in Section 3. In Section 4, we propose two specification tests: one for the no-factor structure and one for the additive structure, and study the asymptotic properties of the test statistics. In Section 5, we show the finite sample performance of our method via Monte Carlo simulations. In Section 6, we apply our method to two datasets: one is to study how Tobin's q and cash flows affect corporate investment and whether firm's external investment to its internal financing exhibits heterogeneity structure, and the other is to study the relationship between economics growth, foreign direct investment and unemployment. Section 7 concludes. All proofs are related to the online supplement.
Notation. $\left\Vert \cdot \right\Vert _{1}$, $\left\Vert \cdot \right\Vert _{op}$, $\left\Vert \cdot \right\Vert _{\infty }$, $\left\Vert \cdot \right\Vert _{\max }$ $\left\Vert \cdot \right\Vert _{2}$, $\left\Vert \cdot \right\Vert _{F}$, $\left\Vert \cdot \right\Vert _{\ast }$ denote the matrix norm induced by 1-norms, the matrix norm induced by 2-norms, the matrix norm induced by $\infty $-norms, the maximum norm, the Euclidean norm, the Frobenius norm and the nuclear norm. $\odot $ is the element-wise product. $\lfloor \cdot \rfloor $ and $\lceil \cdot \rceil $ denote the floor and ceiling functions, respectively. $a\vee b$ and $a\wedge b$ return the max and the min of $a$ and $b,$ respectively. The symbol $\lesssim $ means \textquotedblleft the left is bounded by a positive constant times the right\textquotedblright . Let $A=\{A_{it}\}_{i \in [n], t\in [T]}$ be a matrix with its $(i,t)$-th entry denoted as $A_{it}$, where $[n]$ to denote the set $ \{1,\cdots ,n\}$ for any positive integer $n$. Let $\{A_{j}\}_{j=0}^{p}$ denote the collection of matrices $A_{j}$ for all $j\in \{0,\cdots ,p\}$. When $A$ is symmetric, $ \lambda _{\max }(A)$ and $\lambda _{\min }(A)$ denote its largest and smallest eigenvalues, respectively. The operators $\rightsquigarrow $ and $ \operatornamewithlimits{\to}\limits^{p}$ denote convergence in distribution and in probability, respectively. Besides, we use w.p.a.1 and a.s. to abbreviate \textquotedblleft with probability approaching 1\textquotedblright\ and \textquotedblleft almost surely\textquotedblright , respectively.
In this section, we introduce the PQR model and estimation algorithm.
Consider the PQR model
where $i\in \left[ N\right] ,$ $t\in \left[ T\right] ,$ $\tau \in (0,1)$ is the quantile index, $Y_{it}$ is the dependent variable, $X_{j,it}$ is the $j$ -th regressor for individual $i$ at time $t$, $\{\Theta _{j,it}^{0}\}_{j\in \lbrack p]}$ is the corresponding slope coefficient, $\Theta _{0,it}^{0}$ is the intercept, and $\mathscr{Q}_{\tau }\left( Y_{it}\bigg|\left\{ X_{j,it}\right\} _{j\in \lbrack p],t\in \lbrack T]},\left\{ \Theta _{j,it}^{0}\left( \tau \right) \right\} _{j\in \lbrack p]\cup \{0\},t\in \lbrack T]}\right) $ denotes the conditional $\tau $-quantile of $Y_{it}$ given the regressors $\left\{ X_{j,it}\right\} _{j\in \lbrack p],t\in \lbrack T]}$ and\ the parameters $\left\{ \Theta _{j,it}^{0}\left( \tau \right) \right\} _{j\in \lbrack p]\cup {0},t\in \lbrack T]}$.\footnote{ We will assume that both the intercept term $\Theta _{0,it}^{0}$ and the slope coefficients $\{\Theta _{j,it}^{0}\}_{j\in \lbrack p]}$ have low-rank structures, and follow the convention in the panel data literature by treating the factors to be random. Therefore, $\{\Theta _{j,it}^{0}\}_{j\in \lbrack p]\cup \{0\}}$ are random as well.} Alternatively, we can rewrite the above model as
where $\epsilon (\tau)$ is the idiosyncratic error matrix with the $(i,t)$ -th entry being $\epsilon _{it} (\tau)$. Similarly, $X_{j}$, $\Theta _{j}\left( \tau \right) $, and $Y$ are matrices with the $(i,t)$-th entry being $X_{j,it}$, $\Theta _{j,it}\left( \tau \right) $, and $Y_{it}$, respectively. In this model, we assume $p$, the number of regressors, is fixed and both $N$ and $T$ pass to infinity. In Assumption (ref) below, we characterize the dependence of the data, under which (ref) holds.
In the paper, we focus on the panel quantile regression for a fixed $\tau $ and thus suppress the dependence of $\Theta _{j}^{0}(\tau )$ and $\epsilon (\tau )$ on $\tau $ for notation simplicity. In addition, we impose low-rank structures for the intercept and slope matrices, i.e., $\text{rank}(\Theta _{j}^{0})=K_{j}$ for some positive constant $K_{j}$ and for each $j\in \{0,\cdots ,p\}$. By the singular value decomposition (SVD), we have
where $\mathcal{U}_{j}^{0}\in \mathbb{R}^{N\times K_{j}}$, $\mathcal{V} _{j}^{0}\in \mathbb{R}^{T\times K_{j}}$, $\Sigma _{j}^{0}=\text{diag}(\sigma _{1,j},\cdots ,\sigma _{K_{j},j})$, $U_{j}^{0}=\sqrt{N}\mathcal{U} _{j}^{0}\Sigma _{j}^{0}$ with each row being $u_{i,j}^{0\prime }$, and $ V_{j}^{0}=\sqrt{T}\mathcal{V}_{j}^{0}$ with each row being $v_{t,j}^{0\prime }$.
The low-rank structure assumption includes several popular cases. For the intercept term, one commonly assumes that $\Theta _{0,it}^{0}$ to take the forms $\alpha _{i}^{0},$ $\mu _{t}^{0},$ or $\alpha _{i}^{0}+\mu _{t}^{0}$ in classical PQRs. Then the matrix $\Theta _{0}^{0}$ has rank 1, 1, and 2, respectively. It is also possible to assume $\Theta _{0,it}^{0}$ to take an interactive form, say, $\Theta _{0,it}^{0}=\lambda _{0,i}^{0\prime }f_{0,t}^{0},$ where both $\lambda _{0,i}^{0}$ and $f_{0,t}^{0}$ are $K_{0}$ -vectors. For the slope matrix $\Theta _{j}^{0}$, $j\in \left[ p\right] ,$ the early PQR models frequently assume that $\Theta _{j,it}^{0}$ is a constant across $\left( i,t\right) \ $ to yield a homogenous PQR model. Obviously, such a model is very restrictive by assuming homogenous slope coefficients. It is possible to allow the slope coefficients to change over either $i$, or $t$, or both. See the following examples for different low-rank structures.
Example 1. When $\Theta _{j,it}^{0}=\Theta _{j,i}^{0}$ $ \forall t\in \left[ T\right] ,$ or $\Theta _{j,it}^{0}=\Theta _{j,t}^{0}$ $ \forall i\in \left[ N\right] ,$ or $\Theta _{j,it}^{0}=\Theta _{j}^{0}$ $ \forall \left( i\text{,}t\right) \in \left[ N\right] \times \left[ T\right] $ , and this holds for all $j\in \left[ p\right] ,$ we have the PQR models with only individual heterogeneity, with only time heterogeneity, and with homogeneity, respectively. We observe that $K_{j}=1$ for these three cases.
Example 2. When $\Theta _{j,it}^{0}=\lambda _{j,i}^{0}+f_{j,t}^{0}$, we notice that
Let $\Sigma _{A,j}:=A_{j}^{\prime }A_{j}$ and $\Sigma _{B,j}:=B_{j}^{\prime }B_{j}$. Let $\Sigma _{A,j}^{\frac{1}{2}}$ (resp. $\Sigma _{B,j}^{\frac{1}{2} }$) be the symmetric square root of $\Sigma _{A,j}$ (resp. $\Sigma _{B,j}$). By eigendecomposition, we have $\Sigma _{A,j}^{\frac{1}{2} }=P_{j,1}S_{j,1}P_{j,1}^{\prime }$ and $\Sigma _{B,j}^{\frac{1}{2} }=P_{j,2}S_{j,2}P_{j,2}^{\prime }$. Besides, we apply singular value decomposition to matrix $S_{j,1}P_{j,1}^{\prime }P_{j,2}S_{j,2}$: $ S_{j,1}P_{j,1}^{\prime }P_{j,2}S_{j,2}=Q_{j,1}R_{j}Q_{j,2}^{\prime }$. Then it follows that
where $\mathcal{U}_{j}^{0}=A_{j}\Sigma _{A,j}^{-\frac{1}{2}}P_{j,1}Q_{j,1}$, $\Sigma _{j}^{0}=R_{j}$ and $\mathcal{V}_{j}^{0}=B_{j}\Sigma _{B,j}^{-\frac{1 }{2}}P_{j,2}Q_{j,2}$. Given $P_{j,1}$, $P_{j,2}$, $Q_{j,1}$ and $Q_{j,2}$ are orthonormal matrices, it's easy to that $\mathcal{U}_{j}^{0}$ and $ \mathcal{V}_{j}^{0}$ are also orthonormal so that $\mathcal{U}_{j}^{0\prime } \mathcal{U}_{j}^{0}=\mathcal{V}_{j}^{0\prime }\mathcal{V}_{j}^{0}=I_{2}$. When $j=0$, $\left\{ \lambda _{0,i}^{0}\right\} _{i=1}^{N}$ and $\left\{ f_{0,t}^{0}\right\} _{t=1}^{T}$ are usually referred to as the individual and time fixed effects, respectively, so that the intercept term exhibits an additive fixed effects structure.
Example 3. Let $\Theta _{j,it}^{0}=\sum_{k\in \lbrack K_{j,t}]}\alpha _{j,kt}\mathbf{1}\{i\in G_{j,kt}\}$, where $\left\{ G_{j,kt}\right\} $ forms a partition of $[N]$ for each specific time $t$ and $K_{j,t}$ is the number of groups at time $t$. Moreover, let
where $K_{j}^{(1)}$ and $K_{j}^{(2)}$ are the number of groups before and after the break point $T_{b}$. If $K_{j}^{(1)}=K_{j}^{(2)}$, it is clear that $rank(\Theta _{j}^{0})=1$. If the group structure does not change after the break but $\alpha _{j,k}^{(1)}=c\alpha _{j,k}^{(2)}$ for some constant $ c $, we also have $rank(\Theta _{j}^{0})=1$. Except for these two cases, we can show that
where $\iota _{T_{b}}$ is a $T_{b}\times 1$ vector of ones and $\mathbf{0} _{T_{b}}$ is a $T_{b}\times 1$ vector of zeros. In this case, we notice that $rank(\Theta _{j}^{0})=2$.
Example 4. When $\Theta _{j,it}^{0}=\lambda _{j,i}^{0\prime }f_{j,t}^{0}$ with $\lambda _{j,i}^{0}$ and $f_{j,t}^{0}$ being two $K_{j}$-vectors, we have the IFEs structure. This is the most general example without further restrictions.
Like chernozhukov2019inference, we assume that for each $j\in \left[ p \right] ,$ $X_{j,it}$ exhibits a factor structure: $X_{j,it}=\mu _{j,it}+e_{j,it}=l_{j,i}^{0\prime }w_{j,t}^{0}+e_{j,it},$ where $w_{j,t}^{0}$ and $l_{j,i}^{0}$ are the factors and factor loadings of dimension $r_{j}.$
In this subsection we provide the estimation algorithm by assuming that $ K_{j}$ are all known for all $j$. In the next subsection, we will introduce a rank estimation method to estimate $K_{j}$ consistently.
Define the check function $\rho _{\tau }(u)=u\left( \tau -\mathbf{1}\{u\leq 0\}\right) $. The estimation procedure goes as follows:
In order to obtain the final estimators for the full sample, we propose to switch the role of each subsample for the low-rank estimation, row- and column-wise QR and debiasing, then repeat Steps 1-3 to obtain $\left\{ \hat{u }_{i,j}^{(a,b)}\right\} _{j=0}^{p}$ and $\left\{ \hat{v}_{t,j}^{(a,b)}\right \} _{j=0}^{p}$ for $a\in \lbrack 3]$ and $b\in \lbrack 3]\setminus \{a\}$. Here $(a,b)$ denotes the final estimates for subsample $I_{a}$ obtained from the first step NNR estimates with subsample $I_{b}$. Table (ref) shows the final estimators we obtain by using different combination of subsamples.
Several remarks are in order. First, we randomly split the full sample into three subsamples, each playing a significant role in the algorithm. We use the first subsample for the low-rank estimation to obtain the preliminary NNR estimators of the submatrices of the intercept and slope matrices. But these estimators are only consistent in terms of Frobenius norm, and one cannot derive the pointwise or uniform convergence rates for them. With the low-rank estimates, we use the second subsample to do the row- and column-wise QRs and can now establish the uniform convergence rates for each row of factor and factor loading estimators. Then we use the remaining subsample to debias the second-stage estimator and to obtain the final estimators that have the desirable asymptotic properties.
Second, to reduce the randomness of sample splitting, one can run the estimation algorithm several times with different splittings in practice. Once one obtains factor and factor loading estimates, one can construct estimators for $\Theta _{j}^{0}$ under different splittings and then choose the one specific splitting which yields the minimum quantile objective function.
Third, the bias in the second-stage estimator is inherent from the first-stage NNR estimator. We follow the lead of chernozhukov2019inference to assume that $X_{j,it}$ has a factor structure with an additive idiosyncratic term, and remove the bias by a QR with the demeaned $X_{j,it}$ as regressors. In the least squares panel regression framework, the objective function is smooth and one has closed-form solutions in the last stage so that chernozhukov2019inference only need to split the sample into two subsamples. In contrast, in the PQR framework, the objective function is non-smooth, we do not have closed-form solutions in any stage. In order to remove the bias from the early stage estimation and to derive the distributional results, we need to split the sample into three subsamples.
To save space, we relegate the detailed algorithm for the nuclear norm regularization to the online supplement.
In this subsection we discuss how to estimate the ranks $K_{j}$ consistently. To estimate the ranks, we consider the full sample NNR QR estimation:
For $j\in \left\{ 0,\cdots ,p\right\} $, we estimate $K_{j}$ by the popular singular value thresholding (SVT) as follows
It is standard to show that $\mathbb{P}(\hat{K}_{j}=K_{j})\rightarrow 1$ as $ \left( N,T\right) \rightarrow \infty $ under some regularity conditions given in the next section; see also Proposition D.1 in chernozhukov2019inference and Theorem 2 in Hong_Su_Jiang2022. Since the ranks can be estimated consistently, we assume that they are known in the asymptotic theory below.
In this section, we study the asymptotic properties of the estimators introduced in the last section.
Recall that $X_{j,it}=\mu _{j,it}+e_{j,it}=l_{j,i}^{0\prime }w_{j,t}^{0}+e_{j,it}$ for each $j\in \lbrack p]$. Let $ X_{it}=(X_{1,it},...,X_{p,it})^{\prime }$ and $ e_{it}=(e_{1,it},...,e_{p,it})^{\prime }.$ Define $\epsilon _{i}=\left( \epsilon _{i1},\cdots ,\epsilon _{it}\right) ^{\prime }$, $e_{j,i}=\left( e_{j,i1},\cdots ,e_{j,iT}\right) ^{\prime }$, $W_{j}^{0}$ as the $T\times r_{j}$ matrix with each row being $w_{j,t}^{0\prime }$, and $V_{j}^{0}$ as the $T\times K_{j}$ matrix with each row being $v_{t,j}^{0\prime }$. Further define $a_{it}=\tau -\mathbf{1}\left\{ \epsilon _{it}\leq 0\right\} $ with $ a_{i}=\left( a_{i1},\cdots ,a_{iT}\right) ^{\prime }$ and $a=\left( a_{1},\cdots ,a_{N}\right) ^{\prime }$. Throughout the paper, we treat the factors $\{v_{t,j}^{0}\}_{t\in \lbrack T],j\in \lbrack p]\cup \{0\}}$ and $ \{w_{j,t}^{0}\}_{t\in \lbrack T],j\in \lbrack p]}$ as random and their loadings $\{u_{i,j}^{0}\}_{i\in \lbrack N],j\in \lbrack p]\cup \{0\}}$ and $ \{l_{j,i}^{0}\}_{i\in \lbrack N],j\in \lbrack p]}$ as deterministic.
Table (ref) defines several $\sigma$-fields. We use $\mathscr{D} $ to denote the minimal $\sigma$-field generated by $\left\{V_{j}^{0}\right \}_{j\in [p]\cup \{0\}}\bigcup \left\{W_{j}^{0}\right\}_{j\in [p]};$ the superscripts $I_{1}$ and $I_{1}\cup I_{2}$ are associated with the first subsample and the first two subsamples, respectively. For example, $ \mathscr{D}_{e_{i}}^{I_{1}}$ denotes the minimal $\sigma$-field generated by $\mathscr{D},$ $\left\{e_{it}\right\}_{t\in \left[T\right] }$ and $\left\{ \epsilon_{it},e_{it}\right\}_{i\in I_{1},t\in \left[T\right] }.$
Let $M$ denote a generic bounded constant that may vary across places. Let $ \mathscr{G}_{i,t-1}$ denote the minimal $\sigma $-field generated by $ \mathscr{D}\cup \{e_{ls}\}_{l\leq i-1,s\in \lbrack T]}\cup \{e_{is}\}_{s\leq t}\cup \{\epsilon _{ls}\}_{l\leq i-1,s\in \lbrack T]}\cup \{\epsilon _{is}\}_{s\leq t-1}$. Let $\mathsf{F}_{it}(\cdot )$ and $\mathsf{f} _{it}(\cdot )$ be the conditional cumulative distribution function (CDF) and probability density function (PDF) of $\epsilon _{it}$ given $\mathscr{G} _{i,t-1}$, respectively. Similarly, let $\mathfrak{F}_{it}(\cdot )$ and $ \mathfrak{f}_{it}(\cdot )$ denote the conditional CDF and PDF of $\epsilon _{it}$ given $\mathscr{D}_{e_{i}};$ $F_{it}(\cdot )$ and $f_{it}(\cdot )$ denote the conditional CDF and PDF of $\epsilon _{it}$ given $\mathscr{D} _{e}.$ Let $\mathsf{f}_{it}^{\prime }\left( \cdot \right) $, $\mathfrak{f} _{it}^{\prime }\left( \cdot \right) ,$ and $f_{it}^{\prime }\left( \cdot \right) $ denotes the first derivative of the density $\mathsf{f}_{it}\left( \cdot \right) $, $\mathfrak{f}_{it}\left( \cdot \right) ,$ and $f_{it}\left( \cdot \right) ,$ respectively.
We make the following assumptions.
Assumptions (ref)(i) imposes conditional independence of the error terms and covariates $X_{j,it}$ given the fixed effects. Assumptions (ref)(ii) imposes the moment condition for QR. Assumptions (ref) (iii) imposes the weak dependence assumption along the time dimension via the use of the notion of conditional strong mixing. See Prakasa_Rao2009 for the definition of conditional strong mixing and Su_Chen2013 for an application in the panel setup. Assumptions (ref) (iv)-(v) essentially imposes some conditions on the moments and tail behavior of the both covariates and errors. Note that we allow $X_{j,it}$ to have an infinite support. Assumptions (ref)(vi)-(viii), which are used in the proofs of Theorems (ref), (ref) and (ref), respectively, specify conditions on the conditional density of $\epsilon _{it}$ given different $\sigma $-fields. Assumption (ref)(ix) imposes some restrictions on $N$, $T$ and $\xi _{N}$ in order to obtain the error bound of NNR estimators and to achieve the unbiasedness. It allows not only the case that $N$ and $T$ diverge to infinity at the the same rate, but also the case that $N$ diverges to infinity not too faster than $T$, and vice versa.
Assumption (ref) is the low-rank assumption for the intercept and slope matrices, which is the key assumption for the NNR. The uniform boundedness of elements of these matrices facilitates the asymptotic analysis, but can be relaxed at the cost of more lengthy argument. See ma2020detecting for a similar condition.
Assumption (ref) imposes some conditions on the singular values of the coefficient matrices. It implies that we only allow pervasive factors when these matrices are written as a factor structure. Such an assumption is common in the literature; see, e.g., Assumption 3 in ma2020detecting.
To introduce the next assumption, we need some notation. Let $\Theta _{j}^{0}=R_{j}\Sigma _{j}S_{j}^{\prime }$ be the SVD for $\Theta _{j}^{0}$. Further decompose $R_{j}=\left( R_{j,r},R_{j,0}\right) $ with $R_{j,r}$ being the singular vectors corresponding to the nonzero singular values, $ R_{j,0}$ being the singular vectors corresponding to the zero singular values. Decompose $S_{j}=\left( S_{j,r},S_{j,0}\right) $ with $S_{j,r}$ and $ S_{j,0}$ defined analogously. For any matrix $W\in \mathbb{R}^{N\times T}$, we define
where $\mathcal{P}_{j}\left( W\right) $ and $\mathcal{P}_{j}^{\bot }\left( W\right) $ are the linear projection of matrix $W$ onto the low-rank space and its orthogonal space, respectively. Let $\Delta _{\Theta _{j}}=\Theta _{j}-\Theta _{j}^{0}$ for any $\Theta _{j}$. With some positive constants $C_{1}$ and $C_{2}$, we define the following cone-like restricted set:
Assumption (ref) parallels the restricted strong convexity (RSC) condition in Assumption 3.1 of chernozhukov2019inference who also provide some sufficient primitive conditions.
For any $j\in \left\{ 0,\cdots ,p\right\} $, define $\tilde{\Delta}_{\Theta _{j}}=\tilde{\Theta}_{j}-\Theta _{j}^{0}$ and $\tilde{\Delta}_{\Theta _{j}}^{(1)}=\tilde{\Theta}_{j}^{(1)}-\Theta _{j}^{0,(1)}$, where $\Theta _{j}^{0,(1)}=\left\{ \Theta _{j,it}^{0}\right\} _{i\in I_{1},t\in \lbrack T]} $. The following theorem establishes the convergence rates of the NNR estimators of the coefficient matrices.
Remark 1. Theorem (ref)(i) reports the \textquotedblleft rough\textquotedblright\ convergence rates of the NNR estimators of the coefficient matrices in terms of Frobenius norm for both the full-sample and sub-sample estimators. Unlike the traditional $\left( N\wedge T\right) ^{-1/2}$-rate in the least squares framework, NNR estimators' convergence rates in the PQR framework usually have an additional $\sqrt{\log (N\vee T)}$ term due to the use of some exponential inequalities. The extra term $\xi _{N}^{2}$\ in our rate is due to the upper bound of $|X_{j,it}|$, and it disappears in case $X_{j,it}$'s are uniformly bounded. Theorem (ref)(ii)-(iii) report the convergence rates for the estimators of the factors and factor loadings of $\Theta _{j}^{0}$, which are inherited from those in Theorem (ref)(i). To derive these results, we establish the symmetrization inequality and contraction principle for the sequential symmetrization developed by rakhlin2015sequential. See Lemmas (ref) and (ref) in the online supplement for more detail.
To study the asymptotic properties of the second-stage estimators, we add some notation. Define
where $\Phi_{it}^{0}=(v_{t,0}^{0\prime},v_{t,1}^{0\prime}X_{1,it},\cdots,v_{t,p}^{0\prime}X_{p,it})^{\prime}$ and $\Psi_{it}^{0}=(u_{i,0}^{0\prime},u_{i,1}^{0\prime}X_{1,it},\cdots,u_{i,p}^{0\prime}X_{p,it})^{\prime}.$ Let $K=\sum_{j=0}^{p}K_{j}.$ Note that $\Phi_{i}$ and $\Psi_{t}$ are $K\times K$ matrices. We add the following two assumptions.
Assumption (ref) is similar to Assumption 8 in ma2020detecting. To introduce Theorem (ref), we define
Theorem (ref) below gives the uniform convergence rate and linear expansion of the factor loading estimators from second stage estimation.
Remark 2. Theorem (ref)(i) reports the uniform convergence rate for the factor loading estimators of $\Theta _{j}^{0}$ for $i\in I_{2}\cup I_{3};$ Theorem (ref)(ii) reports the uniform convergence rate for the factor estimators of $\Theta _{j}^{0}$ for $t\in \left[ T\right] $; Theorem (ref)(iii) reports the linear expansion for the factor loading estimators of $\Theta _{j}^{0}$ for $i\in I_{3}$. However, the $\mathbb{J}_{i}\left( \left\{ \dot{\Delta}_{t,v}\right\} _{t\in \lbrack T]}\right)$ term is not mean-zero and represents the bias induced by the first stage NNR. In the third stage below, we aim to remove such a bias from the linear expansion.
In the debiasing stage, we first apply PCA to all independent variables $X_{j,it}$, and then run the row- and column-wise QRs to obtain the final estimators. Below we give Assumptions (ref)-(ref) for the PCA procedure and establish the asymptotic linear expansions of PCA estimates in the online supplement. Theorem (ref) below gives the asymptotic distribution of our final factor and factor loading estimates.
Assumptions (ref)-(ref) are stronger than those in bai2020simpler because we strengthen their Assumptions A1(c) and A3 to hold uniformly. Assumption (ref) imposes some moment and mixing conditions. Even though $f_{it}(\cdot )$ (the PDF of $\epsilon _{it}$ given $\mathscr{D} _{e}$) is a function of $ \left\{ e_{j,it}\right\} _{j\in \lbrack p],i\in \lbrack N],t\in \lbrack T]}$ , we can show that Assumption (ref) holds under some reasonable conditions. For example, we consider the location scale model:
where $u_{it}$ is independent of $\{w_{j,t}^{0},e_{j,it}\}_{j\in \lbrack p],t\in \lbrack T]}$ and $l_{j,i}^{0}$ and $\beta _{j,it}$ are nonrandom. In this case, $\Theta _{j,it}^{0}=\beta _{j,it}+\gamma _{j,it}\mathscr{Q}_{\tau }(u_{it})$, $\epsilon _{it}=\left( \gamma _{0,it}+\sum_{j\in \lbrack p]}X_{j,it}\gamma _{j,it}\right) \left[ u_{it}-\mathscr{Q}_{\tau }(u_{it}) \right] $, where $\mathscr{Q}_{\tau }(u_{it})$ is the $\tau $-quantile of $ u_{it}$. It is clear that $f_{it}(\cdot )$ is the function of $\left\{ e_{j,it}\right\} _{j\in \lbrack p],i\in \lbrack N],t\in \lbrack T]}$ and all factors. However, if $u_{it}$ is independent of sequence $\left\{ e_{j,it}\right\} _{j\in \lbrack p],t\in \lbrack T]}$, we observe that $ f_{it}(0)$ is the PDF of $u_{it}-\mathscr{Q}_{\tau }(u_{it})$ evaluated at zero point, which is independent of $\left\{ e_{j,it}\right\} _{j\in \lbrack p],t\in \lbrack T]}$. Therefore, Assumption (ref) (i) holds under mild conditions that $u_{it}$ is independent of the sequence $ \left\{ e_{it}\right\} _{t\in \lbrack T]}$ and $\mathbb{E}\left( e_{it}\big| \mathscr{D}\right) =0$.
Define
Let $\Sigma_{u_{j},i}=O_{j}^{(1)}V_{u_{j},i}^{-1} \Omega_{u_{j},i}V_{u_{j},i}^{-1}O_{j}^{(1)\prime}$, $ \Sigma_{v_{j}}^{(3)}=O_{j}^{(1)}\left(V_{v_{j}}^{(3)}\right)^{-1} \Omega_{v_{j}}\left(V_{v_{j}}^{(3)}\right)^{-1}O_{j}^{(1)\prime}$, $ b_{j,it}^{0}=e_{j,it}v_{t,j}^{0}(\tau -\mathbf{1}\left\{ \epsilon_{it}\leq 0\right\} )$ and $\xi_{j,it}^{0}=e_{j,it}u_{i,j}^{0}\left(\tau -\mathbf{1} \left\{ \epsilon_{it}\leq 0\right\} \right) $. The following theorem establishes the asymptotic properties of the third-stage estimators.
Remark 3. Theorem (ref) reports the linear expansions for the factor and factor loading estimators for each slope matrix obtained in Step 3. Compared with chernozhukov2019inference, Theorem (ref) obtains the uniform convergence rate rather than the point-wise result for the reminder terms $\mathcal{R}_{i,u}^{j}$ and $\mathcal{R}_{t,v}^{j}.$ In addition, since the regressors in the debiasing step are obtained from Step 2 instead of Step 1, we don't have independence between the regressors and error terms, which makes the proof more complex than that in chernozhukov2019inference. See the proof in the appendix on how to handle the dependence. Assumption (ref) in the online supplement is a regularity condition on the density of $\epsilon_{it}$.
Following Theorem (ref) and estimators defined in Table (ref), we have that $\forall j\in \left[ p\right] $, $ \forall i\in \lbrack N]$ and $\forall t\in \left[ T\right] $,
where $\hat{V}_{v_{j},t}^{(a)}=\frac{1}{N_{a}}\sum_{i\in I_{a}}f_{it}(0)e_{j,it}^{2}u_{i,j}^{0}u_{i,j}^{0\prime }$, $a\in \lbrack 3]$ and $b\in \lbrack 3]\setminus \{a\}$.
Given the above estimates for the factors and factor loadings, we can estimate $\Theta _{j,it}^{0}$ by
where $\mathbf{1}_{ia}=\mathbf{1}\left\{ i\in I_{a}\right\} $ for $i\in \lbrack N]$. Let $\Xi _{j,it}^{0}=\frac{1}{T}v_{t,j}^{0\prime }\Sigma _{u_{j},i}v_{t,j}^{0}+\sum_{a=1}^{3}\frac{1}{N_{a}}\mathbf{1} _{ia}u_{i,j}^{0\prime }\Sigma _{v_{j}}^{a}u_{i,j}^{0}$. The following proposition studies the asymptotic properties of $\hat{\Theta}_{j,it}$.
Remark 4. Proposition (ref) establishes the distribution theory for the slope estimators. Recall that we remove the principle component from the independent variables $X_{j,it}$ which is the key point in the debiasing step and why we don't have the distribution theory result for the intercept estimates $\hat{\Theta}_{0,it}$ in the current framework. However, once we have the distribution theory for the slope estimates, we can follow chen2021quantile and obtain a new estimator for $\Theta_{0,it}^0$ from the smoothed quantile regression and establish its distribution theory. We leave this for the further research.
To make inference for $u_{i,j}^{0},$ $v_{t,j}^{0},$ and $\Theta _{j,it}^{0},$ one needs to estimate their asymptotic variances $\Sigma _{u_{j},i},$ $\Sigma _{v_{j}}$ and $\Xi _{j,it}^{0}$ consistently. Let $k(\cdot )$ be a PDF-type kernel function and $K(\cdot )$ be its survival function such that $\int k(u)du=1$ and $K(u):=\int_{u}^{\infty }k(v)dv$. Let $h_{N}$ be the bandwidth such that $h_{N}\rightarrow 0$ with $N\rightarrow \infty $. Define $ K_{h_{N}}(\cdot )=K(\frac{\cdot }{h_{N}})$, $k_{h_{N}}(\cdot )=\frac{1}{h_{N} }k(\frac{\cdot }{h_{N}})$. Let $\hat{\epsilon}_{it}=Y_{it}-\hat{\Theta} _{0,it}-\sum_{j\in \lbrack p]}X_{j,it}\hat{\Theta}_{j,it}$, $\hat{\mathrm{v}} _{t,s,j}=\frac{1}{6}\sum_{a\in \lbrack 3]}\sum_{b\in \lbrack 3]\setminus \{a\}}\hat{v}_{t,j}^{(a,b)}\hat{v}_{s,j}^{(a,b)\prime }$, and $\hat{\mathrm{u }}_{i,i,j}=\frac{1}{2}\sum_{a\in \lbrack 3]}\sum_{b\in \lbrack 3]\setminus \{a\}}\hat{u}_{i,j}^{(a,b)}\hat{u}_{i,j}^{(a,b)\prime }\mathbf{1}_{ia}$. Define
where $S_{j,its}=\hat{e}_{j,it}\hat{e}_{j,is}\hat{\mathrm{v}}_{t,s,j}\left[ \tau -K\left( \frac{\hat{\epsilon}_{it}}{h_{N}}\right) \right] \left[ \tau -K\left( \frac{\hat{\epsilon}_{is}}{h_{N}}\right) \right] .$ We further define
Let $F_{i,ts}(\cdot ,\cdot )$ and $f_{i,ts}(\cdot ,\cdot )$ denote the joint CDF and PDF of $(\epsilon _{it},\epsilon _{is})$ given $\mathscr{D}_{e},$ respectively. To justify the consistency of the variance estimators, we add the following assumption.
Assumption (ref)(i)-(iv) are standard for consistent estimation of the asymptotic variance matrix; see, e.g., chen2019two and galvao2016smoothed. Assumption (ref)(v) imposes the homogeneity moment condition across individuals, and Assumption (ref)(vi) assumes the moments calculated from subsamples are close to those from the full sample given the random splitting. Under Assumption (ref), following the idea of chen2019two, we establish in Lemma (ref) of the online supplement the consistency of $\hat{\Sigma}_{u_{j}}$ and $\hat{ \Sigma}_{v_{j}}$. Similar conclusions hold for the other estimates.
In this section, we consider two specification tests under different rank conditions.
When $K_{j}=1$ for some $j\in [p]$, it is interesting to test whether the matrix $\Theta_{j}^{0}$ is homogeneous across individuals (i.e., row-wise) or across time (i.e., column-wise). For these two cases, we can write factors and factor loadings as
where $u_{j}=\frac{1}{N}\sum_{i=1}^{N}u_{i,j}^{0}$ and $v_{j}=\frac{1}{T} \sum_{t=1}^{T}v_{t,j}^{0}.$ For the homogeneity across individuals, the null and alternative hypotheses can be written as
Similarly, for the homogeneity across time, the null and alternative hypotheses can be written as
Note that we aim to test the two null hypotheses separately. That is, we can test for homogeneous slope across individuals while allowing for heterogeneous slopes across time and vice versa. This is different from the majority of the literature which either tests for slope homogeneity across individuals while assuming the slopes are homogeneous across time or tests for structural breaks across time while assuming the slopes are homogeneous across individuals.
We first consider testing $H_{0}^{I}$. Following the lead of castagnetti2015inference, we define\footnote{ Alternatively, we can also define $S_{u_{j}}^{o}=\max \left( S_{u_{j}}^{(3,2)},S_{u_{j}}^{(2,1)},S_{u_{j}}^{(1,3)}\right) .$ It is easy to show that this statistic shares the same asymptotic null distribution as $ S_{u_{j}}.$ But due to the unknown dependence structure between the two, we cannot take the maximum or the other continuous function of $S_{u_{j}}$ and $ S_{u_{j}}^{o}$ as a new test statistic.}
where $\hat{\bar{u}}_{j}^{(a,b)}=\frac{1}{N_{a}}\sum_{i\in I_{a}}\hat{u} _{i,j}^{(a,b)}$. Similarly, to test for $H_{0}^{II}$, we construct
where
$\hat{\bar{v}}_{j}^{(a,b)}=\frac{1}{T}\sum_{t=1}^{T}\hat{v} _{t,j}^{(a,b)},$ and $\mathsf{b}(n)=\log n-\frac{1}{2}\log \log n-\log \Gamma (\frac{1}{2})$ for $n\in \left\{ N,T,NT\right\} .$
To proceed, we introduce some notation. Recall that $ b_{j,it}^{0}=e_{j,it}v_{t,j}^{0}(\tau -\mathbf{1}\left\{ \epsilon_{it}\leq 0\right\} )$ and $\xi_{j,it}^{0}=e_{j,it}u_{i,j}^{0}\left(\tau -\mathbf{1} \left\{ \epsilon_{it}\leq 0\right\} \right) $. Define
Let $\mathfrak{B}_{j,t}^{(\ell )}=\left(\mathfrak{b}_{j,1t}^{(\ell )\prime},\cdots ,\mathfrak{b}_{j,Nt}^{(\ell )\prime }\right) ^{\prime}$ for $ \ell \in \left[ 4\right] .$ Define
We add the following two assumptions.
Assumption (ref) implies that both $\Sigma _{u_{j}}$ and $ \Sigma_{v_{j}}$ are well behaved. Assumption (ref) imposes that we can approximate high dimensional vectors $\frac{1}{\sqrt{T}}\sum_{t\in[T]} \mathfrak{B}_{j,t}^{(1)}$, $\frac{1}{\sqrt{N_{3}}}\sum_{i\in I_{3}}\mathfrak{ B}_{j,i}^{(2)}$, $\frac{1}{\sqrt{N_{2}}}\sum_{i\in I_{2}}\mathfrak{B} _{j,i}^{(3)}$ and $\frac{1}{\sqrt{N_{1}}}\sum_{i\in I_{1}}\mathfrak{B} _{j,i}^{(4)}$ by four Gaussian vectors. Similar conditions have been imposed in the literature; see, e.g., Assumption SA3 lu2021uniform.
The following theorem reports the asymptotic properties of $S_{u_{j}}$ and $ S_{v_{j}}$ under the respective null and alternative hypotheses.
Remark 5. Theorem (ref) implies that our test statistics follow the Gumbel distributions asymptotically under the null, are consistent under the global alternatives, and have non-trivial power against the local alternatives. The power function of $S_{u_{j}}$ approaches 1 as long as $\frac{T}{\log N}\max_{i\in [N]}\left\Vert c_{i,j}^{u}\right\Vert_{2}^{2}$ diverges to infinity as $\left(N,T\right) \rightarrow \infty .$
When $K_{j}=2$ for some $j\in [p]$, it is interesting to test whether $ \Theta_{j,it}^{0}$ exhibits the additive structure which is widely assumed in a two-way fixed effects model. That is, one may test the following null hypothesis
The alternative hypothesis $H_{1}^{III}$ is the negation of $H_{0}^{III}.$
Let $\bar{\Theta}_{j,i\cdot }=\frac{1}{T}\sum_{t\in \lbrack T]}\Theta _{j,it}^{0}$, $\bar{\Theta}_{j,\cdot t}^{I_{a}}=\frac{1}{N_{a}}\sum_{i\in I_{a}}\Theta _{j,it}^{0}$, and $\bar{\Theta}_{j}^{I_{a}}=\frac{1}{N_{a}T} \sum_{i\in I_{a}}\sum_{t\in \lbrack T]}\Theta _{j,it}^{0}$ for $a\in \lbrack 3]$. Define
Note that $\Theta _{j,it}^{\ast }=0$ $\forall \left( i,t\right) \in \lbrack N]\times \lbrack T]$ under $H_{0}^{III}.$ So we can propose a test for $ H_{0}^{III}$ based on estimates of $\Theta _{j,it}^{\ast }.$ Define
for $a\in \lbrack 3]$. Then, we can define the sample analogue of $\hat{ \Theta}_{j,it}^{\ast }$ as
Its corresponding asymptotic variance can be estimated by $\hat{\Sigma} _{j,it}^{\ast }$ defined as
where $\hat{\bar{u}}_{j}^{(a,b)}=\frac{1}{N_{a}}\sum_{i\in I_{a}}\hat{u} _{i,j}^{(a,b)}$ and $\hat{\bar{v}}_{j}^{(a,b)}=\frac{1}{T}\sum_{t\in \lbrack T]}\hat{v}_{t,j}^{(a,b)}$. Then, the final test statistic is $$S_{NT}=\max_{i\in \lbrack N],t\in \lbrack T]}\left( \hat{\Theta} _{j,it}^{\ast }\right) ^{2}/\hat{\Sigma}_{j,it}^{\ast }.$$
The following theorem studies the asymptotic properties of $S_{NT}$ under the null and alternatives.
Similar remark after Theorem (ref) holds here. In particular, $ S_{NT}$ has the desired asymptotic Gumbel distribution under the null and is consistent under the global alternative.
In this section, we conduct a set of Monte Carlo simulations to show the finite sample performance of our low-rank quantile regression estimates and specification tests.
Below we will consider the following data generating process (DGP):
where $X_{it}=(X_{1,it},X_{2,it})^{\prime }$, $\Theta _{it}=(\Theta _{1,it},\Theta _{2,it})^{\prime }$, $\Theta _{0,it}$ is the intercept term which will be specified via the IFEs.
First, we consider four DGPs where the rank of each slope matrix is 1:
For the case that the rank of the slope matrix is 2, we consider two DGPs which have the additive structure for the slope coefficient of one regressor and the factor structure with two factors for the slope coefficient of another regressor. Specifically,
For $\Theta \in \mathbb{R}^{N\times T}$, define $RMSE(\Theta )=\frac{1}{ \sqrt{NT}}\left\Vert \Theta -\Theta ^{0}\right\Vert _{F}$. Table (ref) shows the RMSEs of the full-sample low rank matrix estimates under different quantiles for each DGP. As Theorem (ref)(i) predicts, the RMSEs decrease as both $N$ and $T\ $increase. Given the fact that $ N\wedge T=T$ in the simulations, the decrease of the RMSEs is largely driven by the increase of $T.$
Table (ref) reports the frequency of correct rank estimation by the singular value thresholding (SVT) approach based on 1000 replications. Note that the true ranks of the intercept and slope matrices in DGPs 1-4 and 5-6 are 1 and 2, respectively. The results show that the SVT can accurately determine the correct rank of the coefficient matrices in all DGPs for all three quantile indices under investigation.
In Section 4, we define $S_{u_{j}}$\ and $S_{v_{j}}$\ as the sup-type test statistics. Table (ref) reports the empirical size and power at the 5% nominal level for the null hypothesis that the slope coefficient is homogeneous across either $i$ or $t$. The results in DGPs 1 and 3 give the empirical size, and those in DGPs 2 and 4 give the empirical power. As the results in Table (ref) indicate, our tests have reasonable size despite the fact that they are slightly conservative like most extreme-value based sup-tests in the literature. In terms of power, out tests have superb power in both DGPs across all three quantile indices.
Table (ref) shows the empirical size and power of our test for DGPs 5 and 6. The findings are similar to those in Table (ref). In particular, our tests are a bit conservative under the null. The empirical power tends to 1 quickly as $T$ increases.
In this section we consider two empirical applications: the heterogeneous investment equation and the heterogeneous quantile effect of foreign direct investment on unemployment.
In this subsection, we revisit the investment equation. fazzari1988financing point out that investment may show sensitivity to movements in cash flow when firms face constraints for external finance. Since fazzari1988financing, there has been a large literature on the effect of cash flow on the corporate investment; see devereux199011, gilchrist1995evidence, kaplan1995financing, cleary1999relationship, rauh2006investment, and almeida2007financial, among others. Using the panel dataset, we consider the scaled version of the investment equation as follows:
where $I$ is the corporate investment, $CF$ is the cash flow, $q$ is the Tobin's q, $K$ is the capital stock and $u$ is the innovation. $\Theta _{0,it}$ refers to the fixed effects (FEs). Rather than the mean estimation, galvao2015efficient estimate the effects of the firm's cash flow and Tobin's q on investment at different quantiles. By using the panel quantile regression with individual FEs, they show that the slope estimates change across $\tau $. However, they do not allow the slope coefficients, $\Theta _{1}$ and $\Theta _{2},$ to change either over $i$ or $ t.$ Inspired by galvao2015efficient, we estimate the following model
where $IK_{it}=\frac{I_{it}}{K_{i,t-1}}$, and $CFK_{it}=\frac{CF_{it}}{ K_{i,t-1}}$. Here we don't restrict the specific structure on the FEs and they can be either additive or interactive.
The data are taken from the China Stock Market & Accounting Research (CSMAR) Database. We use quarterly data for 195 manufacturing firms in China from 2003 to 2020. Based on the model ((ref)), we define corporate investment as $I_{it}=LI_{it}-LI_{i,t-1}$, where $LI_{it}$ is the total value of long-term corporate investment as the sum of long-term equity investment, long-term bound investment, fixed assets and immaterial assets. The investment measures the change of firm's total investment compared to the last period. All these four variables can be easily obtained from the balance sheet. We directly use Tobin's q from the CSMAR database, where by definition $q=\frac{MV}{K}$ and $MV$ is the market value of the firm. We obtain a balanced panel dataset with 195 firms and 72 time periods. The units of corporate investment, capital and cash flow are measured by billions of Chinese RMB.
By using the SVT approach, we obtain the estimates of the ranks of $\Theta _{1}$ and $\Theta _{2}:$ $\hat{r}_{1}=\hat{r}_{2}=1$ for each $\tau =\{0.25,0.5,0.75\}$. Consequently, we can consider the test that whether $ \Theta _{j,it}$ is constant over $i$ or constant over $t$ for both $j=1,2$. Specifically, we want to test whether the effect of cash flow and Tobin's q on the firm's investment is homogeneous over $i$ or across $t$ with market imperfection. That is, for $j\in \{1,2\},$ we shall test
Figure (ref) shows the estimation results for the factor and factor loadings of two slope coefficient matrices under different quantiles. In each sub-figure, the first and second rows report the results for $\Theta _{1}$ and $\Theta _{2},$ respectively. Specifically, the first row of Figure (ref)(a) gives the plot of $\left\{ \hat{u}_{i,1}\right\} _{i\in \lbrack N]}$ as a catenation of $\{\hat{u}_{i,1}^{(1,2)},\hat{u} _{i,1}^{(2,3)},\hat{u}_{i,1}^{(3,1)}\}$ at the left and as a cantenation of $ \{\hat{u}_{i,1}^{(1,3)},\hat{u}_{i,1}^{(2,1)},\hat{u}_{i,1}^{(3,2)}\}$ at the right in the first row, and similarly the plot of $\left\{ \hat{u} _{i,2}\right\} _{i\in \lbrack N]}\ $in the second row. Similarly, the first row of Figure (ref)(d) shows $\{\hat{v}_{t,1}^{(a,b)}\}_{t\in \lbrack T]}$ for $a\in \lbrack 3],$ $b\in \lbrack 3]\setminus \{a\}$ in the first row and $\{\hat{v}_{t,3}^{(a,b)}\}_{t\in \lbrack T]}$ for $a\in \lbrack 3],$ $b\in \lbrack 3]\setminus \{a\}$ in the second row.
Table (ref) reports the test statistics, critical values, and $p$-values. Tobin's q can measure a firm's investment demand. After controlling the Tobin's q and the intercept FEs, the coefficient of cash flow captures a firm's potential for external investment with the variation of internal finance. It is clear that we can reject the homogeneous hypotheses for both $ i$ and $t$ at the 1% significance level for each $\tau \in \{0.25,0.5,0.75\} $. This indicates that with high probability, the slope coefficient of both $CFK$ and Tobin's q follow the factor structure with one factor.
The above study shows strong evidence that under imperfect market, the sensitivity of corporate investment to cash flow exhibits both individual heterogeneity and time heterogeneity across quantiles. It implies that neither the usual homogenous panel QR model nor the panel QR model with either cross-section or time heterogeneity alone in the slope coefficients fails to fully capture the unobserved heterogeneity in the investment equation.
Investment is one of the major driving forces for economic growth and employment. Among the investment, foreign direct investment (FDI) is an important contributor to the employment. See craigwell2006foreign, aktar2009can, karlsson2009foreign, mucuk2013effect, and strat2015fdi, among others. Controversially, mucuk2013effect argue that FDI may have both positive and negative effects on employment. On the one hand, FDI adds to the net capital and creates jobs through forward and backward linkages and multiplier effects in local economy. On the other hand, acquisitions may rely on imports or displacement of existing firms which may result in job loss.
To study the relationship of FDI, economic growth rate and unemployment at the country level, we consider the following panel quantile regression model,
where $U_{it}$ is the unemployment rate of country $i$ at year $t$, $ G_{i,t-1}$ is the economic growth measured by the growth of real GDP. $ \Theta _{0,it}$ is the FEs of country $i$ and year $t$, $\Theta _{1,it}$ is the elasticity of the economic growth in the previous year to the unemployment this year, and $\Theta _{2,it}$ is the elasticity of FDI to the unemployment.
We draw the data for 126 countries from 1992-2019. The data for the unemployment rate are taken from International Labor Organization (ILO) and GDP growth and FDI are from the World Bank Development Indicators (WDI) historical database. The rank estimation procedure shows that $\hat{r}_{1}=2$ and $\hat{r}_{2}=1$. Consequently, we can test whether the elasticity of FDI to the unemployment rate is homogeneous across individual countries and over years 1992-2009, and whether the elasticity of growth rate to unemployment follows the additive structure, i.e.,
Table (ref) reports the test results under quantiles 0.25, 0.5 and 0.75 for the above three null hypotheses. Figure (ref) gives the estimation results for the factor and factor loading estimates of the slope coefficient $\Theta _{2}$. As Table (ref) suggests, we can reject all the above three null hypotheses safely at the conventional 5% significance level. This means that the effect of FDI on the unemployment rate is different across both countries and time even though the estimated rank of $\Theta _{2}$ is one, and the effect of economic growth rate on the unemployment is heterogeneous across both countries and time and it does not exhibit an additive structure.
This paper considers panel QR model with heterogeneous slopes over both $i$ and $t$. Compared to chernozhukov2019inference, to remove the bias from the nuclear norm regularization, we split the full sample into three subsamples. We then use the first subsample to compute initial estimators via NNR, the second sample to refine the convergence rate of the initial estimator, and the last subsample to debias the refined estimator. Our asymptotic theory shows that the factor estimates, factor loading estimates and the slope estimates all follow the normal distributions asymptotically. By constructing the consistent estimator for the asymptotic variance, we also conduct two specification tests: (1) the slope coefficient is constant over time or individuals under the case that true rank of slope matrix equals one and (2) the slope coefficient exhibits the additive structure under the case that true rank of the slope coefficient matrix equals two. Our test statistics are shown to follow the Gumbel distribution asymptotically under the null, consistent under the global alternative and have non-trivial power against local alternatives. Monte Carlo simulation and empirical studies illustrate the finite sample performance of our algorithm and test statistics.