Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
62,596 characters · 11 sections · 81 citation commands
Kernel Estimation for Panel Data with Heterogeneous Dynamics
The characteristics of heterogeneity across economic units are informative for many econometric applications. For example, there is an interest in heterogeneity in the dynamics of price deviations or changes (e.g., klenow2010microeconomic; CruciniShintaniTsuruga15). As another example, allowing for the presence of heterogeneity may make a crucial difference in identification and estimation of production functions (e.g., AckerbergBenkardBerryPakes07; Kasahara2017). Thus, there are many econometric studies that investigate the degree of heterogeneity using panel data (e.g., HsiaoPesaranTahmiscioglu99; FernandezValLee13; JochmansWeidner2019; Okui2017).
This paper proposes kernel-smoothing estimation for panel data to analyze heterogeneity across cross-sectional units.\footnote{An { R} package to implement the proposed procedure is available from the authors' websites. } After estimating the mean, autocovariances, and autocorrelations of each unit, we compute the kernel densities based on these estimated quantities. This easy-to-implement procedure provides useful visual information for heterogeneity in a model-free manner. For example, the densities of the heterogeneous mean, variance, and first-order autocorrelation of the price deviations indicate visually the characteristics of heterogeneity in the long-run level, variance, and persistence of the price deviations across items (goods and services) that are cross-sectional units in this example. Indeed, several empirical studies have used such estimation for various applications (e.g., Kasahara2017 and roca2017learning), but there is no theoretical foundation for kernel-smoothing to examine heterogeneity in long-panel data.
We show consistency and asymptotic normality of the kernel density estimator based on double asymptotics under which both the cross-sectional size $N$ and the time-series length $T$ tend to infinity with the bandwidth $h$ shrinking to zero (denoted by $N, T \to \infty$ and $h \to 0$).\footnote{More precisely, the double asymptotics $N, T \to \infty$ are any monotonic sequence $T=T(N) \to \infty$ as $N \to \infty$ and the bandwidth $h \to 0$ is any monotonic sequence $h = h(N, T(N)) = h(N) \to 0$ as $N, T \to \infty$ in our setting. Note that each theoretical result in this paper specifies additional conditions on the relative magnitudes of $N$, $T$, and $h$. } The asymptotic properties exhibit several unique features that have not been well examined in the long-panel literature. Most importantly, asymptotic bias of even very high order affects the conditions on the relative magnitudes of $N$, $T$, and $h$ required for consistency and asymptotic normality. As a result, the different orders of asymptotic expansion we can execute have different relative magnitude conditions. This unique feature contrasts our analysis with the existing analyses without kernel smoothing where the required relative magnitude conditions do not depend on the order of expansions (e.g., $N / T^2 \to 0$ for asymptotic normality in HsiaoPesaranTahmiscioglu99 and Okui2017). The weakest condition (i.e., how small $T$ can be compared with $N$) can be obtained by executing an infinite order expansion. Even in that case, the required condition is stronger than those typically observed in the literature without kernel smoothing. Moreover, it requires nontrivial discussions for the expansion (e.g., the summability of the infinite-order series). We clarify that these unique features are caused by the presence of the bandwidth $h$ and by using the estimated quantities.
Based on an infinite-order expansion, we show three asymptotic biases for the density estimation. The first is the standard kernel-smoothing bias of order $O(h^2)$ (see, e.g., LiRacine2007). The second is caused by the incidental parameter problem (NeymanScott48 and Nickell1981) and is $O(1/T)$. The third results from the nonlinearity of the kernel function and the difference between the estimated quantity and the true quantity. We show that this is $O(1/(Th^2)) + \sum_{j=3}^{\infty} O(1/\sqrt{T^j h^{2j}})$, which is obtained only if we execute an infinite-order expansion. By showing these asymptotic biases, we prove that the relative magnitude conditions for consistency and asymptotic normality are $N^2 / T^5 \to 0$ and $N^2 / T^3 \to 0$, respectively, when using the standard bandwidth $h \asymp N^{-1/5}$ in the density estimation with second-order kernels.
We propose to apply a split-panel jackknife method in DhaeneJochmans15 to reduce these biases. In particular, we formally show that the half-panel jackknife (HPJ) corrects the incidental parameter bias and the second-order nonlinearity bias without inflating the asymptotic variance. While the jackknife is useful in bias reduction especially when $T$ is small, we also show that it does not weaken the relative magnitude conditions for consistency and asymptotic normality.
We also develop confidence interval (CI) estimation and selection of bandwidth. To construct CI, we extend the robust bias-corrected (RBC) procedure in calonico2018effect to split-panel jackknife bias-corrected estimation. This method explicitly corrects all three biases above. For the bandwidth selection, we can apply any standard procedures in the literature. This is because, under the relative magnitude conditions, the asymptotic mean squared error (AMSE) and asymptotic distribution of the split-panel jackknife bias-corrected estimator are the same as those of the infeasible estimator based on the true quantity.
We also examine the properties of the cumulative distribution function (CDF) estimator constructed by integrating the kernel density estimator. This kernel CDF estimator also exhibits asymptotic bias that varies in the order of the asymptotic expansion that we execute. We also derive the closed form formula for the asymptotic bias. This is an interesting result from theoretical viewpoint because the formula for asymptotic bias for the empirical distribution is available only for Gaussian errors JochmansWeidner2019 and has not been derived in general form Okui2017. However, the required conditions on $N$ and $T$ for the kernel CDF estimation turn out to be stronger than those for the empirical CDF estimation derived in those studies.
We illustrate our procedures by an empirical application on heterogeneity of price deviations from the law of one price (LOP). Our procedures reveal significant heterogeneity in the price deviations dynamics. The split-panel jackknife bias-corrected density estimates imply much more volatile and persistent dynamics than the estimates without bias correction and the difference is visually noticeable. This result highlights the importance of the bias correction in that the bias-corrected densities can provide distinct visual information for heterogeneity from the densities without bias correction.
\paragraph{Related literature.} Our setting and motivation closely relate to Okui2017, but there are several important distinctions in both theoretical and practical aspects. First, our relative magnitude conditions are different from Okui2017 in which second-order expansions suffice to derive the conditions on estimating the moments of the quantities (e.g., the variance of the heterogeneous mean). This feature in particular contrasts the theoretical contributions in both papers, and indeed our relative magnitude conditions are new in the literature. Second, we show the new insight that the split-panel jackknife is applicable even to kernel estimation. Third, because it is well known that bootstrap inferences do not capture kernel-smoothing bias (see, e.g., hall2013simple), we extend the RBC inference in calonico2018effect instead of the cross-sectional bootstrap in Okui2017.\footnote{The failure of the cross-sectional bootstrap inference in our kernel estimation is formally shown in the previous version of this study uploaded to arXiv (arXiv:1802.08825v2).} Finally, while Okui2017 do not clarify asymptotic biases for their empirical CDFs, we formalize those of our kernel estimators.
Our CDF estimation relates to JochmansWeidner2019 who derive the bias of the empirical distribution based on noisy measurements (e.g., estimated quantities) for the true variables of interest. Their results are complementary to ours. They consider a situation where observations exhibit Gaussian errors. We do not assume such errors. The kernel smoothing allows us to derive bias under much weaker distributional assumptions at the price of creating additional higher-order biases.
Many econometric studies examine heterogeneity in panel data (e.g., PesaranSmith95; HsiaoPesaranTahmiscioglu99; PesaranShinSmith99; FernandezValLee13). Among them, HorowitzMarkatou96, ArellanoBonhomme12, and MavroeidisSasakiWelch15 propose to estimate the densities of heterogeneous quantities with short-panel data based on deconvolution techniques under some model specifications. Compared with them, we propose model-free kernel-smoothing estimation with long-panel data.
Several studies propose model-free analyses for panel data, but do not focus on the degree of heterogeneity in the dynamics. For example, Okui08, Okui11, Okui14 and LeeOkuiShintani13 consider homogeneous dynamics, and GalvaoKato14 study the properties of the possibly misspecified fixed effects estimator in the presence of heterogeneous dynamics.
Kernel density estimation using estimated quantities is also examined in the literature on structural estimation of auction models. For example, mamarmershneyerov2019 and gpv2000 estimate the density of individual evaluations of auctioned goods. In their first stage, individual evaluations of auctioned goods are estimated nonparametrically and their second stage is the kernel density estimation applied to estimated evaluations. They also observe that the estimation errors from the first stage affect the asymptotic behavior of the second stage estimator in a nonstandard way. However, their problems are different from ours. Their main issue is the cross-sectional correlation caused by the use of the same set of observations to estimate individual evaluations. As a result, their estimation errors affect the precision and the convergence rate of the second stage estimator. In our case, estimation errors in the first stage are cross-sectionally independent and affect the bias but not the (first-order) variance of the second stage estimator.
\paragraph{Paper organization.} Section (ref) introduces our setting and density estimation. Section (ref) develops the asymptotic theory, bias correction, CI estimation, bandwidth selection, and CDF estimation. Section (ref) presents the application. Section (ref) concludes. The supplementary appendix contains the proofs of the theorems, technical lemmas, other technical discussions, and Monte Carlo simulations.
This section describes the setting and the proposed estimation. We explain our setting and motivation in a succinct manner because they are similar to those in Okui2017.\footnote{Several remarks and possible extensions can be found in the previous version of this study and Okui2017. For example, we can consider the presence of covariates and time effects and estimation based on other heterogeneous quantities, such as random coefficients in linear models, with minor modifications. In this paper, we explain our estimation briefly to save space.}
We observe panel data $\{ \{ y_{it} \}_{t=1}^T \}_{i=1}^N$ where $y_{it}$ is a scalar random variable. We assume that $y_{it}$ is strictly stationary across time and that each individual time series $\{y_{it}\}_{t=1}^T$ is generated from some unknown probability distribution $\mathcal{L}(\{y_{it}\}_{t=1}^T; \alpha_i)$, where $\alpha_i$ is a (possibly infinite dimensional) random variable specifying the dynamics of $y_{it}$. We note that $\alpha_i$ is an abstract parameter and it does not appear in the actual implementations of our proposed procedure. Characterizing heterogenous dynamics using this abstract parameter is mathematically convenient because it allows us to keep an i.i.d. assumption. Existing studies without model specifications also employ this approach (e.g., GalvaoKato14). We denote the conditional expectation given $\alpha_i$ by $E(\cdot | i)$.
Our goal is to examine the degree of heterogeneity of the dynamics of $y_{it}$ across units in a model-free manner. To this end, we focus on estimating the density of the mean $\mu_i \coloneqq E(y_{it} | i)$, $k$-th autocovariance $\gamma_{k,i} \coloneqq E( (y_{it} - \mu_i)(y_{i,t-k} - \mu_i) | i )$, and $k$-th autocorrelation $\rho_{k,i} \coloneqq \gamma_{k, i} / \gamma_{0,i}$. We first estimate $\mu_i$, $\gamma_{k,i}$, and $\rho_{k,i}$ by the sample analogues: $\hat \mu_i \coloneqq \bar y_i \coloneqq T^{-1} \sum_{t=1}^{T} y_{it}$, $\hat \gamma_{k,i} \coloneqq (T - k)^{-1} \sum_{t=k+1}^T (y_{it} - \bar y_i) (y_{i, t-k} - \bar y_i)$, and $\hat \rho_{k,i} \coloneqq \hat \gamma_{k,i} / \hat \gamma_{0,i}$. Throughout the paper, we use the notation $\xi_i$ to represent one of $\mu_i$, $\gamma_{k,i}$, or $\rho_{k,i}$ and the notation $\hat \xi_i$ for the corresponding estimator. The kernel estimator for the density $f_{\xi}(x)$ is given by:
where $x \in \mathbb{R}$ is a fixed point, $K:\mathbb{R} \to \mathbb{R}$ is a kernel function, and $h > 0$ is a bandwidth satisfying $h \to 0$.\footnote{We can consider estimating the joint density for $\mu_i$, $\gamma_{k, i}$, and $\rho_{k,i}$ in the same manner.} This is a standard estimator except that we replace the true $\xi_i$ with the estimated $\hat \xi_i$.
This section develops our asymptotic theory, CI estimation, bandwidth selection, and CDF estimation based on the density estimator $\hat f_{\hat \xi}(x)$. We define the notations $w_{it} \coloneqq y_{it} - \mu_i = y_{it} - E(y_{it}|i)$ and $\bar w_i \coloneqq T^{-1} \sum_{t=1}^T w_{it}$. By construction, $y_{it} = \mu_i + w_{it}$. Note that $\hat \mu_i = \bar y_i = \mu_i + \bar w_i$, $E(w_{it}|i) = 0$, and $\gamma_{k,i} = E(w_{it} w_{i,t-k}|i)$.
Before formally showing the asymptotic properties, we explore the unique features of our asymptotic investigations in an informal manner. By doing so, we clarify the mechanism behind the observation that even very high orders of asymptotic bias matter for our asymptotic analysis.
We here focus on the density estimator for $\hat \mu_i$, but similar discussions are also relevant for $\hat \gamma_{k,i}$ and $ \hat \rho_{k,i}$. Noting that $\hat \mu_i - \mu_i = \bar w_i$, we examine the $J$-th order Taylor expansion of $\hat f_{\hat \mu}(x)$:
where $K^{(j)}$ denotes the $j$-th order derivative and $\tilde \mu_i$ is between $\mu_i$ and $\hat \mu_i$.
The first term in (ref) is the infeasible density estimator based on the true $\mu_i$, and its asymptotic behavior is standard and well known in the kernel-smoothing literature. It converges in probability to the density of interest $f_{\mu}(x)$ as $N\to \infty$ and $h\to 0$ with $Nh \to \infty$. In addition, when $Nh^5 \to C \in [0, \infty)$ also holds, it can hold that:
where $\kappa_1 \coloneqq \int s^2 K(s) ds$ and $\kappa_2 \coloneqq \int K^2(s) ds$ and $\mathcal{N}(\mu, \sigma^2)$ is a normal distribution with mean $\mu$ and variance $\sigma^2$.
The unique features in our situation are caused from the second and third terms in (ref). For the second term, under regularity conditions, the mean can be evaluated as:
where we used $E( (\bar w_i)^j | \mu_i = x) = O(T^{-j/2})$ (see Assumption (ref) below and Lemma (ref) in the supplement). Noting that $E(\bar w_i| \mu_i) = 0$, the bias caused from the second term in (ref) can be written as $\sum_{j = 2}^{J-1} O(1/\sqrt{T^j h^{2j}})$. This bias is negligible when $1/(Th^2) \to 0$, which is identical to the relative magnitude condition $N^2/T^5 \to 0$ when using the standard bandwidth $h \asymp N^{-1/5}$ in the density estimation with second-order kernels. For the third term in (ref), the absolute mean can be evaluated as:
where $0 < M < \infty$ denotes a generic positive constant and we use $E|\bar w_i|^J = O(T^{-J/2})$ (see Lemma (ref)). Hence, the third term in (ref) is $O_p(1/\sqrt{T^J h^{2J+2}})$ by Markov's inequality. Remarkably, this term does not vanish even when $1/(Th^2) \to 0$ under which the lower-order terms are negligible. This term can be negligible only if $1/(Th^{2+2/J}) \to 0$, which implies $N^{(2 + 2/J)} / T^5 \to 0$ when $h \asymp N^{-1/5}$. Note that $N^{(2 + 2/J)} / T^5 \to 0$ is “stronger” than $N^2 / T^5$.
The asymptotic investigation above exhibits several unique features. First, it implies that the relative magnitude condition for consistency (and also that for asymptotic normality) varies in the order of the expansion. Specifically, we need $1/(Th^{2+2/J}) \to 0$ to achieve the consistency of $\hat f_{\hat \mu}(x)$ based on the $J$-th order expansion. Second, we can obtain the “weakest” relative magnitude condition $1/(Th^2) \to 0$ for consistency, only if we execute the infinite-order expansion (that is, as $J \to \infty$). Finally, while we can derive the suitable condition $1/(Th^2) \to 0$ via the infinite-order expansion, it requires the existence of higher-order moments of $w_{it}$. The evaluation based on the infinite-order expansion demands the existence of $E|w_{it}|^j$ for any $j$. Hence, there is a trade-off between the relative magnitude condition and the existence of higher-order moments.
Asymptotic normality requires a further stronger condition. Because the rate of convergence of the kernel estimator is $\sqrt{Nh}$, it requires $N/(T^J h^{2J+1})\to 0$. This condition is at best $N^2/T^3 \to 0$, which is obtained under an infinite order expansion with standard bandwidth ($h \asymp N^{-1/5}$). Note that, as in the density estimation above, the highest order of the expansion determines the required condition for asymptotic normality. Such a very high order of bias cannot be corrected in practice, even though methods to correct the first few orders of bias are available in the long-panel literature (e.g., DhaeneJochmans15). This result is in stark contrast to the existing studies in which bias correction improves the conditions on the relative magnitudes of $N$ and $T$.
The main reason behind these unique features is that the curvature of the summand (i.e., $K((x - \hat \mu_i)/h)$) depends on the bandwidth $h$. Roughly speaking, as $h \to 0$, the summand function becomes steeper and more “nonlinear.” It exacerbates the bias caused by the nonlinearity and it turns out that even a very high order derivative of $K$ affects the bias. Alternatively, we may also interpret this problem based on the equation $K((x -\hat \mu_i)/h) = K((x - \mu_i)/h + (\mu_i -\hat \mu_i)/h))$. The contribution of the error by using the estimated $\hat \mu_i$ is $(\mu_i - \hat \mu_i)/h$ and it increases as $h\to 0$. Hence, the bias of the density estimator heavily depends on the magnitude of $h$ and the nonlinearity of $K$.
We here formally show the presence of asymptotic biases of the kernel density estimator in (ref). We conduct asymptotic investigations based on an infinite-order expansion under which the weakest possible condition on the relative magnitude of $N$ and $T$ is obtained.
We assume the following basic conditions for the data-generating process. These are essentially the same as the assumptions in Okui2017.
Assumptions (ref) and (ref) require that the individual time series given $\alpha_i$ is strictly stationary across time but i.i.d. across units. The identical distribution across $i$ is essential for our analysis. The independence assumption across $i$ makes our asymptotic investigations tractable, while the consistency result and the same asymptotic biases could be derived even under weak cross-sectional dependence. Note that the i.i.d. assumption does not exclude the presence of heterogeneity in panel data. In our setting, heterogeneity is caused by differences in the realized values of $\{\alpha_i\}_{i=1}^N$ across units. Assumption (ref) also restricts the degree of persistence of the individual time series. The conditions for stationarity and degree of persistence require that the times series for each unit is not a unit root process and that the initial value of each time series is generated from a stationary distribution. Assumption (ref) requires the existence of the moments of $w_{it}$, and it allows us to derive the asymptotic biases of the estimators. While we can develop the theoretical properties of the estimators in situations where Assumptions (ref) and (ref) do not hold for some numbers $r_m$ and $r_d$, we cannot derive the higher-order biases based on infinite-order expansions in such situations. As a result, in such situations, we demand stronger conditions on the relative magnitudes as discussed in the previous section. Assumption (ref) allows us to derive the asymptotic properties of the kernel estimators for $\rho_{k,i}$. All of the assumptions can be satisfied in popular panel data models. For example, they all hold when $y_{it}$ follows a heterogeneous stationary panel autoregressive moving--average model with a Gaussian error term (e.g., $y_{it} = c_i + \phi_i y_{i,t-1} + u_{it} + \theta_i u_{i,t-1}$ with $u_{it} \sim \mathcal{N}(0, \sigma^2)$).
We also assume the following additional conditions.
Assumption (ref) includes the standard conditions for the kernel function, except for infinite differentiability. We require the differentiability in order to expand the kernel estimator for the estimated $\hat \xi_i$ at the true $\xi_i$ based on the infinite-order expansion. Note that the symmetry of $K$ implies that $\int K^{(j)}(s) d s = 0$ for any odd $j$.
Assumption (ref) requires that $\xi_i$ is continuously distributed without probability mass. The continuity of the random variable is essential for implementing kernel-smoothing estimation as it rules out situations where there is no heterogeneity for $\xi_i$ (that is, the situation where $\xi_i = \xi$ for any $i$ with some constant $\xi$) and where there is finitely grouped heterogeneity (that is, $\xi_{i_1} = \xi_{i_2}$ for any $i_1, i_2 \in \mathbb{I}_g$ with some sets $\mathbb{I}_1, \mathbb{I}_2, \dots, \mathbb{I}_G$ satisfying $ \bigoplus_{g=1}^G \mathbb{I}_g = \{1, 2, \dots, N\}$).
Assumption (ref) states the existence and smoothness of the conditional expectations. This assumption allows us to derive the exact forms of the asymptotic biases. The convergence rates of the terms are standard and guaranteed by Lemmas (ref) and (ref) in Appendix (ref). For example, the assumption requires that $T \cdot E((\bar w_i)^2 | \mu_i = \cdot) = O(1)$ and the convergence rate is consistent with the result in Lemma (ref).
The following theorem shows that the kernel density estimators are consistent and asymptotically normal but exhibit asymptotic biases. While the theorem assumes an infinite-order Taylor expansion and the summability of the infinite series of the asymptotic biases directly, we can show their validity under unrestrictive regularity conditions. Because these discussions are highly technical and demand lengthy explanations, they appear in Appendices (ref) and (ref).
The density estimator can be written as the sum of the infeasible estimator based on the true $\xi_i$, say $\hat f_{\xi}(x) \coloneqq (Nh)^{-1} \sum_{i=1}^N K((x - \xi_i)/h)$, and the asymptotic biases. The convergence rate of the estimator is the standard order of $O_p(1/ \sqrt{Nh})$, and the asymptotic distribution is the same as that of the infeasible estimator $\hat f_{\xi}(x)$. However, the feasible estimator exhibits asymptotic biases. These results also require the relative magnitude conditions of $N$, $T$, and $h$; that is, $1/(Th^2) \to 0$ and $N / (T^3 h^5) \to 0$ for consistency and asymptotic normality, respectively.
The density estimator for $\mu_i$ has two main asymptotic biases given $A_{\mu, 1}(x) = 0$, but the density estimators for $\gamma_{k,i}$ and $\rho_{k,i}$ have three main asymptotic biases, in addition to the higher-order biases. The first bias of the form $h^2 \kappa_1 f_\xi''(x)/2$ is the standard kernel-smoothing bias. The second bias of the form $A_{\xi, 1}(x) / T$ is the incidental parameter bias caused from estimating $\gamma_{k,i}$ and $\rho_{k,i}$ by $\hat \gamma_{k,i}$ and $\hat \rho_{k,i}$, respectively. The estimation of $\hat \gamma_{k,i}$ and $\hat \rho_{k,i}$ involves estimating $\mu_i$ by $ \hat \mu_i = \bar y_i$ for each $i$, which becomes a source of the incidental parameter bias. The third bias of the form $A_{\xi, 2}(x) / (Th^2)$ is the second-order nonlinearity bias caused by expanding $K((x - \hat \xi_i )/h)$ for $K(( x - \xi_i )/h)$ by Taylor expansion. Moreover, the $j$-th order nonlinearity bias exhibits the form $A_{\xi, j}(x) / \sqrt{T^j h^{2j}}$ for $j \ge 3$.
We need the two conditions, $1/(Th^2) \to 0$ and $N / (T^3 h^5) \to 0$, to ensure the asymptotic negligibility of the higher-order nonlinearity biases. If we use the standard bandwidth $h \asymp N^{-1/5}$ with second-order kernels, the conditions $1 / (T h^2) \to 0$ and $N / (T^3 h^5) \to 0$ imply that $N^2 / T^5 \to 0$ and that $N^2 / T^3 \to 0$, respectively, which are integrated to $N^2 / T^3 \to 0$. Note that while the incidental parameter bias and the second-order nonlinearity bias are also asymptotically negligible under these conditions, the practical magnitudes of these biases would be larger than those of the higher-order nonlinearity biases.
We have already discussed the source of the nonlinearity bias in Section (ref) so here we provide a slightly more detailed discussion of the incidental parameter bias. It does not appear in $\hat f_{\hat \mu}(x)$ because the estimation error in $\hat \mu_i$ (that is, $\bar w_i$) has zero mean. However, errors in $\gamma_{k,i}$ and $\rho_{k,i}$ are not mean-zero. For example, $\hat \gamma_{0,i} = \sum_{t=1}^T (y_{it} - \bar y_i)^2 /T = \gamma_{0,i} + \sum_{t=1}^T (w_{it}^2 - \gamma_{0,i}) /T - (\bar w_i)^2$ and $(\bar w_i)^2$ is not mean-zero although it converges to zero at the rate $1/T$. This is the source of the incidental parameter bias $A_{\gamma_0,1}/T$ and the order $1/T$ comes from the fact that $(\bar w_i)^2$ is $O_p(1/T)$.
As the incidental parameter bias and the nonlinearity biases in $\hat f_{\hat \xi}(x)$ may be severe in practice, we propose adoption of the split-panel jackknife to correct them. Among split-panel jackknifes, here we consider half-panel jackknife (HPJ) bias correction. For simplicity, suppose that $T$ is even.\footnote{The bias correction with odd $T$ is similar. See DhaeneJochmans15 for details. } For $\xi_i = \mu_i$, $\gamma_{k,i}$, or $\rho_{k,i}$, we obtain the estimators $\hat f_{\hat \xi,(1)}(x)$ and $\hat f_{\hat \xi,(2)}(x)$ of $f_\xi(x)$ based on two half-panel data $\{\{y_{it}\}_{t=1}^{T/2}\}_{i=1}^N$ and $\{\{y_{it}\}_{t=T/2+1}^T \}_{i=1}^N$, respectively. The HPJ bias-corrected estimator is $\hat f_{\hat \xi}^H(x) \coloneqq \hat f_{\hat \xi}(x) - (\bar f_{\hat \xi}(x) - \hat f_{\hat \xi}(x) )$ where $\bar f_{\hat \xi}(x) \coloneqq [\hat f_{\hat \xi,(1)}(x)+ \hat f_{\hat \xi,(2)}(x)]/2$. The term $\bar f_{\hat \xi}(x) - \hat f_{\hat \xi}(x)$ estimates the bias in the original estimator $\hat f_{\hat \xi}(x)$. Importantly, the bandwidths for computing $\hat f_{\hat \xi,(1)}(x)$ and $\hat f_{\hat \xi,(2)}(x)$ must be the same as that for the original estimator $\hat f_{\hat \xi}(x)$ to reduce the biases.
The next theorem formally shows that the HPJ bias-corrected estimator $\hat f_{\hat \xi}^H(x)$ does not suffer from incidental parameter bias and second-order bias, and does not alter the asymptotic variance of the estimator.
Note that HPJ bias correction does not weaken the relative magnitude condition of $N$, $T$, and $h$ for asymptotic normality in Theorem (ref); that is, $N / (T^3 h^5) \to 0$. This is because HPJ bias correction cannot eliminate higher-order nonlinearity biases. This result is in stark contrast to the existing literature where bias correction typically weakens the condition on the relative magnitudes of $N$ and $T$ DhaeneJochmans15.
This section considers CI estimation and the selection of optimal bandwidth for density estimation.
\paragraph{CI estimation.} We propose to apply the RBC procedure in calonico2018effect for CI estimation. It allows us to construct a valid $1 - \alpha$ CI of $f_{\xi}(x)$ while correcting the kernel-smoothing bias $\mathcal{B}_{\xi}(x) \coloneqq h^2 \kappa_1 f_{\xi}^{''}(x) / 2$.
The RBC procedure based on the naive estimator $\hat f_{\hat \xi}(x)$ is almost the same as the original procedure in calonico2018effect. We first note that the kernel-smoothing bias can be estimated by $\hat{\mathcal{B}}_{\hat \xi}(x) \coloneqq h^2 \kappa_1 \hat f_{\hat \xi}''(x)$ where $\hat f_{\hat \xi}''(x) \coloneqq (Nb^3)^{-1} \sum_{i=1}^{N} L''((x - \hat \xi_i)/b)$ with a kernel function $L$ and bandwidth $b \to 0$. Then, the estimator that corrects the kernel-smoothing bias is:
where $\lambda \coloneqq h / b$. Choosing $b$ such that $\lambda \to c$ for some $c \in (0, \infty)$ enables us to capture variance inflation caused by the bias correction while successfully removing the kernel-smoothing bias. In practice, one can set $\lambda = 1$ by following the suggestion in calonico2018effect. The RBC $t$ statistic is given by:
where $\hat \sigma_{RBC}^2(x)$ is the estimator of the nonasymptotic variance of $\hat f_{\hat \xi}(x) - \hat B_{\hat \xi}(x)$:
It holds that $T_{RBC}(x) \stackrel{d}{\longrightarrow} \mathcal{N}(0, 1)$ under similar conditions in Theorem (ref), so that we can construct the $1 - \alpha$ CI of $f_{\xi}(x)$ in the usual manner.
The RBC procedure based on the split-panel jackknife bias-corrected estimator demands some modifications. To see this, the HPJ bias-corrected estimator that also reduces the kernel-smoothing bias can be written as follows:
where $\hat \xi_i^{(1)}$ and $\hat \xi_i^{(2)}$ are the estimators based on the half-series $\{y_{it}\}_{t=1}^{T/2}$ and $\{y_{it}\}_{t=T/2+1}^T$, respectively. Then, the nonasymptotic variance of $\hat f_{\hat \xi}^H(x) - \hat{\mathcal{B}}_{\hat \xi}(x)$ can be estimated by:
As a result, the RBC $t$ statistic based on the HPJ estimator is:
Note that $\hat \sigma_{RBC}^H(x)$ is different from $\hat \sigma_{RBC}(x)$ above because the former also captures the finite-sample variability of HPJ bias correction. We can construct the $1 - \alpha$ CI of $f_{\xi}(x)$ based on $T_{RBC}^H(x)$ in the usual manner. We can also consider similar RBC procedures based on higher-order split-panel jackknife bias correction.
\paragraph{Bandwidth selection.}
We can select the bandwidth $h$ for the density estimation using any standard procedures based on the estimated $\hat \xi_i$. This is because Theorem (ref) shows that the AMSE and asymptotic distribution of the HPJ bias-corrected estimator $\hat f_{\hat \xi}^H(x)$ are identical to those of the infeasible estimator $\hat f_{\xi}(x) = (Nh)^{-1} \sum_{i=1}^N K((x - \xi_i)/h)$. In our application and Monte Carlo simulations, we apply the coverage error optimal bandwidth selection procedure in calonico2018effect because of its desirable properties as shown in the paper. Furthermore, their bandwidth tends to be larger than the bandwidth that minimizes AMSE and would be more suitable in our context because a larger bandwidth makes the nonlinearity biases smaller. Our Monte Carlo simulations also confirm the appropriate finite-sample properties of the procedure.
In this section, we consider the smoothed CDF estimator and derive its asymptotic biases. The CDF $F_{\xi}(x) \coloneqq \Pr(\xi_i \le x)$ can be estimated by integrating the kernel density estimator: $\hat F_{\hat \xi}(x) = \int^{x}_{-\infty} \hat f_{\hat \xi}(v) dv$. It is convenient to write this kernel CDF estimator as:
where $x \in \mathbb{R}$ is a fixed point, and $\mathbb{K}:\mathbb{R} \to [0,1]$ is a Borel-measurable CDF (or $\mathbb{K} (a) = \int^{x}_{-\infty} K(v) dv$).
For the CDF estimation, we need the following condition instead of Assumption (ref). The continuity of the random variable is essential, even for the kernel-smoothing CDF estimation.
The following theorem shows the presence of asymptotic biases for the kernel CDF estimator.
The CDF estimator can be rearranged as the sum of the infeasible estimator based on the true $\xi_i$ and the asymptotic biases. We present the result based on an infinite-order expansion because it yields the best possible condition of the relative magnitudes of $N$ and $T$ but it requires the validity of the infinite-order expansion, in particular the summability of the infinite series and they hold under technical regularity conditions as in the case of the density estimation in Theorem (ref). The biases of the forms $B_{\xi, 1}(x) / T$ and $B_{\xi, 2}(x) / T$ are the incidental parameter bias and the second-order nonlinearity bias, respectively. Note that $\hat F_{\hat \mu} (x)$ does not exhibit the incidental parameter bias as in the case of the density estimation. We also note that the standard kernel-smoothing bias of order $O(h^2)$ does not exist under asymptotic normality because it is asymptotically negligible under $Nh^3 \to C$ (see Lemma (ref) in Appendix (ref)). Consistency and asymptotic normality require the conditions $1 / (T h^2) \to 0$ and $N / (T^3 h^4) \to 0$, respectively, which asymptotically eliminate the higher-order biases. When using the standard bandwidth $h \asymp N^{-1/3}$ in the CDF estimation with second-order kernels, the conditions $1 / (T h^2) \to 0$ and $N / (T^3 h^4) \to 0$ are the same as $N^2 / T^3 \to 0$ and $N^7 / T^9 \to 0$, respectively, which are integrated to $N^7 / T^9 \to 0$. Note that we can weaken the relative magnitude condition by using higher-order kernels, which leads to a larger bandwidth, as in the density estimation.
The relative magnitude conditions for the kernel CDF estimation with second-order kernels are stronger than those for the empirical CDF estimation in JochmansWeidner2019 and Okui2017. The empirical CDF estimation is also easier to implement in practice. Hence, one should probably employ empirical CDF estimation in practice, and here we do not explore split-panel jackknife, CI estimation, and bandwidth selection for the kernel CDF estimation (although they are feasible). Nonetheless, the asymptotic biases for the kernel CDF estimation in Theorem (ref) are new in the literature, and they would be interesting in their own right.
We apply our procedure to panel data on prices of items in US cities. Our procedure allows us to examine the heterogeneous properties of the deviations of prices from the LOP across items and cities, and the difference in the degree of heterogeneity between goods and services.
Many empirical studies examine the heterogeneous properties of the level and variance of price deviations and the speed of price adjustment toward the long-run LOP deviation (see AndersonVanWincoop04 for a review). For example, EngelRogers01, ParsleyWei01, and CruciniShintaniTsuruga15 examine such heterogeneous properties and find that the LOP deviation dynamics are significantly heterogeneous across items and cities based on regression models. Our investigation below complements such empirical analyses by using our model-free procedure, as it provides visual information concerning the degree of heterogeneity.
We estimate the densities of the mean $\mu_i$, variance $\gamma_{0,i}$, and first-order autocorrelation $\rho_{1,i}$. We use the Epanechnikov kernel with the coverage error optimal bandwidth in calonico2018effect.\footnote{We also observed similar results with different kernels and different bandwidths.} The codes to compute the CIs and the optimal bandwidths are developed based on the { nprobust} package for { R} nprobust.
\paragraph{Data.} We use data from the American Chamber of Commerce Researchers Association Cost of Living Index produced by the Council of Community and Economic Research.\footnote{Mototsugu Shintani kindly provided us with the data set ready for analysis.} The same data set is used by ParsleyWei96, YazganYilmazkuday11, CruciniShintaniTsuruga15, LeeOkuiShintani13, and Okui2017. The data set contains quarterly price series of 48 consumer price index items (goods and services) for 52 US cities from 1990Q1 to 2007Q4.\footnote{While the original data source contains price information for more items in additional cities, we restrict the observations to obtain a balanced panel data set, as in CruciniShintaniTsuruga15.} The categorization of goods and services can be found in Okui2017.
We define the LOP deviation for item $k$ in city $i$ at time $t$ as $y_{i k t}=\ln P_{i k t}-\ln P_{0kt}$, where $P_{ikt}$ is the price of item $k$ in city $i$ at time $t$ and $P_{0kt}$ is that for the benchmark city of Albuquerque, NM. We regard each item--city pair as a cross-sectional unit, such that we focus on the degree of heterogeneity of the LOP deviations across item--city pairs. The number of cross-sectional units is $N = 48 \times (52 - 1) = 2448$ and the length of the time series is $T = 18 \times 4 = 72$.
\paragraph{Results.} Figure (ref) depicts the density estimates for $\mu_i$, $\gamma_{0,i}$, and $\rho_{1,i}$. In each panel, the solid black line indicates the density estimates without split-panel jackknife bias correction, the red dashed line shows the HPJ estimates, and the blue dotted line shows the TOJ estimates.
The estimation results with and without bias correction show that the LOP deviation dynamics are significantly heterogeneous across items. The density estimates without bias correction for $\mu_i$ are similar to those with bias correction. The results for $\mu_i$ also show that the mode of the heterogeneous long-run LOP deviations is close to zero, with a nearly symmetric, unimodal distribution. In contrast, the estimates without bias correction for $\gamma_{0,i}$ and $\rho_{1,i}$ are very different from the bias-corrected estimates. The bias-corrected estimates for $\gamma_{0,i}$ demonstrate larger variances for the LOP deviation dynamics, while the bias-corrected estimates for $\rho_{1,i}$ show more persistent dynamics with a more left-skewed distribution. These results suggest the severe impact of the incidental parameter biases, which highlights the importance of bias correction methods.
Figure (ref) depicts 95% point-wise confidence bands based on the HPJ estimates. The confidence bands are narrow, implying that our HPJ estimates seem to be precise and reliable.
Figure (ref) illustrates the HPJ estimates of $\mu_i$, $\gamma_{0,i}$, and $\rho_{1,i}$ for goods and services separately. The solid black lines are the HPJ estimates for goods, and the dashed red lines are those for services. The estimated densities and CDFs show that the heterogeneous properties are significantly different between goods and services. The densities for $\mu_i$ show that the long-run LOP deviation for goods generally tends to be larger than that for services (in an absolute sense). The estimation results for $\gamma_{0,i}$ and $\rho_{1,i}$ show that the LOP deviation for goods tends to be more volatile but less persistent than that for services. These results suggest that goods tend to have more volatile processes with faster adjustment speeds toward the nonnegligible long-run LOP deviation.
If we seek to examine the degree of heterogeneity of the LOP deviations across items and cities as in CruciniShintaniTsuruga15, our model-free results are informative in their own right. There are several possible sources of differences in the degree of heterogeneity, including the differences in trade costs across items (e.g., AndersonVanWincoop04) and differences in sale and nonsale prices across goods and services (e.g., NakamuraSteinsson08). Furthermore, our model-free results also suggest how we should model heterogeneity when implementing structural estimation for price deviations or change. For example, as our procedure demonstrates that the heterogeneous properties of goods and services differ, we should model unobserved heterogeneity differently for goods and services.
This paper presented nonparametric kernel-smoothing estimation to examine the degree of heterogeneity in panel data. The kernel density and CDF estimators are consistent and asymptotically normal under the relative magnitude conditions on the cross-sectional size $N$, time-series length $T$, and bandwidth $h$. Because of the presence of incidental parameter bias and nonlinearity biases, the relative magnitude conditions vary in the order of the expansions. Via infinite-order expansions, we derived the relative magnitude conditions that are suitable for microeconometric applications. We discussed the split-panel jackknife to correct biases, the construction of CIs, and the selection of bandwidth. We also illustrated our procedure based on an application on price deviations.
The authors greatly appreciate the assistance of Mototsugu Shintani in providing the price panel data. The authors would also like to thank Stephane Bonhomme, Kazuhiko Hayakawa, Koen Jochmans, Shin Kanaya, Hiroyuki Kasahara, Yoonseok Lee, Oliver Linton, Jun Ma, Yukitoshi Matsushita, Martin Weidner, Yohei Yamamoto, Yu Zhu and the participants of many conferences for helpful discussions and comments. Sebastian Calonico kindly helped us better understand the use of the { nprobust} package. All remaining errors are our own. Part of this research was conducted while Okui was at Vrije Universiteit Amsterdam, Kyoto University, and NYU Shanghai and while Yanagi was at Hitotsubashi University. This work was supported by the New Faculty Startup Fund from Seoul National University, JSPS KAKENHI Grant Numbers JP25780151, JP25285067, JP15H03329, JP16K03598, JP15H06214, and JP17K13715.
\setcounter{page}{1}
This supplementary appendix contains technical discussions omitted from the main text and Monte Carlo simulation results. Appendix (ref) presents the proofs of the theorems in the main body of the paper. Appendix (ref) presents the technical lemmas used in the proofs of the theorems. Appendices (ref) and (ref) present the technical discussions on the validity of the infinite-order expansions in Theorem (ref). Appendix (ref) presents the simulation results.