Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
103,975 characters · 21 sections · 89 citation commands
Recovering latent linkage structures and spillover effects with structural breaks in panel data models
Innovation is widely acknowledged as a cornerstone of technological progress and plays a pivotal role in enhancing productivity Romer1990. Empirical research consistently demonstrates that innovation not only bolsters domestic productivity but also generates spillover effects, positively influencing productivity in other countries Coe1995, Potterie2001, Coe2009, Ertur2007. Despite the extensive body of literature on technology spillovers, the dynamics of these effects have received surprisingly little attention. This lack of focus contrasts with empirical evidence, which strongly suggests that spillover effects are not static but instead exhibit significant time-varying patterns. For instance, OECD2012 highlights that the global financial crisis and the public debt crisis had a pronounced negative impact on business R&D and innovation activities across a wide range of countries in 2008. Also documented is a remarkable decrease in foreign direct investment (FDI) and a disproportionate collapse in trade during the financial and debt crisis OECD2010, OECD2010outlook.\footnote{OECD2010 reports a significant 15% decrease in the volume of OECD imports and exports in the first two quarters of 2009 compared to the same period in 2008. This decline is attributed to the disproportionate fall in domestic demand, outputs, trade of capital goods, and the temporary drying up of short-term trade finance. OECD2010outlook shows that FDI inflows sharply declined between 2008 and 2009 in many countries.} Given the disruption in R&D activities by the crisis and considering that FDI and international trade are commonly regarded as key channels for technology spillovers keller2004, Coe1995, Potterie2001, a natural question is if and how economic and political shocks influence R&D spillovers?
Addressing this empirical question presents several econometric challenges. First, how to identify the latent spillover structure, i.e., which countries generate spillovers and to whom? Existing studies typically assume a predefined functional form of the spillover structure, constructing the adjacency matrix that summarizes cross-country spillovers as an exponential or linear function of an observable bilateral factor, e.g., geographic distance, trade volume, FDI, or linguistic similarity. While the use of each observed factor may be justified by a specific theory, spillover channels are likely influenced by multiple factors simultaneously and also by unobserved variables. Furthermore, the relationship between these factors and spillover structures may be nonlinear and complex. Consequently, the postulated adjacency matrix based on a predetermined function of an observable can lead to misspecification, introducing biases in parameter estimation hardy2020. Second, even with a correctly specified spillover structure, estimating the magnitude of pair-specific spillover effects poses additional challenges, because the dimensionality of these effects grows with the number of countries, often surpassing the available time-series observations. Lastly, capturing the temporal dynamics of the latent spillover structure is critical. Most existing studies on R&D spillovers assume that spillover structures and effects are static over time. In practice, cross-country R&D spillovers may undergo abrupt changes in response to economic and political shocks. Failing to account for such time variation obscures important insights into network dynamics and may render subsequent analyses based on the estimated spillover structure misleading.
To address these challenges, we propose a novel model and estimation approach that identifies time-varying latent spillover structures without relying on predefined determinants or functional forms of network formation. Specifically, we consider a linear panel data model where the outcome of each unit is influenced by the characteristics of other unit, whose coefficients are referred to as spillover effects, as well as covariates that do not generate spillover effects, whose coefficients are referred to as private effects. Both spillover effects and private effects, or either one, may vary over time. To capture this temporal variability, we model it as abrupt structural breaks. Our approach allows the spillover structure to remain fully unspecified, introducing high-dimensional parameters. Consequently, our methodology is designed to detect structural breaks in these high-dimensional parameters.
Our estimation strategy involves multiple steps. First, we apply a penalized least squares method to obtain preliminary estimates of the breakpoint and other high-dimensional parameters, under the assumption of sparse spillover structures. In the second step, we refine the breakpoint estimates using non-penalized least squares to improve accuracy. With the refined breakpoint estimate, we re-estimate the spillover and private effects within each regime. To eliminate the regulation bias and enable proper inference for the private effect estimator, we employ double machine learning (DML) in the similar essence of Belloni2014 and Chernozhukov2018. Departing from conventional DML, we propose a post-double-Lasso (least absolute shrinkage and selection operator) procedure to estimate the parameters needed for constructing the Neyman orthogonal score, and this additional post-double-Lasso step is essential for achieving a fast convergence rate of the private effect estimator.
We show that the proposed breakpoint estimator exhibits super-consistency, meaning that the probability of the estimator coinciding exactly with the true breakpoint approaches one as the cross-sectional dimension ($N$) and time-series dimension ($T$) grow. This property ensures that the estimation of unknown breakpoints has a negligible impact on the accuracy of spillover and private effect estimators. Consequently, the effects estimated under the unknown breakpoint are asymptotically equivalent to those obtained under the true breakpoint. Furthermore, we rigorously analyze the asymptotic properties of the private effect estimator. When the private effect is assumed to be homogeneous across units and constant over time, the estimation can theoretically leverage $NT$ observations. However, the accuracy of this estimation is influenced by the error in estimating the spillover effects, which diminishes at a rate no faster than $\sqrt{T}$. This limitation inherently restricts the convergence rate of the private effect estimator under standard DML inference theory. Hence, we employ post-double-Lasso and establish the conditions under which the proposed private effect estimator achieves $\sqrt{NT}$-consistency.
Using the proposed techniques to examine the relationship between R&D expenditure and total factor productivity (TFP), we identify a structural break in the cross-country R&D spillover structure in 2009, aligning closely with the onset of the global financial crisis. Notably, the spillover structure becomes markedly sparser after this break, with many R&D spillover channels originating from continental Europe, such as those from France, Germany, and Italy, disappearing. This finding indicates that financial crises weaken technology diffusion among countries, resulting in fewer spillover channels in the post-crisis regime. Further analysis reveals that the diminished R&D spillovers following the financial crisis can be partly attributed to reduced R&D expenditures in key European countries that had previously served as major spillover generators. This finding aligns with evidence from OECD reports OECD2012 and other studies using different datasets Babina2022,hardy-sever2021. It is also justified by microeconomic foundation that smaller and private firms with weaker balance sheets, as well as firms facing higher financial constraints, tend to cut back disproportionately on R&D investments during crisis episodes peia2022. Our study is the first to analyze the time-varying dynamics of cross-country latent R&D spillovers, demonstrating that financial crises affect the structure and magnitude of these spillovers. Importantly, our estimated spillover structure, while correlated with factors such as geographical distance, language similarity, and bilateral economic activities to some extent, cannot be fully explained by these observables. This suggests that the spillover structure is complex, driven by a combination of both observable and unobservable factors.
In addition to its empirical contribution, our work also enriches multiple strands of econometric literature. First, it relates to the growing body of research on modeling and estimating unknown spillover effects, or more broadly, recovering latent networks. Several identification and estimation techniques have been developed to uncover latent network structures, including network formation models Goldsmith-Pinkham2013, Hsieh2016, Hsieh2020, penalized estimation methods manresa2016,Paula:2020, and other approaches hardy2020, Lewbel2023, among others. Notably, these studies typically assume that the adjacency matrix remains constant over time.
Recently, efforts have been made to relax the static assumption of networks. For instance, Comola2021 explore the identification and estimation of treatment effects when networks change after an intervention, though their framework assumes that the change points and network structures are known. In different contexts involving network or graph data, several studies have focused on break detection using various methods, such as tensor decomposition of longitudinal network data Park2020, break detection in observed networks WangYuRinaldo2021, and modeling smooth or abrupt changes in Markov random fields Kolar2010. We differ from these studies by considering break detection in a regression model with a latent network structure, without pre-specifying the network models.
Two closely related studies in terms of topic, though differing greatly in techniques, are Han2021 and Chen2022. Han2021 model a time-varying adjacency matrix within a spatial panel data framework, capturing temporal variation by linking the binary adjacency matrix to a set of time-varying observed variables via a logistic function, along with unobserved factors modeled by a Markov process. They employ a Bayesian approach for estimation. In contrast, our model allows for both weighted and unweighted adjacency matrices and does not require specifying the functional form of network formation or the dynamics of unobservables, though it imposes sparsity on the adjacency matrix. Chen2022 focus on high-dimensional vector autoregressive (VAR) models, using penalized local linear approaches to estimate smoothly varying transition and precision matrices. We differ from this work by considering discrete structural changes. To the best of our knowledge, this paper is the first to analyze structural breaks in a panel regression model with a latent network structure.
Our work also contributes to the growing literature on detecting structural breaks in panel data models. Various methods have been proposed to estimate common structural breaks in the (low-dimensional) slope coefficients, where breaks affect all units simultaneously bai2010, baltagi&kao&liu2017, qian&su2016, li&qian&su. okui&wang2018 propose a penalized estimation technique to detect heterogeneous structural breaks, where the magnitudes and timing of breaks may vary across latent groups. Lumsdaine2022 extend this by studying break detection when group membership also changes after the break. Our paper differs from these studies by addressing breaks in high-dimensional parameters, which requires a distinct estimation approach and theoretical treatment.
The topic of structural break detection in high-dimensional models has recently received increasing attention, primarily in the statistical literature. For example, li2020 explore break detection in large covariance matrices. safikhani2020 detect break in large VAR models using fused-type penalization with a focus on computational efficiency, which differs from ours in both estimation and theoretical analysis. Lee2016 provide a theoretical analysis of high-dimensional break detection using Lasso under Gaussian errors, while Lee2018 extend this framework to quantile regressions. Our break detection and asymptotic analysis of the breakpoint estimator builds on and extends that of Lee2018 to a panel setting. A key contribution of our paper is the establishment of the super-consistency of the estimated breakpoint by exploring the cross-sectional variation, which is crucial for the inference of spillover and private effect estimators. The use of cross-sectional information also alleviates the requirement on the length of time series to a large extent.
Our work is related to the burgeoning literature on DML Belloni2014,Chernozhukov2018. A key difference is that the low-dimensional parameter in our framework, namely the private effect, can be estimated using both cross-sectional and time-series observations, much larger than the number of observations used to estimate the high-dimensional parameters. This increased sample size provides the potential to achieve a faster rate of convergence for the private effect estimator. Building on the DML literature, our estimator solves the Neyman orthogonal score and involves cross-fitting. However, we further innovate by proposing the use of a post-double-Lasso procedure as a preliminary estimate to compute the Neyman orthogonal score. Remarkably, we demonstrate that the proposed private effect estimator attains $\sqrt{NT}$-consistency, even though the spillover effect estimators converge at a slower rate no faster than $\sqrt{T}$.
The remainder of the paper is organized as follows. Section (ref) introduces the model. Section (ref) presents the estimation method. Section (ref) examines the asymptotic properties of the proposed estimators. Extensions, including the specification of break types and the heterogeneity of private effects, are explored in Section (ref). Section (ref) evaluates the finite sample performance via simulation. Section (ref) presents the empirical study, and Section (ref) provides concluding remarks.
A widely used model for linking the TFP and innovation employs the Cobb-Douglas function Potterie2001, Coe2009 as
where $f_{it}$ represents the TFP of country $i$ in year $t$, $\alpha_i$ accounts for unobserved country-specific heterogeneity, and $u_{it}$ denotes the idiosyncratic error term. The determinants of TFP include the domestic and foreign R&D capital stock, represented by $S_{it}^d$ and $S_{it}^f$, respectively, as well as human capital, denoted by $H_{it}$. These determinants are allowed to correlate with $\alpha_i$, and their contributions to the TFP are captured by the parameters $\gamma$, $\tau$, and $\delta$. The foreign R&D capital stock, $S_{it}^f$, is typically modeled as a weighted average of R&D capital stock of other countries, i.e., $S_{it}^f=\sum_{j\neq i}\omega_{ij}S_{jt}^d$, where $\omega_{ij}$ reflects the potential cross-country R&D spillover channels. Existing literature often assumes a specific, time-invariant spillover structure, meaning that $\omega_{ij}$ is predefined and remains constant over time. For example, the spillover effect are commonly modeled as a function of geographic distance Ertur2017, language similarity Musolesi2007, or a time aggregated measure of bilateral economic activities, such as foreign direct investment (FDI) Potterie2001 and international trade Coe2009. However, in practice, R&D spillovers are likely influenced by a combination of observable and unobservable factors, including geopolitical dynamics, historical ties, and other complex relationships. The relative contribution of domestic and foreign R&D to the TFP also vary markedly across countries. Furthermore, both the spillover channels and effects can shift due to major economic and political events.
To address these complexities, we propose a model that allows the spillover structure to remain latent and evolve dynamically over time, while accounting for country-specific R&D effects. Specifically, we extend the TFP Cobb-Douglas function as
where $S_{it}^f=\sum_{j\neq i}\omega_{ij,t}S_{jt}^d$, with $\omega_{ij,t}$ being unknown and time-varying. The R&D effects, $\gamma_{it}$ and $\tau_{it}$, are allowed to vary across countries and over time. Taking the logarithm of (ref) results in the following linear model:
where $\gamma_{ii,t}=\gamma_{it}$ and $\gamma_{ij,t}=\tau_{it}\omega_{ij,t}$. With a slight abuse of terminology, we refer to $\{\gamma_{ij,t}\}$ as the spillover parameter, as it encapsulates the impact of the spillover covariate, with $\gamma_{ii,t}$ capturing the heterogeneous direct effect of a country's own R&D, and $\gamma_{ij,t}$ capturing the pair-specific spillover effect from $j$ to $i$ for $i\neq j$. In contrast, we define $\delta$ as the private effect parameter, since it reflects the influence of non-spillover covariates. Following the majority of studies on the determinants of TFP Miller2000, Vandenbussche2006, we adopt the change specification, namely using the above model to explore the relationship among the changes in the variables.\footnote{Under the change specification, the slope parameters capture the short-run relationship, or put it differently, the deviations from the long-run equilibrium Coe1995, Potterie2001.} For notational simplicity, we let $y_{it}=\Delta\log f_{it}$, $x_{jt}=\Delta\log S_{jt}^d$, and $z_{it}=\Delta\log H_{it}$. We assume that the time variation of spillover effects is characterized by structural breaks. For the demonstration purpose, we consider the case of a single break occurring at an unknown time $b$. Thus, we obtain the following benchmark econometric model:
where $\mathbf{1}(\cdot)$ is an indicator function, $\gamma_{ij,B}$ and $\gamma_{ij,A}$ are the spillover effects before and after the break, respectively. The spillover network is typically sparse due to, e.g., distance decay Keller2013, and thus many elements of $\{\gamma_{ij,B}, \gamma_{ij,A}\}$ are expected to be zero. While $x_{it}$ and $z_{it}$ are both scalars, the estimation method and theory can easily accommodate a vector of covariates. In Section (ref), we extend model (ref) by allowing the private effect to be time-varying and heterogeneous (with a latent group pattern) and discuss the specification of breaks.
To write the model in a more compact form, we denote $X_{it}:= (x_{1t}, \dots, x_{Nt})'$ as the vector of all units' spillover covariate that potentially affects unit $i$'s outcome, $X_{it} (b):= (X_{it}'\mathbf{1} (t \leq b ) , X_{it}' \mathbf{1}(t > b))'$, and the coefficient of $X_{it} (b)$ is $\gamma_{i}:=(\gamma_{i1, B}, \dots, \gamma_{iN, B}, \gamma_{i1, A}, \dots, \gamma_{iN, A})^\prime$. Further, denote $W_{it} (b):= (1, X_{it}(b)', z_{it})'$, whose associated coefficient is denoted as $\beta_i:= (\alpha_i, \gamma_{i}^\prime, \delta)'$ for $i=1,\ldots,N$. Then, model (ref) can be written as
This representation indicates that fixed effects are addressed by incorporating individual-specific dummy variables. Our objective is to estimate the unknown breakpoint $b$, the spillover and private-effect parameters in $\beta_i$, with their true values denoted by $b^0$ and $\beta_i^0$, respectively.
This model generalizes the widely studied panel data models with common structural breaks baltagi&qu&kao2016,baltagi&kao&liu2017,qian&su2016 by incorporating spillover effects. Unlike standard panel data models, where breaks affect a small number of parameters, our framework addresses breakpoint detection in high-dimensional parameters, adding complexity to both estimation and theoretical analysis. Additionally, our model can be viewed as a generalization of certain network models, such as network autoregressive models Zhu:2017, by allowing unit connections to be both unknown and time-varying. Furthermore, our model extends the work of manresa2016 by permitting spillover effects to vary over time.
Our estimation strategy is implemented in multiple stages. First, we estimate the breakpoint. Given the estimated breakpoint, we then re-estimate the spillover effect for each regime and the private effect. Each stage consists of several sub-steps.
We begin by estimating the breakpoint. To achieve this, we first obtain preliminary estimates of the breakpoint, fixed effects, and slope parameters by Lasso. Specifically, denoting $\beta:=(\beta_1',\dots,\beta_N')'$, we solve the following optimization problem:
where
and $D(\beta, b)$ incorporates the $L_1$-penalty, given by
Here, $\lambda_{NT,B}$ and $\lambda_{NT,A}$ are tuning parameters, and $\phi_{ij,B}$ and $\phi_{ij,A}$ are adaptive weights, which can be computed, e.g., as $\phi_{ij,B}^2=\left[\sum_{t=1}^{b}x_{jt}^2/b\right]^{-1}\quad \textrm{and}\quad \phi_{ij,A}^2=\left[\sum_{t=b+1}^Tx_{jt}^2/(T-b)\right]^{-1}$. The penalty term $D(\beta, b)$ depends on $b$ through $\phi_{ij,B}^2$ and $\phi_{ij,A}^2$. Although Lasso effectively handles high-dimensional parameters and yields a breakpoint estimator $\hat{b}$ near its true value, $\hat{b}$ is not guaranteed to be consistent.
To obtain a breakpoint estimator with desirable theoretical properties, we refine it in a second step by re-estimating the breakpoint using least squares, based on the preliminary estimator $\hat{\beta}$,\footnote{In this refining step, we may also use the post-Lasso estimator of $\beta$, namely re-estimating $\beta$ based on the selected nonzero spillover effects from (ref). While this approach is expected to yield a breakpoint estimator with properties similar to those stated below, it requires adjustments in the theoretical analysis. To maintain theoretical transparency, we update the breakpoint estimator using $\hat{\beta}$ as in (ref).} without imposing any penalty terms, i.e.,
In the next stage, with the updated breakpoint estimate, we re-estimate the spillover effects of each regime using adaptive Lasso and estimate the private effect using DML in a similar spirit of CHHW2021 Belloni2014, manresa2016, Chernozhukov2018. This approach helps mitigate the bias of the private-effect estimator introduced by penalized estimation.
To implement DML, we further split each of the two regimes into two sub-samples: the main sample and the auxiliary sample. Let $\mathcal{T}_m$ be the main sample containing both the pre- and post-break regimes and $\mathcal{T}_a$ be the auxiliary sample. We apply post double Lasso separately to each sample. Specifically, using the auxiliary sample, we perform sparse regressions of $z_{it}$ and $y_{it}$ on $X_{it}(\tilde{b})$, respectively, i.e.,
where $\eta_i$ and $\alpha_i^*$ are the individual fixed effect, $\nu_{i}$ and $\gamma^*_i$ are both $2N\times 1$ vectors, and $e_{it}$ and $\tilde{u}_{it}$ are the error terms. We assume that $\nu_i$ is sparse, i.e., $\sum_{j\neq i}\textbf{1}(\nu_{i}\neq 0)<s_\nu$ for a positive and small constant $s_\nu$, and that $\operatorname{\text{E}}(e_{it}|X_{i1},\ldots,X_{iT})=0$. This relationship between $z_{it}$ and $X_{it}(\tilde{b})$ implies that $\alpha_i^* = \alpha_i + \eta_i'\delta$, $\gamma_i^* = \gamma_i + \nu_i'\delta$ and $\tilde u_{it} = u_{it} + e_{it}'\delta$. The sparsity of $\gamma_i$ and $\nu_i$ implies that $\gamma_i^*$ is also sparse. The double Lasso procedure identifies the sets of relevant covariates for each of the two regressions in (ref) using adaptive Lasso, with the adaptive weights computed based on the assumed structure of the error covariance. For example, if the errors are heteroscedastic and uncorrelated with $x_{it}$, the adaptive weights can be calculated as $\varphi_{ij}^2=1/|\mathcal{T}_a|\left(\sum_{t\in\mathcal{T}_a}x_{jt}\hat{e}^*_{it}\right)^2$, where $|\mathcal{T}_a|$ denotes the cardinality of $\mathcal{T}_a$, and $\hat{e}^*_{it}$ is a consistent estimator of $e_{it}$, obtained from a preliminary estimation using conservative weights that depend only on $x_{it}$ and $z_{it}$. If the errors are both heteroscedastic and autocorrelated, a HAC-type estimator may be used instead; see manresa2016 for a more detailed discussion. Let $\hat{\mathbf{s}}_i^{v*}$ and $\hat{\mathbf{s}}_i^{g*}$ be the sets of covariates selected by the adaptive Lasso for the two models in (ref). Then we combine these sets by taking their union to construct the final set of covariates used for estimation, namely $\hat{\mathbf{s}}_i := \hat{\mathbf{s}}_i^{v*} \cup \hat{\mathbf{s}}_i^{g*}$. Denote $X_{it, \mathbf{s}_i}(b)$ as the subvector of $X_{it}(b)$ containing the elements indexed by $\mathbf{s}_i$. Finally, we perform a post-Lasso step, re-estimating the two models in (ref) using OLS with the covariates selected in $ \hat{\mathbf{s}}_i$, i.e.,
The resulting estimators are denoted as $\tilde{\alpha}_i^a$, $\tilde{\gamma}_i^a$, $\tilde{\eta}_i^a$ and $\tilde{\nu}_i^a$, where the superscript “$a$” represents estimates obtained from the auxiliary sample.
With these estimates from the auxiliary sample, we estimate the private effect $\delta$ with the main sample $\mathcal{T}_m$ based on the Neyman orthogonal score. Specifically, we solve the following equationfor $\delta$
and can obtain a closed form solution
where the superscript “$m$” in $\tilde \delta^m$ indicates that the estimator is obtained with the main sample. We then swap the role of the main and auxiliary samples to obtain $\tilde{\delta}^a$. The final estimator of $\delta$ is obtained as the average of $\tilde\delta^m$ and $\tilde\delta^a$, namely $\tilde \delta = ( \tilde \delta^m + \tilde \delta^a)/2$. In the final step, we estimate the spillovers effects using post-Lasso with $y_{it}-z_{it}^\prime\tilde{\delta}$ as the outcome variable.
We summarize the estimation procedure in the following algorithm.
This section studies the asymptotic properties of the proposed method. Our key result establishes the super-consistency of the breakpoint estimator defined in (ref). This super-consistency property allows us to treat the breakpoint as known when justifying the penalized estimator of spillover effects and the DML estimator of the private effect. We then analyze the asymptotic distribution of the DML estimator for private effects. We use the superscript “0” to denote true values, e.g., $\beta_i^0$ represents the true value of $\beta_i$.
Some regularity conditions are needed. Recall that $W_{it} (b):= (1, X_{it}(b)', z_{it} ')'$, and denote $W_{it} := (1, X_{it}', z_{it}')'$.
Assumption (ref) requires that the covariates are exogenous. Given the construction of $W_{it}(b)$, this assumption can be decomposed into two parts: the error term having a zero mean, and the absence of correlation between the error term and any covariate, i.e., $E(u_{it})=0$ and $E(x_{jt}u_{it}) =0$ for any $j=1,\ldots,N$. Note that this assumption only imposes restrictions on the contemporaneous correlation between the error term and covariates, but not strict exogeneity or predeterminedness of the covariates. It also limits the degree of cross-sectional correlation by ensuring that the variance of the cross-sectional average of $u_{it} X_{it}$ is of order $1/N$. Assumption (ref) requires thin tails for the distributions of both the covariates and the error term, implying the existence of finite moments of any order. Assumption (ref) guarantees a limited degree of serial dependence in the covariates and the error term. Assumption (ref) resembles the compactness assumption commonly applied to the parameter space in asymptotic analysis of extremum estimators. However, due to the high-dimensional nature of the problem, it is expressed differently in this context. Assumption (ref) bounds the maximum eigenvalue of $W_{it}W_{it}'$ uniformly over $i$ and $t$. This condition precludes scenarios where any element of $W_{it}$ diverges, even as the dimension of $W_{it}$ grows with the sample size.
Conditions related to the structural break are also required. Let $\beta_{i,B} := (\alpha_i, \gamma_{i,B}', \delta')'$ and $\beta_{i,A} := (\alpha_i, \gamma_{i,A}', \delta')'$ denote the parameter vectors before and after the break, respectively.
Assumption (ref) is essential for establishing the asymptotic properties of the preliminary penalized coefficient estimators derived from (ref). It ensures that the error introduced by estimating the unknown breakpoint has a limited impact on the coefficient estimators. This assumption holds, for instance, when $W_{it}$ is strictly stationary. A similar assumption is also used in Lee2018 (2018, Assumption A.6). Assumption (ref) is needed for identifying the breakpoint. It requires that the (true) coefficients before and after the breakpoint differ sufficiently, and that a significant number of units are affected by the structural break. Finally, Assumption (ref) excludes structural breaks near the sample boundaries. It ensures sufficiently long pre- and post-break periods, allowing precise estimation of spillover effects in both regimes.
Since the estimation procedure involves penalization, assumptions regarding the tuning parameters and adaptive weights are necessary. Define the following quantities:
Let $p$ be the dimension of $W_{it} (b)$, which is $p=2 + 2N$ in our benchmark model (ref). The number of slope and fixed effects parameters is $N + 2N^2 + 1 $. Let $J(\beta)$ denote the index set of the nonzero elements of $\beta \in \mathcal{B}$, and define $s:=|J(\beta^0)|$ as the cardinality of $J(\beta^0)$, i.e., the number of nonzero elements in the true coefficient vector $\beta^0$. Let $\beta_J$ be the vector that matches $\beta$ in the entries indexed by $J(\beta^0)$ and has all other entries set to zero. Similarly, let $\beta_{J^c}$ be the vector that matches $\beta$ in the entries indexed by the complement of $J(\beta^0)$, with all other entries set to zero. Note that $\beta = \beta_{J} + \beta_{J^c}$ and that $\beta^0 = \beta_J^0$. Finally, let $|\cdot |_1$ denote the $L_1$-norm and $\Vert \cdot \Vert_2$ denote the $L_2$ (Euclidean) norm.
Assumption (ref) imposes restrictions on the degree of sparsity, the relative scales of the cross-sectional and time dimensions, and the order of the tuning parameter in the penalty term in (ref). This assumption comprises three key parts. The first part requires that the average number of nonzero coefficients per unit is sufficiently large. The second part imposes lower and upper bounds for $D_{NT}$. On the lower bound, $D_{NT}$ must be sufficiently large, with an order greater than $\log (Np^2)/\sqrt{T}$, to ensure sparsity in the estimated coefficients. On the upper bound, $D_{NT}$ is constrained by $s D_{NT} N^{c} \to 0$, where the constant $c$ can be arbitrarily small. For Gaussian $x_{it}$, $N^c$ may be replaced with $\log N$. If the support of $x_{it}$ is uniformly bounded over $i$ and $t$, $c=0$ is also valid. Overall, this part of condition ensures that $s \cdot (N^c \log (Np^2) / \sqrt{T}) \to 0$. Given $s \geq \sqrt{N}$ and $p = O(N^2)$, a necessary condition is $(\sqrt{N/T}) N^{c}\log N \to 0$, implying that the time series length must exceed the number of units. Finally, the third part requires a lower bound on $D_{\min}$, akin to the standard Lasso. Assumption (ref) is essentially a compatibility condition commonly used in the Lasso literature Lee2018, but with two notable differences. First, due to the presence of a structural break, the condition is applied separately for the two regimes. Second, the compatibility condition is imposed on the cross-sectional average. A sufficient condition for this compatibility is that $\min_t \min_i \lambda_{\min} (E(W_{it}W_{it}') )$ is bounded away from zero, where $\lambda_{\min}(\cdot)$ returns the minimum eigenvalue. Note that while $\sum_{t=1}^T W_{it}W_{it}'$ may be singular, especially when $W_{it}$ is high-dimensional, $E(W_{it}W_{it}')$ can still be non-singular.
Our goal is to analyze the properties of the refined breakpoint estimator. To this end, we first study the behavior of the preliminary estimators defined in (ref), in a similar spirit of Lee2018. Specifically, we evaluate the excess risk of these estimators, defined as:
Here, the expectations are taken with respect to $y_{it}$ and $W_{it}(b)$, while treating $b$ and $\beta_i$ as fixed. The following lemma establishes the order of this excess risk.
This lemma indicates that the preliminary estimators for the coefficients and the breakpoint lie within an appropriately defined neighborhood of their true values with probability approaching one. However, it does not guarantee that the preliminary breakpoint estimator coincides with its true value with high probability.
Next, we analyze the properties of the least-squares breakpoint estimator defined in (ref). Leveraging the fact that the preliminary coefficient estimator resides in the neighborhood of its true value, we build on the (low-dimensional) panel structural break literature Lumsdaine2022 to study the behavior of the refined breakpoint estimator.
This theorem demonstrates the super-consistency of the refined breakpoint estimator, meaning that the estimator converges to the true breakpoint with a probability approaching one. As a result, the DML procedure and the spillover effect estimator in the following stages can be analyzed as if the breakpoint were known.
While the super-consistency property allows us to disregard the error in breakpoint estimation when conducting inference for the DML estimator of the private effect $\tilde{\delta}$, the analysis becomes non-standard if we aim to achieve $\sqrt{NT}$-consistency. More specifically, although $NT$ observations can be utilized to estimate the private effect, $\tilde{\delta}$ is influenced by the estimation uncertainty of the spillover effects, which diminishes at a rate no faster than $\sqrt{T}$. This limitation constrains the convergence rate of $\tilde{\delta}$ under standard DML inference theory. In this section, we establish conditions under which $\tilde{\delta}$ attains $\sqrt{NT}$-consistency. For clarity and to emphasize the key ideas without overcomplicating the proofs, we impose the following assumptions, some of which are stronger than necessary and can be relaxed.
Assumption (ref) imposes restrictions on the distribution of the error terms in (ref), and ensures there is no endogeneity. Assumption (ref) requires the coefficients of the double Lasso models, $\gamma_i$ and $\nu_i$, to be sparse, with the numbers of nonzero elements bounded by a finite constant. Note that this sparsity assumption holds for all $i$. Even under this assumption, the set of non-zero parameters in the model is still high-dimensional. Each $i$ contains a finite number of non-zero parameters, and there are $N$ units. Thus, the number of non-zero parameters is of order $O(N)$. Assumption (ref) concerns the sets of Lasso selected covariates, $\hat{\mathbf{s}}_i^{v*}$ and $\hat{\mathbf{s}}_i^{g*}$, ensuring that these sets exhibit no random variation in the limit under sufficiently large time-series sample size. It also requires that the selected sets include all covariates with nonzero true coefficients. Notably, we do not impose variable selection consistency; in other words, the selected sets need not exactly match the true set of nonzero coefficients and may contain redundant variables. But, of course, this condition is more easily satisfied by using methods with variable selection consistency (i.e., oracle property), such as the adaptive Lasso zou2006adaptive and SCAD fan2001variable. Thus, it can also be replaced by lower-level assumptions on the tuning parameters and adaptive weights. Assumption (ref) states that the main and auxiliary samples produce comparable averages of the outer products of the double Lasso selected covariates. These averages converge to their expected values, which, under the stationarity assumption, remain constant over time. It is worth noting that the rate of the left-hand side of Assumption (ref) is typically $O_p (\log N / \sqrt{T}) $. Hence, to achieve the rate stated in the assumption, namely $o_p( \sqrt{T/N} )$, we would need $\sqrt{N} \log N / T \to 0$, which is satisified under Assumption (ref).
Importantly, the theorem establishes $\sqrt{NT}$-consistency of $\tilde \delta$, despite the nuisance parameters, such as $\gamma_i$, being at most $\sqrt{T}$-consistent. This result is made possible by leveraging the sparsity assumptions in the regressions of $y_{it}$ and $z_{it}$ on $X_{it} (b)$, combined with the use of estimators that possess variable selection consistency properties. The post-Lasso OLS in Step 4 precisely serves to eliminate the uncertainty of variable selection, which is crucial to achieve $\sqrt{NT}$-consistency of $\tilde \delta$.
When a structural break occurs, both the spillover and private effects may change. Moreover, considering cross-country heterogeneity in education and economic development, the private effect of human capital is also likely to vary across countries. This section discusses the estimation of time-varying private effects, the specification of break types, and the estimation of heterogeneous private effects.
If a structural break is present, it is reasonable to consider that private effects may also change. Thus, a natural extension is to model structural breaks in both spillover and private effects as
where $\delta_{B}^0$ and $\delta_{A}^0$ are the private effects before and after the break, respectively. To estimate this model, we can adapt Algorithm (ref) by redefining $z_{it}(b):= (z_{it}'\mathbf{1} (t \leq b ) , z_{it}' \mathbf{1}(t > b))'$, $W_{it} (b):= (1, X_{it}(b)', z_{it}(b) ')'$, and $\beta_i:= (\alpha_i, \gamma_{i}^\prime, \delta_B', \delta_A')'$. Using these modified variables, the breakpoint estimate can be obtained in two steps as in (ref) and (ref). With the estimated breakpoint, we implement DML (Steps 3--6) separately for each regime to estimate time-varying private effects. The spillover effect can then be re-estimated by regressing $y_{it}-z_{it}^\prime\mathbf{1} (t \leq b )\tilde{\delta}_B-z_{it}^\prime\mathbf{1} (t > b )\tilde{\delta}_A$ on $X_{it}(\tilde{b})$ using post Lasso.
In practice, researchers often lack prior knowledge about the number of breaks or the parameters affected by them. Therefore, it is essential to detect the presence of breaks and determine which parameters are impacted. One approach is to use an information criterion (IC), defined as
where $\widehat{Q}$ is the average sum of squared errors, $n_p$ is the total number of parameters, and $f(N, T)$ serves as a tuning parameter; for instance, setting $f(N, T)=\ln(NT)/(NT)$ corresponds to the Bayesian information criterion, which we employ in simulations and applications. Both $\widehat{Q}$ and $n_p$ depend on the number of breaks and the specification which parameters have a break. For example, a break in spillover effects alone implies $n_p=2N^2+N+1$, whereas breaks in both spillover and private effects yield $n_p=2(N^2+1)+N$. We evaluate the finite sample performance of the IC in determining the number and type of breaks in Section (ref).
Private effects may vary across units due to individual heterogeneity. To capture this heterogeneity, we employ a latent group approach, assuming that units within a group share the same private effect:
where $\delta_{g_i}$ represents the private effect for group $g_i$, and $g_i$ is the unknown group membership of unit $i$ that takes values from $\{1,\ldots, G\}$ for some pre-specified number of groups $G$. hahn&moon2010 provide sound foundations for group structure in game theoretic models. Bonhomme&lamadon&manresa2017 argue that a group pattern can be a good approximation even in the presence of individual heterogeneity.
In our empirical application, we employ a state-of-the-art clustering algorithm, namely the Sequential Binary Segmentation Algorithm (SBSA) developed by Wang2021, which classifies units based on preliminary estimation of individual-specific private effects. Specifically, we redefine $\beta_i:=(\alpha_i,\gamma_i^\prime,\delta_i)$ and estimate the breakpoint using Steps 1 and 2 outlined in Algorithm (ref). We then split the sample at the estimated breakpoint and estimate the individual-specific private effect with DML as $\tilde\delta_i=(\tilde{\delta}_i^m+\tilde{\delta}_i^a)/2$, where $\tilde{\delta}_i^m$ and $\tilde{\delta}_i^a$ are the estimators from the main and auxiliary samples, e.g., $\tilde{\delta}_i^a$ solves $$ \sum_{t\in \mathcal{T}_m} \left( y_{it} - \tilde{\alpha}_i^a - X_{it} (\tilde b)' \tilde{\gamma}_i^a - (z_{it} - \tilde{\eta}_i^a - X_{it} (\tilde b)'\tilde{\nu}_i^a )' \delta_i \right) \left(z_{it} - \tilde{\eta}_i^a - X_{it} (\tilde b)'\tilde{\nu}_i^a \right) =0. $$ With individual estimators at hand, SBSA then clusters units by minimizing within-group variation, iteratively segmenting groups until $G$ clusters are formed. Specifically, for a specified number of groups $G$, SBSA begins by sorting countries based on $\tilde{\delta}_i$, dividing them into two groups to minimize within-group variation. This process continues iteratively, further segmenting existing groups until $G$ groups are formed. See SBSA 1 in Wang2021 for further details. Once group membership estimates $\tilde{g}_i$ are obtained, we re-estimate the group-specific private effects using DML for each group, yielding $\tilde{\delta}_{\tilde{g}_i}$. Finally, the spillover effect is estimated by regressing $y_{it}-z_{it}^\prime\tilde{\delta}_{\tilde{g}_i}$ on $X_{it}(\tilde{b})$ using post Lasso.
The latent group approach requires specifying the number of groups $G$. To determine this, we adopt a commonly used method of minimizing an information criterion (IC), constructed analogously to (ref) bonhomme&manresa2015,okui&wang2018.
This section assesses the finite-sample performance of the proposed method. We focus on the accuracy of the estimated breakpoint. We also evaluate the estimators of spillover and private effects, as well as the performance of the IC in diagnosing break types.
We generate data according to (ref). The spillover covariate $x_{it}$ is sampled from a standard normal distribution, while the private effect covariate $z_{it}$ is correlated with $x_{it}$ and defined as $z_{it}=1/T\sum_{t=1}^Tx_{it}+\varepsilon_{it}$, where $\varepsilon_{it}$ follows a standard normal distribution. A structural break is introduced at $b_1^0=\lfloor T/3\rfloor$, resulting in different values for $\gamma_{ij,t}$ and $\delta_t$ across the two regimes. We consider two types of networks. The first is a discrete network generated using the Erd\"{o}s-R\'{e}nyi model. The second is a continuous network in which a random subset of the spillover parameters are set to zero and the remaining parameters are drawn from a normal distribution.
We examine six distinct data-generating processes (DGPs), incorporating two error processes and three types of breaks. In DGP 1.X, the error term is generated as i.i.d. standard normal. In DGP 2.X, the error follows an autoregressive process, $u_{it}=0.6u_{it-1}+\epsilon_{it}$, where $\epsilon_{it}$ is i.i.d. standard normal. For each of the two error processes, we consider three types of breaks:
We consider cross-sectional dimensions $N=(15, 30)$ and time series dimensions $T=(50,100)$. The simulations are conducted with 1,000 replications.
We first examine the estimated breakpoint. The accuracy of the breakpoint estimator is measured using the Hausdorff distance (HD), which simplifies to the absolute distance between the true and estimated breakpoints in the case of a single break, specifically $\textrm{HD}(\widehat{k},k^0)\equiv|k-k^0|.$ Table (ref) reports the HD ratio relative to the length of time dimension, i.e., $100\times \textrm{HD}(\widehat{k},k^0)/T$, averaged across the replications. The proposed method demonstrates a high level of accuracy in detecting breakpoints across all designs. Increasing the length of the time periods $T$ directly enhances the convergence of the breakpoint estimator and indirectly improves estimation accuracy by yielding more precise estimates of spillover effects. Consequently, the HD ratio decreases rapidly as $T$ grows. While an increase in cross-sectional sample size $N$ theoretically facilitates convergence, it may sometimes reduce the accuracy of breakpoint estimates when $T$ is small, particularly under autoregressive error terms. This is because a larger $N$ introduces a higher number of spillover effect parameters, complicating their estimation and, consequently, lowers the accuracy of breakpoint estimates. This high-dimensionality challenge explains the increasing HD ratio observed when $N$ grows from 15 to 30 and $T=50$ in some cases of DGP 2. However, when $T$ is sufficiently large (e.g., $T=100$), increasing $N$ reduces the HD ratio, confirming the consistency result in Theorem (ref).
Next, we assess the accuracy of estimated spillover effects. Following Paula:2020, we report three measures: the percentage of zero entries in the true adjacency matrix that are accurately estimated as zeros (“proportion of zeros to zeros”), the percentage of nonzero entries accurately estimated as nonzeros (“proportion of nonzeros to nonzeros”), and the overall root mean squared error (RMSE) of the adjacency matrix for each regime. The RMSE is defined as $\textrm{RMSE}= 1/N(N-1)\sum_{i\neq j}\left(\tilde{\gamma}_{ij,t}-\gamma_{ij,t}^0\right)^2$. Table (ref) presents the accuracy of adjacency matrix estimation for Erd\"{o}s-R\'{e}nyi networks, and Table (ref) displays the result for continuous networks. The adjacency matrix is estimated with high accuracy for Erd\"{o}s-R\'{e}nyi networks, where the proportion of correctly estimated nonzero elements approaches or exceeds 90% in most cases. Although the pre-break estimates of nonzero elements are less accurate when $N=30$ and $T=50$, this proportion improves greatly as $T$ increases. Due to uneven sample sizes before and after the break, the RMSE of pre-break spillover effect estimates is generally higher than that of post-break estimates, as expected. When the adjacency matrix is generated in a continuous structure, estimation becomes more challenging, yet the proportion of accurately identified (non)zeros remains high and improves as $T$ increases. Conversely, increasing $N$ reduces estimation accuracy, as reflected in both the proportions of (non)zeros to (non)zeros and the overall magnitude.
Finally, we evaluate the accuracy of DML estimates of private effects. Table (ref) presents the bias and RMSE of $\tilde{\delta}$ in each regime, averaged across replications. The private effect is estimated with high accuracy, and its precision depends on the accuracy of the breakpoint and adjacency matrix estimators, as discussed above. Both bias and RMSE generally decrease as $N$ and $T$ increase, supporting the validity of DML and our theoretical results.
We also evaluate the performance of the IC introduced in Section (ref) for identifying the type of breaks. Table (ref) shows the empirical probability of selecting each break type, with the bolded values indicating the highest probability in each row for a given DGP. Overall, the IC performs effectively. Specifically, in DGP X.1 (break in spillover effects only) and DGP X.2 (break in the private effect only), the IC correctly identifies the type of break with a probability close to or equal to 1. In DGP X.3 (break in both), the IC occasionally favors a specification with fewer parameters when $T=50$. However, as $T$ increases, the empirical probability of selecting the correct specification becomes dominant.
The role of innovation in driving economic growth has garnered substantial attention from economists and policymakers. Innovation, fueled by knowledge gained from research and development (R&D), expands the overall knowledge base, enabling more efficient resource use and boosting productivity within a country. Consequently, innovation is viewed as a crucial driver of an economy’s productivity growth Romer1990. Moreover, the knowledge embedded in exported goods can generate spillover effects, indirectly stimulating productivity growth in partner countries. Foreign R&D also offers direct benefits, such as facilitating the acquisition of new skills and technologies, including production processes, organizational methods, and materials Coe1995. Although the theory of cross-country R&D spillovers is widely recognized, empirical estimates of R&D spillover effects remain elusive, partly due to the ambiguity surrounding the specific channels through which these spillovers occur.
Existing studies on R&D spillovers typically assume a predefined structure for spillover channels based on economic or geographic measures. While each of these measures can be justified by a specific theory, it is likely that spillover channels are shaped by multiple factors—both observable and unobservable—that may interact in complex, nonlinear ways. Consequently, using an adjacency matrix based on a single measure (e.g., an exponential or linear function) risks misspecifying the spillover structure. Furthermore, existing literature often assumes that spillover effects are constant over time. This assumption is also restrictive given the economic and political shocks that impact international relations and corporate strategies. Ignoring temporal variation may obscure essential information about network dynamics and lead to biased estimates of spillover effects.
Motivated by these challenges, we revisit the study of cross-country R&D spillovers to identify the latent spillover structure without imposing a specific network formation mechanism. In particular, we extend (ref) to enable spillover and/or private effects to vary over time, i.e.,
where $y_{it}=\Delta\log f_{it}$, $x_{it}=\Delta\log S_{jt}^d$, and $z_{it}=\Delta\log H_{it}$. We model the time variation in parameters using structural breaks, with the presence and type of breaks determined through diagnostic tests. Later, we will extend model (ref) to incorporate potential heterogeneity in the private effect of human capital.
Our analysis focuses on 24 OECD countries, a sample widely examined in recent literature Potterie2001, Coe2009, Ertur2017. We use the most up-to-date data from 1981 to 2019, sourced from the OECD and the Penn World Table version 10.0. Total factor productivity (TFP) is measured at a constant national price level, normalized to 1 for 2017. R&D expenditure is calculated by multiplying the percentage of R&D expenditure in GDP by the real GDP at constant 2017 national prices (in millions).\footnote{We construct R&D expenditure in this way because the OECD's R&D expenditure data is not normalized in the same manner as other data from the Penn World Table, while the latter does not provide expenditure level data.} To create a balanced panel, we interpolate missing R&D expenditure values using data from Ertur2017 where available, or apply linear interpolation for values outside the sample period of Ertur2017. Human capital is measured by the average years of schooling for the population aged 25 and older, consistent with the existing literature.
Effective model estimation requires the knowledge of the number and type of structural breaks. We range the number of breaks from 0 (i.e., no breaks) to 2, and detect them sequentially. For each break, we consider three scenarios: breaks only in the spillover effect, only in the private effect, and in both. Figure (ref)(a) plots the IC results for different numbers and types of breaks. The results indicate that the best model features a single break in the spillover effects, while the private effects remain constant. However, the difference between this model and one that includes breaks in both the spillover and private effects is minor.
We proceed to estimate the single break. Both models, whether including or excluding a break in the private effect, identify 2009 as the estimated breakpoint. This estimated breakpoint precisely matches the significant event of the global financial crisis and great recession. The crisis, which began developing in the U.S. in 2007 and was subsequently amplified by the European debt crisis from 2009 onward, impacted most of the countries in our sample. The estimated breakpoint aligns with documented evidence showing that innovation activities declined remarkably following the financial crisis of the late 2000s. For instance, OECD2009 reports that domestic and foreign companies listed on U.S. stock markets reduced their R&D expenditures by 6.6% in the first quarter of 2009. The reduced R&D investment can be explained by tighter financial constraints caused by the crises campello2010, peia2022. It is therefore unsurprising that R&D expenditure and its correlation with total factor productivity (TFP) changed after the crisis. Additionally, the crisis likely altered international relationships through various channels, such as foreign direct investment (FDI) and trade, which, in turn, shifted R&D spillover channels over time. Further explanations of the estimated breakpoint will be provided after discussing the spillover effects before and after the break.
Figure (ref) presents the heatmap of the estimated adjacency matrix across two distinct regimes. The results reveal that the short-run spillover structure is predominantly sparse, with only a limited number of countries exhibiting connections with others. Notably, this network becomes even more sparse following the structural break, with the network density, defined as the ratio of nonzero entries to the total possible edges, decreasing by around 51% after the break. We discuss the spillover effects in the two regimes in turn.
Before the structural break, approximately 2/3 of the countries generated R&D spillovers, with six countries, namely France, Italy, Canada, the Netherlands, and Spain, emerging as key sources, transmitting their R&D benefits to multiple destinations. Around half of the countries received spillovers from various sources, with Japan, Korean, and Sweden standing out as the largest beneficiary. Japan gained not only from its domestic R&D efforts but also from those of Western European nations, e.g., France and the Netherlands. Similarly, Korea profited from its own R&D activities and those of several European partners including Denmark and Spain. These findings underscore the strong connections between these European countries and their significant Asian allies. The identified spillover channels align with the substantial imports of machinery and transport equipment—goods that embody the highest level of technology—by Japan and Korea from Europe. Sweden receives the most R&D spillover from Switzerland, Finland, and Canada. Notably, while the growth literature widely recognizes the private R&D effect on a country's economic growth, our analyses indicate substantial variability across nations. For instance, strong private R&D effects are evident in Japan, Finland, and Korea; moderate effects are observed in Canada and Spain, while for some countries, such effects remain too weak to be detected.
Following the structural break identified in 2009, both the network structure and the magnitude of R&D spillover effects exhibit remarkable changes. Notably, the post-break period is characterized by a more sparse linkage structure. Only about half of the countries in our sample continue to generate and receive R&D spillovers, and most of these countries are connected to only a single partner, either as an influencer or recipient. This stands in sharp contrast to the pre-break period, which featured several major generators and recipients that interact with multiple countries. In fact, a considerable number of pre-break R&D spillover channels originated from key European countries, such as Germany, France, the Netherlands, and Italy, dissipate after the break.
Further examination reveals that the R&D expenditure greatly declined for many European countries, particularly in the years immediately following the structural break. This trend aligns with various institutional reports documenting a sharp drop in business R&D and innovation in most OECD countries around 2009. The fact that the breakpoint year of 2009 coincides with the onset of the European debt crisis lends strong support to the theory and much evidence that financial crises can have profound and potentially long-lasting negative impacts on innovation. The negative impact of crises on R&D expenditure can find its micro foundation, primarily through the channel of tighter financial constraints. Specifically, financial crises are known to create tighter financial constraints for firms. For instance, campello2010, using survey data from a broad range of firms in North America, Europe, and Asia collected at the end of 2008, show that firms worldwide planned significant budget reductions across nearly all policy areas for 2009 in response to the crisis. Especially, the survey reveals a plan of cut in technology spending by over 10% in American and European firms during 2009. Furthermore, financial constraints are widely recognized as a critical factor influencing innovation. For example, Himmelberg1994 identify a strong, significant relationship between internal financing and R&D activities for small high-tech firms in the U.S. Considering the exacerbation of these constraints during financial crises, peia2022 argue that firms facing higher risks of financial strain and greater reliance on external financing tend to cut R&D investments disproportionately during crisis periods. Smaller and private firms with weaker balance sheets are also particularly prone to reducing R&D expenditures during financial disruptions. Historical evidence further supports this view. Babina2022, using a century of U.S. patent data, document significant and sustained declines in patenting following the Great Depression due to disruptions in financial access, especially among young and inexperienced inventors. hardy-sever2021 reach similar conclusions, showing persistent reductions in R&D investment during banking crises using industry-level data from developed countries. This body of evidence aligns with our observation of reduced R&D expenditure in the years following the structural break. It also helps explain our finding that the primary R&D spillover generators in the post-break period were largely countries less affected by the financial crisis, such as Canada and New Zealand.
While there is substantial empirical evidence on how financial crises impact innovation, studies specifically examining the effects of crises on R&D spillovers are scarce. Our analysis shows that the structure of R&D spillovers became more sparse following the global financial crisis. This finding can be attributed partly to the decline in innovation during the crisis, as discussed earlier, and partly to the adverse effects of the crisis on cross-border economic interactions. For instance, numerous studies and reports have highlighted a sharp decline in international trade starting in late 2008. This decline was driven by multiple factors, including tighter credit conditions, disproportionate fall in domestic demand, and decreased output and trade of capital goods OECD2010outlook. Also, severely impacted by the crisis is the FDI, as the liquidity constraints triggered by the crisis affected both firm owners and potential foreign buyers, leading to a marked decrease in cross-border mergers and acquisitions OECD2010outlook. Given that cross-border economic activities are important channels for R&D spillovers, it is not surprising that the spillover structure became more sparse due to the reduced bilateral interactions prompted by the financial crisis.
To better understand the estimated latent spillover structure, we regress the estimated adjacency matrix on several bilateral variables commonly used to construct the adjacency matrix. These include the geographic distance (Dist), a dummy variable indicating whether two countries share a border (Contg), a dummy variable for countries with a common official language (Lang), the trade flow (Trade) and foreign direct investment (FDI) from CEPII.\footnote{The trade flow and FDI are averaged over time in the pre- and post-break regimes, respectively, and are expressed in units of one million dollars. Distance is expressed in units of 1,000 km. Data source: http://www.cepii.fr/CEPII/en/bdd_modele/bdd_modele.asp} Given that the adjacency matrix is bounded below by zero, we estimate a Tobit regression for each regime and present the results in Table (ref). Also reported are the OLS estimates. Overall, we find that geographic and demographic variables partially explain the estimated adjacency matrix, but substantial variation remains unexplained. The effects of these variables also vary remarkably across regimes. In the pre-break regime, the Tobit model produces weakly significant estimates (at 15%) of geographic distance, contiguous and common language dummies, suggesting R&D spillover occurs more likely between countries that are closely located, share the common border and language. The bilateral economic activities explain little variation of the adjacency matrix. In the post-break regime, the significance of all proxies decreases except the contiguous dummy, partly there is less variation in the adjacency matrix due to more sparsity. This result shows that the financial crisis disrupt R&D spillover with spillover remaining only among the highly close countries. Furthermore, the model fit indicates that, in both regimes, the short-term spillover structure cannot be fully explained by geographic, cultural, and economic variables alone, let alone by a single observable. Therefore, it is valuable to consider data-driven methods to uncover latent R&D spillover structures.
The role of human capital in productivity has garnered considerable attention in the growth literature. While human capital is generally regarded as beneficial for productivity, the magnitude and significance of the effect often vary across studies, depending on the specific samples analyzed, the quality of data, and the measures of educational outcomes used. For example, pritchett2001 finds that increases in educational attainment did not lead to corresponding growth in productivity in many developing countries, attributing this to poor institutional and educational quality. In a recent review, psacharopoulos2018, drawing on a meta-analysis of 1,120 estimates in 139 countries from 1950 to 2014, highlight that returns of education can be low or insignificant, especially in contexts where education does not translate into relevant skills or productive employment.
We revisit the role of human capital using Model (ref), controlling for time-varying R&D spillovers. Our diagnostic analysis with IC indicates that the private effect of human capital remains largely constant after the structural break (see Figure (ref)(a)). Thus, we estimate a time-invariant effect of human capital using the DML, while allowing R&D spillovers to shift at the estimated breakpoint. We begin by estimating a homogeneous effect of human capital across countries, and will later explore potential heterogeneity. The DML estimate for the human capital effect is approximately 0.134 but not statistically significant. This result is in line with Ertur2017, who, using the same set of OECD countries (but over a shorter time horizon) and variables, also report an insignificant human capital effect and of the similar magnitude\footnote{Their estimates range from $-0.45$ to 0.23 across different subsamples and specifications.} after accounting for cross-sectional dependence and R&D spillovers. As we measure human capital by years of schooling following the literature, our finding also aligns with the broader insight from the growth empirics that the quantity of education alone does not necessarily drive growth; rather, the quality of education plays a crucial role. For example, Islam1995 found an insignificant effect of education on growth using panel data from various samples, including OECD countries, attributing this to differences in the quality of education. The author also points out the disparity between panel data and cross-sectional analyses, with the former focusing on temporal relationships and often reporting an insignificant link between human capital and growth, whereas the latter often showing a positive relationship. Hanushek&Kimko2000 also underscore that educational quality is more important for growth than educational quantity. In the similar spirit, Vandenbussche2006, using panel data from OECD countries, argue that only skilled human capital contributes to growth, rather than the total human capital. They also find that the impact of human capital on productivity is weaker for the countries further from technology frontier. These arguments partly explain the insignificant estimate of human capital effect in our context after controlling for R&D effects.
Next, we examine potential heterogeneity in the human capital effect. While allowing for country-specific effects seems conceptually appealing, it suffers from the incidental parameter issue and estimation instability. Traditional approaches often involve interacting human capital with an observable, such as a dummy indicating a country's G7 membership Coe2009, Ertur2017, openness, physical capital, or labor force Miller2000. However, the cross-country heterogeneity is likely to be driven by multiple factors jointly, both observed and unobserved, and the postulated group structure (i.e., which countries belong to which group) may not coincide with the true heterogeneity pattern. To flexibly model cross-country heterogeneity and minimize potential misspecification, we adopt the latent group structure model described in (ref), assuming that the human capital effect is homogeneous within each group. Importantly, we do not impose a predefined group structure; instead, we estimate it based on the data. The number of groups $G$ is determined using the Bayesian IC (BIC) similar to (ref) bonhomme&manresa2015,okui&wang2018. By applying the SBSA method from Wang2021 with $G$ ranging from 1 to 6, we plot the IC in Figure (ref)(b). It shows that $G=1$ is the best specification, confirming the reliability of the estimate obtained under the homogeneity assumption.
To assess the robustness of our results, we re-estimate the spillover and human capital effects using $G=2$, the second-best specification indicated by the IC, which allows for some degree of heterogeneity. The resulting human capital effects are positive in both groups, but of distinct magnitudes. However, these estimates remain insignificant, which supports the IC result that favors the homogeneity assumption. Compared to the estimates under $G=1$, a slightly higher number of spillover links are detected in both the pre- and post-break periods. Nonetheless, the overall pattern of R&D spillovers becoming more sparse after the break remains salient. More detailed results are available in Appendix (ref).
This paper investigates the impact of financial crises on R&D spillovers, introducing a novel model and estimation method for analyzing time-varying spillover effects without pre-specifying a linkage structure. Our approach recovers the latent spillover network and identifies the unknown breakpoint at which spillover and/or private effects experience a shift. We establish the super-consistency of the estimated breakpoint, which guarantees that the estimators of spillover and private effects under unknown breakpoints are asymptotically equivalent to those obtained under true breakpoints. We also establish the $\sqrt{NT}$-consistent DML estimator of the private effect, despite the fact that the spillover effects are estimated at a slower rate. An analysis of data from 24 OECD countries over a 38-year period reveals that the cross-country R&D spillover network became sparser following the 2009 financial crisis, suggesting a decline in innovation and disruption to its spillover mechanisms. The ubiquity of spillovers in various economic and social contexts, such as corporate performance, crime, and educational outcomes, extends the applicability of our approach beyond the domain of technology diffusion.
There are numerous avenues for future research, and we highlight two here. First, while we have demonstrated the super-consistency of our breakpoint estimator, this result does not enable statistical inference for the true breakpoint, as the asymptotic distribution of the estimator is not provided. According to our findings, the singleton set $\{\tilde{b}\}$ serves as a valid confidence set for $b^0$ at any confidence level, since $\mathbb{P}(\tilde{b} = b^0) > 1 - \alpha$ asymptotically for any $\alpha > 0$. However, such a confidence set fails to capture the statistical uncertainty associated with $\tilde{b}$. Developing methods for statistical inference on breakpoints in our framework would require entirely different approaches from those used for estimation. Second, our model assumes that the random variables in the system ($x_{it}$, $z_{it}$, and $u_{it}$) are only weakly dependent over time. However, in some applications, variables may exhibit strong serial dependence, such as in the case of random walk processes. Exploring spillover effects involving variables with unit roots or those within cointegrated systems containing multiple integrated variables would be both theoretically insightful and practically valuable.
During the preparation of this work the authors used Grammarly in order to improve the readability and language of the manuscript. After using this tool/service, the authors reviewed and edited the content as needed and takes full responsibility for the content of the published article.