Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
92,264 characters · 7 sections · 115 citation commands
The disruption index suffers from citation inflation and is confounded by shifts in scholarly citation practice
\centerline{{ \bf Keywords:} Disruption index, innovation, secular growth, measurement bias, quasi-experiment}
{\bf One sentence summary: } The disruption index is susceptible to measurement error, which calls into question a growing body of research employing $CD$ to measure innovation trends and identify co-factors associated with team assembly. \\
\footnotetext[1]{ \ \ Send correspondence to: [email removed]}
Disruptive innovation refers to intellectual and industrial breakthroughs that sidestep conventional theory or practice by appealing to new value networks, to the extent that the disruptive entrants can quickly and unexpectedly overcome the competitive advantages characteristic of established incumbents christensen2013disruptive. In the case of scientific advancement, the process of disruptive innovation manifests as intellectual contributions that appeal to novel configurations of concepts and methods belonging to the knowledge network pan2016memory,BMConvergence_2021,RecombInnov_2022, thereby substituting prior combinatorial knowledge \`a la Schumpeter's theory of creative destruction schumpeter2013capitalism. Against this backdrop, scholars recently developed an index for quantifying citation disruption (denoted by $CD$) according to the implicit value of intellectual attribution that is encoded within the local structure of citation networks funk2017dynamic, with the objective of identifying intellectual contributions that appeal to new streams of intellectual attribution while subverting established ones.
Identifying the micro-level processes underlying disruption and quantifying its overall rate are fundamental to understanding scientific progress, and can in principle guide the management of institutions and policies that accelerate innovation. As such, the $CD$ index has received considerable attention and inspired a significant volume of follow-up research. However, there is a growing literature that challenges the definition and application of $CD$ to real scientific and patent citation networks bornmann2020disruption,ruan2021rethinking,petersen2023disruption,macher2023illusive,bentley2023disruption,leibel2023we,holst2024dataset. One stream of critique calls into question the long-term temporal decline in $CD$ reported by park2023papers, and the morose interpretation of its implications on the status and outlook of the scientific endeavor kozlov2023disruptive,holst2024dataset. In particular, Macher et al. macher2023illusive identify a substantial number of missing patent citations in the data analyzed in park2023papers, both at the beginning of their data sample and towards the end, which effectively reduces the number of backwards (i.e. references) and forward citations that were analyzed. Once correcting for the data omission, the re-analysis reveals an {\it increasing rate} of disruption in several patent domains central to the techno-informatic revolution of the last 30 years BMConvergence_2021.
Similarly, a recent and independent re-analysis by Holst et al. holst2024dataset shows that missing citations in the scientific publication data, which are more prominent in the early years of the sample, give rise to a subsample of publications with 0 references (which by definition correspond to maximum disruption value of $CD =1$). They show that the prevalence of these temporally biased data anomalies are entirely sufficient to generate the negative trend in $CD$ reported by park2023papers; upon correcting for these anomalies, they show that $CD$ trends insubstantial for both patents and publications. A third independent study of scientific publications by Bentley et al. also report an accelerating rate of disruption at the end of their data sample after developing a weighted variant of $CD$ that accounts for temporal shifts in the connectivity of real citation networks bentley2023disruption. In addition to these studies reporting an increasing rate of innovation according to temporal patterns in the citation network, a complementary approach based upon measuring combinatorial innovation in the knowledge network also reports a persistent innovation rate over time RecombInnov_2022.
A second stream of critique focuses on how $CD$ is defined, and the implications of its mathematical formulation on its reliability in bibliometric analysis wu2019confusing,bornmann2020disruption,leydesdorff2021proposal,ruan2021rethinking,petersen2023disruption. Taken together, these considerations further call into question a number of studies reporting statistical relationships between $CD$ and various covariates related to research production and team assembly -- such as career productivity, team size, citation impact, and the geographic dispersion of team members li2024productive,wu2019large,park2023papers,wang2023evaluating,Lin_remote_2023,li2024productive -- most notably because each of these covariates also grows with time. Statistical relationships between $CD$ and other time-dependent covariates are susceptible to omitted variable bias, which is a formidable source of measurement error in statistical analysis. As such, the connection between these two streams is the role of secular growth, which manifests as a temporal bias that underlies the data artifacts generating declining trends in $CD$ and the susceptibility of $CD$ to measurement error due to its non-linear dependence on the structure and rate of backwards citations.
Against this backdrop, here we contribute to these two streams by demonstrating how systematic measurement bias deriving from the inextricable densification of empirical citation networks, combined with omitted variables capturing confounding shifts in scholarly citation practice, further contributes to the mismeasurement of a “decline in disruptiveness” park2023papers. Our critique is centered upon the role of reference list in the definition of $CD$, and the implications of multifold increases in reference list lengths over the last half century, which is a fundamental source of `citation inflation' (CI) petersen_citationinflation_2018. Specifically, we apply four complementary methodologies that expose the underlying bias in $CD$ definition and its application -- deductive quantitative reasoning, computational modeling, a quasi-experimental test, and multivariate regression -- the results of which are consistent with a companion study petersen2023disruption based upon different data sources and regression model specification.
In order to test and validate the CI hypothesis -- that increasing reference list lengths confound the measurement of trends in $CD$ and covariate relationships -- we developed a quasi-experiment based upon the entire corpus of research published in {\it Proceedings of the National Academy of Sciences of the United States of America} (PNAS) over the 5-year period 2011-2015. Our identification strategy is based around comparing research articles published in the traditional format (print and online publication) to those published as long-form {\it PNAS Plus} articles (online-only publication) verma2012pnas,schekman2010creating -- as these two publication formats are nearly indistinguishable, aside from the longer reference list lengths of {\it PNAS Plus} articles. We conclude with a large-scale multivariable regression analysis of 7.8 million articles from 1995-2015, which accounts for the data quality and measurement biases identified, thereby improving upon the methodological designs employed in wu2019large,park2023papers. Results show that (a) the net effect size of temporal and team-size trends are at the level of noise and therefore inconsequential to science and innovation policy; and (b) by including appropriate controls and focusing on the metric itself instead of percentile values, the relationship between $CD$ and team size is instead increasing.
{\bf Definition of CD and its susceptibility to secular growth}\\ The disruption index $CD_{p}$ funk2017dynamic,park2023papers measures the degree to which an intellectual contribution $p$ (e.g. a patent or academic research publication) supersedes the sources cited in its reference list, denoted by the set $\{r\}_{p}$. The argument for $CD_{p}$ is that if future contributions cite $p$ but do not cite members of $\{r\}_{p}$, then $p$ plays a disruptive role in the citation network. As such, disruption can be inferred according to the local structure of the subnetwork $\{r\}_{p}\cup p \cup \{c\}_{p}$ that includes the set of citing nodes $\{c\}_{p}$ connecting to either the focal node $p$ or any member of $\{r\}_{p}$. According to its definition funk2017dynamic,park2023papers reformulated as a ratio wu2019large, $CD_{p}$ is calculated by identifying three non-overlapping subsets of $\{c\}_{p} = \{c\}_{i} \cup \{c\}_{j} \cup \{c\}_{k} $, of sizes $N_{i}$, $N_{j}$ and $N_{k}$, respectively -- see {\bf Fig. (ref)}(a,b) for a schematic illustration. In practice, a citation window (CW) is used to temper the effects of right-censoring bias, such that only citations occurring within a CW-year period are included in the subnetwork $\{c\}_{p}$. In what follows, we employ a CW = 5-year window denoted by $CD_{p,5}$, as in prior research park2023papers,wu2019large; however, the fundamental issues with the definition of $CD$ are independent of CW petersen2023disruption, and so for brevity we represent the general definition by $CD_{p}$.
The subset $i$ refers to members of $\{c\}_{p}$ that cite the focal $p$ but do not cite any elements of $\{r\}_{p}$, and thus measures the degree to which $p$ disrupts the flow of attribution to members of $\{r\}_{p}$. The subset $j$ refers to members of $\{c\}_{p}$ that cite both $p$ and $\{r\}_{p}$, measuring the degree of consolidation that manifests as triadic closure in the subnetwork (i.e., triangles formed between $\{r\}_{p}$, $p$, $\{c\}_{j}$). The subset $k$ refers to members of $\{c\}_{p}$ that cite $\{r\}_{p}$ but do not cite $p$. As such, wu2019large show that $CD_{p}$ can be calculated as the ratio
where the second equivalence is a simple re-organization of the equation to highlight the extensive quantity $R_{k} = N_{k}/(N_{i}+N_{j}) \in [0, \infty)$, which measures the rate of extraneous citation. Such re-organization facilitates deducing the scaling behavior of $CD$ associated with the number of articles $n(t)$ and the average reference list length per article $r(t)$ per year $t$. According to the scaling of network growth, $R_{k} \propto r(t)$ because $N_{k} \sim n(t)r(t)$ and $N_{i}+N_{j} \sim n(t)$, which is empirically validated in petersen2023disruption. Hence, because the ratio $CD^{\text{nok}} = (N_{i}-N_{j})/(N_{i}+N_{j}) \in [-1,1]$ is an intensive measure, $CD_{p}$ thus features a numerator that is bounded and a denominator that is unbounded -- and so $CD_{p}(t)$ converges to 0 over time because $R_{k}$ grows proportional $r(t)$. Moreover, even alternate disruption definitions such as $CD^{\text{nok}}$ are biased, because shifts in scholarly practice that manifest as network autocorrelations, such as self citation and journal impact-factor boosting, increase the overall rate of consolidation (triadic closure) measured by the term $N_{j}$. \\
{\bf Citation inflation: a measurement bias deriving from secular growth}\\ Citation inflation (CI) refers to the exponential growth of citations produced via the secular growth of the scientific endeavor petersen_citationinflation_2018. {\bf Figure (ref)(c)} illustrates how CI arises through the combination of increasing reference lists, denoted by $r(t)$, and increasing publication (or patenting) rates, denoted by $n(t)$, which significantly increases the density of citation networks over time. In the case of scientific research production, empirical growth rates estimated from the entire Clarivate Analytics Web of Science citation network show that total volume of citations generated by the scientific literature, $C(t)$, grows exponentially with annual rate $g_{C} = g_{n} + g_{r} = 0.051$; hence, the number of links in the citation network is growing by roughly 5% annually, corresponding to a doubling period of just $\ln(2)/g_{C} = 13.6$ years pan2016memory.
Moreover, as references tend to increasingly extend further back in time pan2016memory, the impacts of CI are not constrained to contemporaneous layers of the citation network, but instead are cross-generational. As such, the increasing density of citation networks manifests at both the source (reference) and destination (citation) of each new link, which is a temporal bias that is challenging to neutralize in the development of standardized network metrics. For example, in the 1980s the median-cited paper received 2 citations within 5 years of publication; however by the 2000s, this nominal quantity increased to 9 pan2016memory. In addition, there has been a paradigm shift towards online publishing that has facilitated greater publication volumes, faster publication times, and longer articles with longer reference lists. Take for example the journal {Nature} for which $n(t)$ has been roughly constant over the last 60 years: in 1970 the average number of references per articles was $\overline{r_{p}}=7$; by 2000 $\overline{r_{p}}$ increased to 24; and by 2020 $\overline{r_{p}} = 51$, a 7-fold increase over the 60-year period -- see {\bf Fig. (ref)}(d) and {\bf Fig. (ref)} in the {\it Supplemental Information}.\\
{\bf Issues with prior works analyzing trends in CD}\\ Park et al. park2023papers develop four robustness check approaches: comparison with alternative definitions of $CD$, normalization of $N_{k}$ in $CD$, regression adjustment, and synthetic randomization of the citation network. Yet each is susceptible to either data quality issues that are temporally biased towards early years, or measurement bias and omitted variable bias associated with the definition of $CD$.
First, because alternative disruption variants, namely $CD^\text{nok}$ and $CD^{*}$ bornmann2020disruption,leydesdorff2021proposal, are constructed around similar ratios, they are also susceptible to data quality issues as well as CI. However, it is notable that the alternative indices are less susceptible to CI, and indeed their trends shown in Extended Data Fig. 7 of park2023papers are markedly less prominent than what is shown for $CD$. In the particular case of $CD^\text{nok}$, the location of the average value is large enough that a vast majority of papers are classified as disruptive -- which begs for improvement. Moreover, in our companion study focusing upon a computational model, we show that even the tempered decline in $CD^\text{nok}$ can be attributed to CI petersen2023disruption .
Second, in order to attenuate the effect of CI, Park et al. park2023papers develop both `paper' and `field x year' normalized variants of $CD$ by modifying the factor $N_{k}$ in Eq. ((ref)). In the first case, they replace $N_{k}$ with $N_{k}-r_{p}$. However, according to scaling behavior arguments, since $N_{k}\sim n(t)r(t)$ then $(N_{k}-r_{p}) \sim (n(t)r(t)-r_{p}) \approx (n(t) -1)r(t)$, which does not generate the intended consequence. Moreover, the reduction of $N_{k}$ by $r_{p}$ still renders these normalized variants susceptible to the scenario exhibited in {\bf Fig. (ref)}(a,b), whereby citing just a single highly-cited paper causes $N_{k}$ to vastly exceed the difference $N_{i}-N_{j}$ in the numerator of $CD$, such that $CD_{p}\rightarrow 0$. In the second case of the field-year normalized variant, the average $r(t)$ for papers from the same field and year are subtracted, which amounts to the same issue, $N_{k}-r(t)\sim n(t)r(t)-r(t) = (n(t) -1)r(t)$.
Third, the regression adjustment implemented by park2023papers is poorly documented, as there is no model specification; moreover, Extended Table 8 only shows the estimated coefficients for the year indicator variable, and does not show the estimates for other controls. And according to Extended Data Table 1 in park2023papers, their model specification does not incorporate available publication-level factors that co-vary with $CD_{p}$, namely $r_{p}$, $c_{p}$ and the number of coauthors, $k_{p}$.
And finally, robustness checks based upon rewired citation networks are insufficient since the degree-preserving randomization holds constant $r_{p}$ or $c_{p}$. Hence, this randomization scheme can only be expected to attenuate biases attributable to correlated citation behavior -- contrariwise, biases deriving from data quality issues and CI can be expected to persist. Moreover, because shuffling the citation network reduces the rate of triadic closure to random chance, then this null model essentially converts all $N_{j}$ links into $N_{i}$ links. Consequently, the expected randomized value is $CD^\text{Rand}_{p} = ((N_{i}+N_{j})-0)/(N_{i}+N_{j}+N_{k}) = 1 / (1+R_{k}) >0$, which is positive definite and converges to 0 as $R_{k}$ increases. Park et al. then calculate a Z-score comparing the real and randomized values, $Z_{p}=(CD_{p} - CD^\text{Rand}_{p})/\sigma[{CD^\text{Rand}_{p}}]$, and plot the average value over time in Extended Data Fig. 8 park2023papers. Their results show extremely negative $Z$-score values (upwards of $2\sigma$ effect sizes). These deviations are also methodological artifacts: since $CD^\text{Rand}_{p}>0$, then all papers with $CD_{p}<0$ deterministically yield $Z_{p}<0$; moreover, the standard deviation $\sigma[CD^\text{Rand}_{p}]$ is extremely small because the chances of the randomization producing triadic closure is extremely small. Hence, with little variation to work with around a systematically small value $CD^\text{Rand}_{p} \approx 1 / (1+R_{k}) \sim 1/ r(t)$ , this randomization approach vastly underestimates the intrinsic scale of variation, i.e., $\sigma[CD^\text{Rand}_{p}] \ll \sigma[CD_{p}]$.
In a different study on the relationship between $CD$ and team size, Wu et al. wu2019large also develop robustness checks based upon multi-variate regression. However there is no clear model specification provided in their Supplementary Table 4; hence, in addition to omitting $c_{p}$ and $r_{p}$, it is unclear how they controlled for publication year. Moreover, the majority of their analysis is based upon descriptive trend analysis using {\it percentile} values of $CD$. This mapping of nominal $CD$ values to percentiles obfuscates the extremely narrow distribution of $CD$. At the same time, this modification of dependent variable generates the appearance of considerable effect sizes. Because most publications are concentrated around relatively small $CD$ values, a small idiosyncratic shift in $CD$ will generate disproportionately large shifts in the percentile value.
A third study analyzes the relationship between $CD$ and collaboration distance among coauthors Lin_remote_2023. This analysis also omits $c_{p}$ and $r_{p}$ from their multivariate robustness check (see Extended Data Table 1), and instead categorize papers as being before or after 2000 (a crude temporal control) and being solo-author or not (a crude team size control). As an example of persistent negligence for confounding factors, Lin et al. write: “For example, the 1953 paper on DNA by Watson and Crick is among the most disruptive works (D = 0.96, top 1%), whereas the 2001 paper on the human genome by the International Human Genome Sequencing Consortium is highly developing (D = -0.017, bottom 6%).” Yet these papers are from vastly different socio-technological eras, with the former produced by two coauthors citing $r_{p} = 6$ prior works, where the latter is attributed to 200+ coauthors and cites $r_{p} = 452$ prior works.
Additionally, results reported in Lin_remote_2023 are based upon the relative rates of $CD>0$ versus $CD<0$. This dependent variable is simply the sign of $CD$, which only depends on the numerator difference, $N_{i}-N_{j}$. While this choice may at first appear to be less susceptible to CI, it is still susceptible to the data quality issues identified by macher2023illusive,holst2024dataset, as well as confounding trends in the rate of triadic closure attributable to shifts in self-citation and other correlated citation behavior, which are more prominent in larger teams -- and larger teams are more likely to be extended across larger distances.
In response to these issues, two streams of critique have emerged regarding research on $CD$ -- one regards data quality issues and the other regards methodological issues. Regarding the former, citation networks based upon publication and patent data are susceptible to missing references, citations, and the classification of non-research oriented content (e.g. editorials) as the products of dedicated research. These data quality issues are more frequent for older publications, and less so for newer ones, as the modern publication industry benefits from information system features that were not available in the past (e.g. Digital Object Identifiers, and the web-based publication). As demonstrated by holst2024dataset, the frequency of publications and patents with 0 references is highly concentrated during early years. They show that this systematic bias in data quality contributes significantly to the decline in $CD$, since papers and patents with $r_{p}=0$ correspond to maximum disruption, $CD_{p} = (N_{i} - 0)/(N_{i}+0+0) =1$. Similarly, macher2023illusive show that left-censoring bias in US patent data means that patents from early years are artificially missing references to patents before the starting date of the dataset; upon correcting for these omitted references, which increases $r_{p}$ closer to their true value, then the decline in $CD$ for patents is greatly reduced.
A second stream of critique regards the methodological choices, e.g. omitted variables and the susceptibility of $CD$ to secular growth. By way of example, Bentley et al. bentley2023disruption modify the definition of $CD_{p}$ to account for CI according to both the number of publications $n(t)$ and the total number of citations produced, $C(t) = n(t)r(t)$. Their re-analysis reveals an increasing weighted $CD$ from the early 1990s through 2013. They also critique the results of park2023papers based upon natural language analysis, which also is susceptible to secular trends affecting the usage frequency and fashion of words.
To summarize, several recent analyses report findings of the following form: as $X$ increases, $CD$ decreases (where $X$ is time, individual publication rate, team size, nominal citations, and collaboration distance) li2024productive,wu2019large,park2023papers,wang2023evaluating,Lin_remote_2023. Accordingly, we conjecture that any variable $X(t)$ that increases over time will generate correlations of this pattern -- yet the degree to which such correlations survive confounders and wether they represent significant effect sizes is a more intriguing matter. To this end, here we seek to consolidate a growing number of critiques -- first by addressing methodological issues, and concluding with a regression framework that also addresses the data quality issues.
{\bf Computational model incorporating CI}\\ We use a tested and validated mechanistic citation network growth model pan2016memory to compare average trends in $CD$ calculated for synthetic networks generated with and without CI, which are otherwise identical in their construction. This generative citation network model belongs to the class of growth and redirection models krapivsky_network_2005,barabasi2016network and implements stochastic link dynamics that mimic preferential attachment barabasi2016network, while also incorporating other latent features of scientific production, in particular the exponential growth of $n(t)$ and $r(t)$. This model reproduces a number of statistical regularities established for real citation networks -- both structural (e.g. a log-normal citation distribution UnivCite) and dynamical (e.g., increasing reference age with time pan2016memory; exponential citation life-cycle decay petersen_reputation_2014).
Network growth in this model is governed by two complementary citation mechanisms that can be completely controlled by tunable parameters: (i) direct citation and (ii) redirected citation pan2016memory,petersen2023disruption. The second mechanism (ii) gives controls the rate of triadic closure in the synthetic citation network, thereby capturing correlated shifts in scholarly citation practice, such as citation trails illuminated by web-based hyperlinks that make it easier to find and cite prior literature; and self-citation among individuals and journals aimed at increasing their prominence in the attention economy. Moreover, the `consolidation' measured by $N_{j}$ in $CD_{p}$ explicitly measures the number of citations that feature triadic closure. In related work focusing on the details of this computation model petersen2023disruption, we find that CI has a stronger role than redirection in explaining the decline in $CD_{5}(t)$, and so in what follows we focus on the effects of CI.
Accordingly, here we focus on the network growth parameters that determine the rate of CI, which thereby facilitates measuring the impact of secular growth on two trends: (a) the decline in the average $CD_{p,5}$ value in year $t$, denoted by $CD_{5}(t)$; and (b) the scaling hypothesis connecting the rate of extraneous citations featured in the denominator of Eq. ((ref)) to the average number of references per paper, $R_{k}(t) \propto r(t)$. Note that the increase in $r(t)$ manifests from secular growth as well as shifts in scholarly citation practice. For example, papers with more authors tend have longer reference lists petersen2023disruption, partly because larger teams tend to write longer papers, but also because there are more authors seeking to benefit from self-citation.
Hence, we use the computational model to the CI hypothesis by generating two distinct network ensembles, each comprised of 10 random networks grown over $t=1...T $ periods (representative of publication years), terminating the network growth at $T \equiv 150$. Each network realization is seeded with common initial conditions -- e.g. the initial cohort features $n(1)=30$ nodes, each with $r(1)=5$ references. These synthetic citation networks are available for cross-validation, and can be used to develop alternative scientometrics that are neutral to CI DryadDisruption2023.
By construction, both network ensembles are statistically identical for the first 107 periods -- see {\bf Fig. (ref)}. The first ensemble incorporates the empirical rates $g_{n}=0.033$ and $g_{r} = 0.018$ for the entire $T \equiv 150$ periods, with networks reaching a final size of $N = 125,270$ nodes and 5,948,492 links with $r_{p} \equiv r(t=150) = 73$ after 150 periods. The second ensemble features the same $g_{n}$ and thus reaches the same final size of $N = 125,270$ nodes. However, the growth of $r(t)$ is suddenly quenched at $T^{*}=108$ to $g_{r}=0$, such that $r(t) = 34$ for $t\geq T^{*}$, which effectively `turns off' CI attributable to growing reference lists -- see {\bf Fig. (ref)(a,b)}. This sudden hault represents a hypothetical scenario in which journals were to impose a hard cap on reference list lengths in order to temper CI. Such caps are not inconceivable, as {\it Nature} provides a soft policy that “articles typically have no more than 50 references” in their \href{https://www.nature.com/nature/for-authors/formatting-guide}{formatting guide} for authors.
{\bf Figure (ref)}(a) compares the average $CD_{5}(t)$ for these two scenarios, which reproduces the magnitude of empirical decline reported in park2023papers for $t<T^{*}$. However, the two curves diverge for $t\geq T^{*}$, with the curve featuring quenched reference lists suddenly reversing course, thereby revealing the acute effect of CI on $CD$. Similarly, {\bf Fig. (ref)}(b) tracks the growth of $R_{k}(t)$, showing that this quantity is extensive when $r(t)$ is growing, and intensive when $r(t)$ is constant. {\bf Figure (ref)}(c-f) confirm the relationship $R_{k}(t) \propto r(t)$, and show that the proportionality is independent of the citation window (CW) used for calculating $CD_{CW}(t)$; see petersen2023disruption for additional empirical validation based upon a different dataset. Accordingly, our simplified network model demonstrates that CI fully controls the trends in $CD(t)$. \\
{\bf Quasi-experimental validation of the citation inflation hypothesis: Empirical analysis of $\vert CD_{p,5} \vert$} \\ Computational `toy models' are designed to capture the essential parameters underlying observed variation, while neglecting those features that are perceived to be non-essential. However, an unavoidable limitation to such approaches is determining what exactly are the essential parameters. For example, in our parsimonious growth model we do not account for various sources of heterogeneity underlying scientific publication, such as team size and reference list lengths, which both extending from just a few to several hundreds. Instead, all synthetic publications in our model from the same year have the same number of references, $r_{p} = r(t) \equiv $ average reference list length, so that we can rule out variation in this feature as a cause for the effect we are seeking to understand. Hence, in order to further test the CI hypothesis in an empirical setting, we again resort to a simplified scenario.
Specifically, we exploit the 2011 launch of a strategic publishing model developed by the journal PNAS, consisting of a long-form online-only publishing option -- called {\it PNAS Plus} -- to complement its traditional print option schekman2010creating. Article submissions were not processed, reviewed or prioritized according to the author-designated print option verma2012pnas, and so publications in these two formats are satisfactory counterfactuals for testing wether publications with larger $r_{p}$ are biased towards smaller $CD_{p}$. To demonstrate this causal link, we juxtapose the two sets of articles that are otherwise indistinguishable, on average, aside from {\it PNAS Plus} articles having longer reference lists.
Because computing $CD_{p,5}$ requires 5 years of post-publication citation data, in what follows we compare the disruptiveness of 18,644 research articles published in PNAS from 2011-2015. Notably, online-only {\it PNAS Plus} articles feature a different page numbering system, and so by inspecting this metadata for each article we identified 12.6% of the total sample as {\it PNAS Plus} articles. While the difference in the average $\vert CD_{p,5}\vert$ is incremental, the {\it PNAS Plus} articles do have smaller disruption values in magnitude across the bulk of the sample distribution -- see {\bf Fig. (ref)}(a,b). In terms of relevant citation network characteristics related to $CD_{p}$, {\it PNAS Plus} articles differ primarily in terms of $r_{p}$, as they feature $100 \times (57 - 41)/41 = 39\%$ more references per article, on average. Otherwise, the two subsamples are nearly indistinguishable in terms of citation impact ($c_{p,5}$) and team size ($k_{p}$) -- see {\bf Fig. (ref)}.
We exploit this quasi-experimental setting in order to distinguish between the following behavioral (i) and statistical (ii & iii) mechanisms that could contribute to declines in $CD$:
Park et al. park2023papers primarily attribute the observed decline in $CD$ to shifting balance of disruptive innovation captured by mechanism (i). However, they do not rule out mechanisms (ii) or (iii), which are not related to the innovation capacity of the scientific enterprise, but instead reflect the susceptibility of the $CD$ metric to statistical bias. Hence, we empirically test the CI hypothesis using a simplified CD metric, $\vert CD_{p,5} \vert$, which is not sensitive to mechanism (i). This modification is not dissimilar to the choice of alternative disruption metric employed by Lin et al. Lin_remote_2023, who base their results on the sign of $CD$, which avoids the measurement bias associated with mechanism (iii).
We test the relationship between $\vert CD_{p,5} \vert$ and various covariates using the model specification employed in our companion study petersen2023disruption (which is based upon a different dataset). Instead, here we use publicly available publication metadata from the {\it SciSciNet} open data repository lin2023sciscinet, which features pre-calculated $CD_{p,5}$, $k_{p}$, $r_{p}$. We use the following multivariate linear regression model,
which accounts for team size and the most relevant scalar citation network quantities relating to $CD$. We estimate the parameters of the model using the STATA 13 package “xtreg fe” using publication-year fixed effects; each covariate enters in logarithm to temper the right-skew in the distribution of each variable. For the full list of parameter estimates see {\bf Table (ref)}. We also tested the robustness of the parameter estimates by applying the same model to a larger sample of journals, comprised of 6.9 million articles from the same period, 2011-2015, which shows consistent results across a larger range of journals -- see {\bf Table (ref)}.
Results show a negative relationship between $\vert CD_{p,5} \vert$ and $r_{p}$: $b_{r} = -0.0039$; $p = 0.002$; 95% CI = [ -0.0055 -0.0023] which further supports the CI hypothesis. Put in real terms, a paper with twice as many references ($2r_{p}$) has a $\vert CD_{p,5} \vert$ value that is $b_r \ln(2) = -0.002$ smaller than if it had $r_{p}$ references. This scenario corresponds to a $0.6\sigma$ effect size, as the joint standard deviation across both PNAS subsamples is $\sigma[\vert CD_{p,5} \vert] = 0.0065$. Notably, the sign, magnitude, statistical significance level of $b_r$ is consistent with the analog coefficient reported in leahey2023types. The relationship between $\vert CD_{p,5} \vert$ and $k_{p}$ are not robust in sign, which is likely attributable to the small effect size compounded by the non-linear increasing relationship between $k_{p}$ and $r_{p}$ over time petersen2023disruption, which we address in the following section.
Moreover, this model facilitates estimating the differences in $\vert CD_{p,5} \vert$ between the {\it PNAS} and {\it PNAS Plus} deriving solely from the differences in $r_{p}$. Our results show that 100% of the difference in the average $\vert CD_{p,5} \vert $ between the two journal subsets are explained by $\delta$, the difference in the average $r_{p}$ across the two subsets -- see {\bf Fig. (ref)}(c). Hence, these empirical results definitively demonstrate that a significant portion of variation in $CD$ is attributable to variation in $r_{p}$. For this reason, the main results reported in park2023papers survived their robustness checks, e.g. the random rewiring they employed conserves $r_{p}$, and Extended Table 1 and Supplementary Table 3 show that they did not include $r_{p}$, $k_{p}$ or $k_{p}$ as publication-level covariates of $CD$.\\
{\bf Empirical analysis of $CD$ from 1995-2015}\\ There is considerable disagreement emerging from research analyzing the relationships between $CD$ and various other factors. For example, wu2019large mainly rely on descriptive methods to establish a negative relationship between $CD_{p}$ and the team size, $k_{p}$. Instead, petersen2023disruption and leahey2023types employ multivariate regression and report a positive relationship, and no relationship between $CD_{p}$ and $k_{p}$, respectively. One reason for the discrepancy emerging in the literature is a lack of consistency in the data and methodological specifications.
Hence, in this section we re-analyze publication-level temporal trends park2023papers and team-size trends wu2019large in $CD_{p,5}$ using pre-generated and publicly available citation network data from {\it SciSciNet} lin2023sciscinet. For consistency, we apply the same general model specification developed in the previous sections and also applied in petersen2023disruption. Furthermore, we restrict our analysis to publications that feature explicit signatures of research outcomes -- namely, those with sufficiently large $r_{p}$ that we can be confident that they are not editorials, commentaries, book reviews or other non-research based content that may be misclassified as such. This selection also excludes publications featuring substantial missing network data macher2023illusive, since these data quality issues effectively reduce $r_{p}$, and consequently give rise to spuriously large $\pm CD$; this selection also avoids the issue deriving from the surprisingly frequent singularity identified by Holst holst2024dataset whereby papers with $r_{p}=0$ generate $CD_{p}=1$. As such, we focus on the components of the citation network that are both conceivably and consequentially disruptive, in line with the originator's definition of disruption representing a form of breakthrough innovation christensen2013disruptive.
With this in mind, we ranked journals over the period 1995-2015 according to the number of publications with $10 \leq r_{p} \leq 200$. We then analyze publications from the top 1000 most prominent journals, which also satisfy the following criteria: team sizes in the range $1 \leq k_{p} \leq 25$, and citation counts in the range $1 \leq c_{p,5} \leq 1000$. This exclusion produces a 0.26% decrease in the sample size, resulting in 7,819,889 publications. By focusing on prominent journals, we can also calculate a journal-year normalized disruption index,
where $\overline{CD}_{j,t}$ is the average $CD$ value and $\sigma[CD]_{j,t}$ is the standard deviation calculated for publications from journal $j$ in year $t$. As such, this normalized disruption metric controls for year-specific factors such as journal publication modality (online, print, hybrid), as well as the characteristic value and variation CD according to the discipline associated with $j$, etc. As a final data quality assurance, we exclude publications with $\vert \text{Norm}CD_{p,5,j,t} \vert \geq 5$, which corresponds to just a 0.66% decrease in the sample size. The resulting sample is comprised of 7,768,207 publications (comprising 99% of the original data sample), with the most productive (least productive) journal featuring 138,883 (respectively, 3371) publications over the 21-year period.
Following these data quality refinements, we then re-analyze the temporal trend in $CD$ using the following model specification,
which incorporates squared terms to account for non-linear relationships. For example, because $N_{k} \sim n(t) r(t)$ appears in the denominator of $CD$, a linear correction for $r_{p}$ is insufficient. The coefficient $b_{r}=-0.0033$ (p-value $<$ 0.001; 95% CI =[-0.0042, -0.0025]) is negative, reflecting the residual impact of CI. For the full list of parameter estimates see {\bf Table (ref)}.
{\bf Figure (ref)}(a) shows the trend in the factor variable $\gamma_{t}$, which captures year-specific trends that persist in spite of the publication-level controls. Note that the regression adjustment robustness checks reported in Supplementary Table 1 of park2023papers does not report any of the field-year and paper-level controls, and so it is not possible to validate our results according to their covariates; in particular, they not include the covariates $r_{p}$, $k_{p}$ and $c_{p,5}$ in their model specification. The results of our reanalysis indicates that the residual trend in $CD(t)$ associated with time is at the level of noise, with the uptick in the regression adjusted $CD(t)$ after 2008 corresponding to just $0.06\sigma$ effect size relative to the baseline level in 1995.
In order to evaluate team-size trends, we leverage the journal-year normalized disruption index $\text{Norm}CD$ to estimate the standardized parameters of the model
As such, coefficients are measured in units of $\sigma[CD]_{j,t}$, which facilitates assessing the relative magnitude of effect sizes. The interaction term represented by $(\ln k_{p} \times t )$ controls for the tendency of larger teams to produce longer papers with longer reference lists petersen2023disruption. After controlling for temporal variation and CI, we find that $CD$ increases (albeit weakly) with team size (for $k_{p} \in [3,25]$) -- which is consistent with a statistically significant and positive coefficient associated with $\ln r_{p}$ identified in our companion study petersen2023disruption. As with the temporal trend, the net effect is at the level of noise, with the difference between $k_{p} = 2$ and $k_{p} = 25$ corresponding to just a $0.09\sigma$ effect size. These results are in stark contrast with wu2019large. One source of discrepancy is the methodology, as their descriptive analysis does not account for multivariable interactions. Moreover, Wu et al. base their analysis upon differentials in the {\it percentile} values of $CD_{p,5}$, which obscures the relatively small magnitude of the effect size obtained for nominal $CD$ values, which are extremely narrowly distributed around $CD=0$, as illustrated in {\bf Fig. (ref)}.
A growing body of research seeks to relate $CD_{p}$ to time-dependent covariates such as team size wu2019large, novelty leahey2023types, the geographic dispersion of team members Lin_remote_2023, and citation impact wang2023evaluating -- all are quantities that have systematically increased over time. A common pattern among these studies is a result of the form: as $X$ increases, $CD$ decreases. However, this class of results naturally follows from the susceptibility of correlations between $X(t)$ and $CD(t)$ to (a) temporal biases associated with the secular growth of the scientific enterprise, and (b) temporal biases associated with increasing data quality of the citation network data over time.
Data quality issues deriving from missing citations and references is a fundamental source of error identified by macher2023illusive,holst2024dataset that explains the anomalous decline in $CD$ reported by park2023papers. In the re-analysis by Macher et al. macher2023illusive, missing references at the beginning of the patent data artificially reduce $r_{p}$ for early patents; upon correcting for their omission, which effectively increases $r_{p}$ for those early patents, the negative trend in $CD$ largely disappears. Similarly, Holst et al. holst2024dataset show that a significant source of systematic error follows from including items with $r_{p}=0$ that generate $CD_{p}=1$ outliers. They show that these anomalies tend to occur earlier in the publication and patent datasets. Upon correcting these issues, they also show that the negative trend in $CD$ largely disappears. This second re-analysis also provides the full set of coefficients estimated in their regression adjustment analysis, which shows a negative correlation between $CD_{p}$ and $r_{p}$ ($\beta_{1}$ in Table S1 of holst2024dataset). Indeed, these data quality issues give rise to the same net effect as CI. Beyond data quality issues, another issue are the small effect sizes between $CD$ and covariates. Our re-analysis of temporal trends and team-size trends generate effect sizes at the 0.06$\sigma$ and $0.09\sigma$ level, respectively; moreover, the directions of the trends are opposite of what was previously reported park2023papers,wu2019large.
To summarize, even in the absence of data quality issues, the $CD$ index decreases over time due to two mechanisms unrelated to innovation -- one behavioral, and the other structural. Importantly, the disruption index does not account for confounding shifts in citation behavior (e.g. self-citation, impact factor boosting) that increase the rate of triadic closure measured by $N_{j}$ in the numerator of $CD$. Thus, decreases in $CD$ could follow from a number of competing mechanisms, some behavioral, others statistical in nature. Hence, in this work we began by testing the `CI hypothesis' that underlies the structural mechanism. Indeed, shifts in strategic behavior and normative practice are challenging to directly measure. For this reason, we confront this issue via computational simulation in our companion work petersen2023disruption. And in order to guide the development of unbiased citation-network metrics, we make available an ensemble of synthetic citation networks so that they can be used to test future citation-based indices for systematic bias DryadDisruption2023.
In short, our mixed method approaches consistently demonstrate that CI causes the denominator of $CD$ defined in Eq. ((ref)) to systematically increase as reference lengths increase over time, which causes $CD$ to converge to 0. According to its present definition, there is no clear way to correct for this dependence, since $CD_{p}$ is non-linearly related to $r_{p}$ via the factor $N_{k}$. This susceptibility is illustrated in {\bf Fig. (ref)}, which shows how a publication (or patent) needs to only cite one highly-cited publication for $N_{k}$ to increase to the extent that $CD\rightarrow 0$ independent of the difference $N_{i}-N_{j}$. The likelihood and magnitude of this one-off mechanism is increasing over time as a result of CI pan2016memory. Yet this issue even affects publications from the same cohort that have significantly different $r_{p}$. As a case example, we juxtaposed the disruptiveness of PNAS versus PNAS Plus articles published from 2011-2015, which differ primarily in their article lengths. Results show that nearly all of the difference in disruptiveness is attributable to the PNAS Plus articles having larger $r_{p}$ on account of their extended online-only publication format. Hence, a significant amount of the variation in $CD$ derives from variation in $r_{p}$, which could follow simply from journal-specific constraints on article lengths. By way of example, our analysis based upon normalized shows that the covariate with the largest effect size is $r_{p}$, which features a $0.14\sigma$ effect size for each unit change in $\ln r_{p}$ -- see {\bf Table (ref)}.
We conclude by suggesting a policy consideration for managing the scientific publishing enterprise, namely caps on reference list lengths to temper the effects of CI in research evaluation (see pan2016memory for computational modeling that instead explores the implications of a sudden increase in $r(t)$, i.e. simulating the emergence of online-only mega-journals and their impact on the citation network). Such limits on $r_{p}$ appear to cure the systematic decrease in $CD(t)$ petersen2023disruption, and could simultaneously address other shortcomings in citation practice such as surgical self-citation by authors ioannidis2019standardized and institutional collectives tang2015there, and journal impact-factor boosting martin2016editors,ioannidis2019user.
AMP designed the research, performed the research, participated in the writing of the manuscript, collected, analyzed, and visualized data. FA and FP contributed to the manuscript. AMP and FP designed the research.
We declare no competing financial interests.
Synthetic citation networks and code for analyzing them are available in the Dryad open data repository: DOI: \href{https://datadryad.org/stash/dataset/doi:10.6071/M3G674}{10.6071/M3G674}. Data used for the multivariate regression analysis, including the disruption index $CD_{p,5}$, $c_{p,5}$, $k_{p}$ and $r_{p}$, were obtained from the {\it SciSciNet} open data repository lin2023sciscinet.
\setcounter{page}{1} \setcounter{section}{0} {S\arabic{section}} \setcounter{table}{0} {S\arabic{table}} \setcounter{figure}{0} {S\arabic{figure}} \setcounter{equation}{0} {S\arabic{equation}}