Search arXiv⌕ Search

arXiv · 2610.07329

Large scientific teams are more, not less, disruptive

Abstract

Modern science has moved decisively toward larger and more collaborative research teams. Yet an influential 2019 study reported that small teams are more disruptive than large teams, a conclusion difficult to reconcile with both the well-known trend toward increasing collaboration and the disproportionately large teams behind widely recognized disruptive research, such as work that has been awarded Nobel Prizes. Here we show that the reported disruptive advantage of small teams is largely an artifact of measurement rather than a fact about teams. Analyzing about 60 million scientific publications, we find that larger teams cite more references and that the standard disruption index declines mechanically as reference lists lengthen. Equalizing reference-count distributions across team sizes reverses the negative gradient within the index's own framework. Evaluating each paper against one reference at a time yields a reference-robust measure that distinguishes recognized breakthroughs from ordinary papers more sharply and more consistently than the standard disruption index across six independently curated benchmarks, including the Nobel Prize. Under this validated measurement, scientific disruption increases with team size across four scientific domains and in every decade from the 1970s to the 2010s. For scientific publications, we conclude that larger teams are more, not less, disruptive.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Junming Huang, Yu Xie. 2026-10-05. Large scientific teams are more, not less, disruptive. https://arxiv.org/abs/2610.07329

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Life after delisting: tracking the publication output and citation impact of Scopus-discontinued journals with OpenAlex

Selective databases such as Scopus periodically re-evaluate the titles they index and remove those that no longer meet their quality criteria. Once a journal is removed, its subsequent output is no longer recorded, and the consequences of this process have therefore received little attention. This study analyzes the publication output and citation impact of 514 sources delisted from Scopus between 2018 and 2024, based on 678,162 documents retrieved from OpenAlex over a window of three years before and after delisting. Differences between periods are tested with the Wilcoxon signed-rank test for paired samples, globally and by geographic region, field of knowledge and SJR quartile. Output grows until the year of delisting and falls sharply afterwards. The median annual output per journal drops from 35 to 22 documents (-37.1%; p < 0.001; r = 0.30), 61.7% of journals publish less, and 84 (16.3%) have no documents recorded in OpenAlex after delisting. Citations per document, measured over a three-year window, barely change when each journal is compared with itself (-2.4%; p = 0.079), and journals are almost evenly split between those whose impact falls and those whose impact holds steady or rises. The results indicate that delisting makes these journals less attractive as publication venues but does not penalize the citation of the work they publish to the same extent, which raises concerns from a research integrity perspective.

cs.DL↗

Persistence Paradox in Dynamic Science: Evidence from the Deep Learning Revolution

Persistence is often regarded as a virtue in science. In this paper, however, we challenge this conventional view by highlighting its contextual nature, particularly how persistence can become a liability during paradigm shifts. We focus on the deep learning revolution catalyzed by AlexNet in 2012. Analyzing the 20-year career trajectories of more than 5,000 scientists active in top machine learning venues during the preceding decade, we examine how their research focus and output evolved. We first uncover a dynamic period in which leading venues increasingly prioritized cutting-edge deep learning developments, displacing traditional statistical learning methods. Scientists responded to these changes in markedly different ways: those who were previously successful or affiliated with established teams adapted more slowly. Such persistence is positively associated with productivity but, after 2012, negatively associated with scientific impact. Most researchers, and the largest share of the field's output, cluster in a band of moderate persistence, pointing to a trade-off between output and impact, as well as to institutional frictions that make larger departures costly. These conclusions are robust to alternative identification strategies and to competing explanations such as topic popularity premiums and survivorship bias. Taken together, our macro- and micro-level findings suggest that, in this case, a paradigm shift creates an opportunity structure by devaluing the very expertise that conferred incumbents' advantage in the first place.

cs.DL↗

The Challenges of PROTAC Permeability Prediction

Cell permeability is a key bottleneck for PROTAC development, and public data available to model it is scarce and inconsistent. We adapt an expert-in-the-loop LLM extraction workflow to mine PAMPA measurements from the primary literature, recovering image-only structures by optical chemical structure recognition and hand-verifying every record, expanding the public record from 31 PROTACs to 87. Ridge models trained on PROTAC-DB 3.0 reach $R^2 = 0.67$ within that resource but collapse on the newly extracted chemistry ($ρ= 0.12$), while models trained on the new compounds transfer back successfully ($ρ= 0.80$). We conclude that the current composition of the published records, and not dataset size, is limiting the construction of more generalizable models, and we outline what would need to change in reporting practices for better data-driven permeability models.

cs.DL↗