Search arXiv⌕ Search

arXiv · 2609.32963

More of the same? Are scientific papers losing originality?

Abstract

Scientific output is growing rapidly, but it is unclear whether the expanding literature remains original or is increasingly repeating itself. Originality has many dimensions. One dimension can now be measured directly: how textually distinct a paper is from the work that came before it. Using semantic language models, we screened more than 26,000 full-text papers across eight fields in the CORE database (2015--2025) for passages resembling recent literature. We compared every paper against a fixed-size sample of earlier work, so increases in resemblance are not a by-product of literature growth. We find that the share of a paper's text resembling recent work was stable before 2021, then rose sharply, more than tripling by 2025. The increase came not from a few papers borrowing more heavily but from resemblance spreading across the literature: more papers now contain passages that echo recent work, and almost none of this growth is explained by cited work. We also found an analogous increase in passages expressing the same scientific idea. The pattern held across robustness checks and replicated in a second, independent database. Along this one measurable axis of originality, scientific writing has become less distinct from the work preceding it.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ilan Doron-Arad, Elchanan Mossel. 2026-09-26. More of the same? Are scientific papers losing originality?. https://arxiv.org/abs/2609.32963

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Which reported inputs govern molecular docking reproducibility? A benchmark from reporting audit to independent re-execution

Computational docking results can be re-executed only when the molecular system and protocol are specified, yet reporting checklists do not quantify how strongly individual method fields affect the reported score. We combined a span-verified audit of 50 open-access papers with controlled one-factor perturbations, a Vinardo scoring-function check and a Vina cross-docking extension covering 12 targets in 7 protein families, and re-execution of 37 published claims on local and independent cloud infrastructure. Search-related fields were sparsely reported, but perturbing box centre, box size, exhaustiveness or random seed produced median absolute score changes of no more than 0.08 kcal mol$^{-1}$; ligand protonation and identity produced median changes of 0.19 and 0.29 kcal mol$^{-1}$, respectively, whereas receptor structure produced the largest change (1.02 kcal mol$^{-1}$; n = 85; P < 0.001), and 22% of 116 unique reported ligand strings remained unresolved or ambiguous after deterministic name normalization. These measurements yielded empirical field weights for the NewScience Evidence score and identified exact receptor structure, machine-resolvable ligand identity and protonation state as primary reporting elements; 29 of 37 local and 27 of 37 cloud re-executions were within 2.0 kcal mol$^{-1}$ of the reported score, but these descriptive rates do not estimate experimental replication or a calibrated probability of reproduction.

cs.DL↗

Submission Responsibility Matters: Role-Aware Submission Quotas under Coauthorship

Author-level submission quotas are increasingly used to control growing peer-review load. Recent coauthorship-sensitive quota rules improve over fixed per-author limits by reducing the quota cost of multi-author submissions, often using harmonic authorship-credit models to prevent simple author-list padding. However, these rules conflate three distinct quantities: review burden, authorship credit, and submission responsibility. As a result, they can penalize genuine solo-authored work, treat all coauthors as equally responsible for a submission, and create bottlenecks for student-led papers when a faculty advisor appears on multiple unrelated submissions. We argue that submission quotas should be designed around the responsibility structure of a paper rather than only its number of coauthors. We formalize desiderata for quota rules, including venue-load control, padding resistance, role sensitivity, solo neutrality, and student non-blocking. We then propose a role-aware quota framework that assigns author-specific quota costs based on constrained roles such as lead author, regular coauthor, and designated advisor. The framework includes fixed, per-capita, and harmonic-style rules as special or limiting cases, while allowing venues to distinguish lead authors, corresponding authors, advisors, and peripheral contributors. We show how simple role constraints can preserve resistance to manipulation while avoiding several structural disadvantages of coauthor-symmetric quota rules. Our analysis suggests that role-aware quota mechanisms provide a more faithful and flexible foundation for managing peer-review load under modern collaborative authorship.

cs.DL↗

Enhancing RAMOSE, a Framework for Implementing REST APIs and Semantic-Actionable Outputs Over Data Sources

Scholarly infrastructures increasingly expose their data through REST APIs that follow shared specifications, such as the Scientific Knowledge Graphs - Interoperability Framework (SKG-IF), which defines a common data model, exchange format, and REST API for research information. Implementing such specifications over existing data sources, however, requires a development effort that many open infrastructures cannot afford. RAMOSE, the RESTful API Manager Over SPARQL Endpoints, is an open-source Python framework that reduces this effort by turning a declarative configuration file into a documented REST API over RDF triplestores. This article presents its second major version, which addresses nine new requirements. The new features include query orchestration across multiple SPARQL endpoints and non-RDF sources, with joins across their results; pluggable output formats and request parameters; pagination and caching; OpenAPI export; and write operations, with authentication for both API consumers and protected endpoints. A built-in module packages the format and filters that SKG-IF prescribes, letting a provider expose a compliant endpoint only through configuration. A functional comparison with nine similar tools, grounded in reproducible tests, shows that RAMOSE is the only one able to serve both RDF and non-RDF sources simultaneously within a single API operation, joining their results on arbitrary keys. Its first version served the OpenCitations REST APIs, which peaked at almost 38 million monthly requests between May 2025 and May 2026. The new version extends this deployment to the OpenCitations SKG-IF endpoint and is adopted by the GRAPHIA project to onboard data sources into its SKG-IF-based federation.

cs.DL↗