Search arXiv⌕ Search

arXiv · 2610.05462

The machine use of human beings

Abstract

Research results are prioritised by search engines, large language models and knowledge graphs chiefly through restatement rather than through the originating source. Although the open-access share of annual scholarly output has recently exceeded half, a substantial subscription corpus remains outside what machine readers may legitimately access under prevailing licences. A mechanism is described whereby subscription content is drawn into the open corpus through its citation and restatement in open-access articles---a process termed human mining, by analogy with the text and data mining performed by machines. An optimistic upper bound on what may thereby be inferred by a machine reader confined to the open literature is modelled, and the completeness, latency and licence constraints of that bound are quantified. A majority of subscription articles that have ever been cited is found to be reachable through at least one open citation, and the median interval between a subscription article's publication and its first open citation is shown to have fallen from thirteen years for work of 1990 to one year for work of 2020. The advantage once conferred by direct subscription access to machine readers has thereby been largely eroded, chiefly as a consequence of the growth of open publishing rather than any change in how researchers cite.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Daniel W. Hook. 2026-10-04. The machine use of human beings. https://arxiv.org/abs/2610.05462

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Errors of LLM-Assisted Literature Retrieval in Environmental Science: A Comparison Study of Abstract versus Full-text Based Prompts

Large language models (LLMs) are increasingly used for literature search and synthesis. However, it is unclear whether they retrieve accurate bibliographic information in environmental science. Therefore, we quantitatively compared the errors of widely used LLM platforms in retrieving references related to original articles from five leading environmental science journals (Energy and Environmental Science, Nature Sustainability, Nature Climate Change, Lancet Planetary Health, and Environmental Science and Technology) published in 2024 to 2025. Claude, ChatGPT, Grok, DeepSeek, Perplexity, and Gemini were used as the LLM platforms. LLMs retrieved 10 references for each of the 50 randomly selected original article using either the article's abstract or its full-text as prompt. The retrieved references were subject to a multimetric score ratio combining validity of bibliographic data, Google Scholar link, digital object identifier, Scopus Electronic Identifier and relevance score (cited by or being the index paper), and the proportion of complete fabrication that failed all metrics. Abstract-only prompt yielded significantly higher accuracy than full-text one. This advantage was confirmed in multilevel mixed-effect multivariable regression after adjusting for journal, platform, and output order. Source journal and the position of a reference within the output list were also independently associated with retrieval accuracy, with lower-listed references associated with lower accuracy. These findings suggest that LLM assisted literature retrieval in environmental science remains moderately accurate and overall inconsistent, varying significantly by platform, journal, prompt type, and output position. Abstract-based prompting, as task-aligned information compression, may outperform full-text one in literature retrieval. Caution should be used when generalizing our findings.

cs.DL↗

MTRACE: Multilingual Retrieval-Augmented Generation for Temporally Diverse Text Corpora

Large multilingual knowledge bases expose temporally diverse information, yet retrieval quality remains sensitive to lexical variation and cross-lingual terminology shifts. We develop and evaluate MTRACE (Multilingual Temporal Retrieval-Augmented Generation with evidence grounding), a pipeline designed to test whether query expansion and multi-query fusion mitigate vocabulary mismatch in temporally layered text corpora, on the French and English subsets of MIRACL. Our approach integrates: (i) semantic query expansion (SQE) and multi-query fusion via Reciprocal Rank Fusion (RRF), targeting retrieval stability under query variation; (ii) a generation prompt enforcing strict grounding in retrieved evidence and explicit abstention when evidence is insufficient; and (iii) a modular architecture enabling systematic component evaluation. Ablation studies on Named Entity Recognition (NER) and embedding model selection demonstrate the importance of syntactic coherence in entity extraction and of self-retrieval and efficiency measurements for retriever selection. Our end-to-end evaluation over 50 constructed queries shows faithful answers for well-supported queries, correct abstention on unanswerable questions, and no re-scored similarity gains from multi-query fusion over single-query dense retrieval. By scoping our claims to a clean, text-only baseline, we separate these effects from OCR-noise confounds; direct measurement of diachronic lexical drift is left to future work. Code and configurations are available at \url{https://anonymous.4open.science/r/MIRAGE-8EAA/

cs.DL↗

The Rise and Fall of the Initial Era

Bibliographic data is a rich source of information that goes beyond the use cases of location and citation---it also encodes both cultural and technological context. This paper uses large-scale analysis of author-name representation in the Dimensions database to identify and characterise an ``Initial Era'' in the scholarly record, running from 1945 to 1983, during which initials were used in preference to full names on scholarly communications. We document this era's emergence, its persistence across countries and disciplines, and its rapid decline from approximately 2002 in conjunction with specific technological and policy changes in the bibliographic infrastructure. We argue that the Initial Era is exceptional in the four-century history of formalised scholarly communication, and we examine its implications for the visibility of researchers---particularly women---in the scholarly record, and for the broader project of bibliometric archaeology over digital research data.

cs.DL↗