arXiv · 2610.05462
The machine use of human beings
Abstract
Research results are prioritised by search engines, large language models and knowledge graphs chiefly through restatement rather than through the originating source. Although the open-access share of annual scholarly output has recently exceeded half, a substantial subscription corpus remains outside what machine readers may legitimately access under prevailing licences. A mechanism is described whereby subscription content is drawn into the open corpus through its citation and restatement in open-access articles---a process termed human mining, by analogy with the text and data mining performed by machines. An optimistic upper bound on what may thereby be inferred by a machine reader confined to the open literature is modelled, and the completeness, latency and licence constraints of that bound are quantified. A majority of subscription articles that have ever been cited is found to be reachable through at least one open citation, and the median interval between a subscription article's publication and its first open citation is shown to have fallen from thirteen years for work of 1990 to one year for work of 2020. The advantage once conferred by direct subscription access to machine readers has thereby been largely eroded, chiefly as a consequence of the growth of open publishing rather than any change in how researchers cite.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Daniel W. Hook. 2026-10-04. The machine use of human beings. https://arxiv.org/abs/2610.05462
Cite the original work for its findings. Save a collection to share your selection of sources.