Search arXivSearch

subject

cs.DL

cs.DL: explore 28 source-linked works published from 2026 to 2026, with original documents and citations.

This collection is a preview while coverage and quality are evaluated.

Search within this collection

Coverage and selection

Includes records with this source-supplied label or an explicit phrase match in their metadata. Matches indicate a mention, not proof that a paper uses a method or tests a material. Source versions are consolidated by DOI.

Sources: arxiv. Collection updated 2026-09-15. Counts describe this index, not the complete source archives.

Traceable Scholarship: Page Anchors and Ariadne's Thread for Humanistic Inquiry in the Age of Generative AI

Generative AI lets large language models produce scholarly-looking text within seconds, yet fluency does not equal valid explanation. The deepest risk is not factual error alone but the appearance that an explanation is already established without clear sources, page numbers, editions, or evidence. We liken the page anchor to Ariadne's thread: within the labyrinth of generative fluency, it is the thread that leads the scholar back to the source. This paper proposes Traceable Scholarship as the minimum normative condition for AI-assisted humanistic research, situating it across the three revolutions of knowledge infrastructure: print, digital, and generative AI. We introduce page anchors, dual page numbers, citation-first generation, NO_EVIDENCE, human verification, four-level compliance, and Scope Contract, and present AIH-Infra as a three-layer reference implementation: Contexture (document structuring), Open WebUI AIH-Infra (traceable knowledge base), and AIH-Infra MCP Server (agent gateway). A case study on a 29-volume Kant Akademie-Ausgabe knowledge base illustrates how traceability supports retrieval correction, evidence grading, and judgment downgrading. Traceability is not a software feature; it is the condition under which humanistic research can remain public and refutable in the age of generative AI.

cs.AI

SurveyAgent-HKA: A multi-agent framework for scientific survey generation with LLMs and human knowledge augmentation

Automatic scientific survey generation has become an important task in scientific document processing. The common approach of retrieving literature from a single source (e.g., arXiv) and generating surveys through a one-pass large language model (LLM) call often leads to limited reference coverage and, more importantly, fails to replicate the expert-driven revision process that is crucial for writing high-quality surveys. In this paper, we introduce SurveyAgent-HKA, a multi-agent framework that improves end-to-end scientific survey generation by incorporating knowledge derived from published surveys and peer-review comments. The framework decomposes survey generation into well-defined sub-tasks handled by LLM-powered agent. It first retrieves relevant papers from multiple sources and identifies key topics through clustering to construct an initial outline, which is then refined using outlines from related human-written surveys. Based on the refined outline, topic-focused papers are retrieved and re-ranked to select for drafting a well-grounded survey. Then, we identify common issues raised by experts in peer-review comments from published surveys to guide the revisions and finalize the survey. Experiments on two domains show that our approach outperforms mainstream baselines in citation quality, structural consistency, and content quality. Furthermore, our framework is efficient in both time and cost, making it a practical solution for broader AI-assisted scientific writing applications.

cs.CL

The Generative AI Gold Rush in Theoretical and Computational Research

Generative AI is changing the production conditions of theoretical and computational research, but its sys tem level effects require measures that separate plat form growth, field specific divergence, and production structure. We assemble 2,080 monthly observations for twenty arXiv archives from January 2018 through Au gust 2026 and a separate pseudonymized Mathematics author panel. A regularized convex synthetic control fitted through December 2025 identifies the January August 2026 anomaly, while spatial placebos, prior year pseudo holdouts, donor refits, and alternative preperiods assess comparative robustness. Mathematics recorded 47,127 list entries, 33.5% above 2025 and 11.9% above a synthetic counterfactual of 42,113 entries. Qualified donor and preperiod designs yield 9.6% to 14.9%, and Mathematics has the largest RMSPE ratio among fifteen eligible placebo archives. Subfield growth is broad, with 29 of 30 primary math.* categories expanding. The author panel shows a marked thickening of the repeated output tail. The share of active author units produc ing at least five submissions rose from 2.45% to 3.80%, while the ten submission tail rose from 0.21% to 0.49%. These results document a new and unusually large 2026 Mathematics production regime shift. Its timing and production structure, combined with independent evi dence on AI diffusion and verifiable research tasks, are consistent with delayed diffusion and capability thresh old mechanisms. The comparative design identifies the anomaly, and separate triangulation evaluates AI related explanations. The findings locate verification, selection, and attention as central constraints for research gover nance.

cs.DL

Westlake Scholar: AI-Enhanced Scholarly Discovery over an Institutional Repository

Institutional repositories (IRs) provide mature infrastructure for preserving and disseminating research outputs, but conventional record- and document-centric interfaces provide limited support for connecting deposited papers to related research and people. We present Westlake Scholar, an open-source, institution-grounded platform that adds four complementary artificial intelligence (AI) services to repository infrastructure: contextual paper reading, research-direction-guided paper discovery, publication-grounded expert discovery, and AI-generated research chronologies for scholars. The services draw on a shared institutional knowledge layer connecting approved publication records, paper content, and scholar--publication relationships. This allows the same paper to support contextual reading, cross-paper discovery, expert matching, and longitudinal views of scholarly work. Westlake Scholar provides an open and governable implementation of an institution-controlled AI layer that connects repository content, scholarly discovery, and researcher relationships while preserving provenance, human review, and institutional governance. A deployment at Westlake University, in operation since April 2026, demonstrates that the integrated system can operate in a live institutional setting.

cs.DL

Measuring the Novelty of Biomedical Papers Using the Latent Distances between Knowledge Units

Measuring the novelty of scientific papers is a central concern in research evaluation and scientometrics. From a recombination perspective, prior studies have largely focused on the co-occurrence of knowledge units to assess the novelty of scientific papers. However, these studies often overlook other relationships between knowledge units. This narrow view may result in inaccurate or incomplete evaluations of novelty for scientific papers. To fill this gap, this study introduces a comprehensive novelty measurement that incorporates three types of relationships between knowledge units: network, semantic, and hierarchical. These relationships are used to quantify the latent distances among knowledge units. Using a dataset of 142,036 articles published in PLoS ONE and a validation dataset from the H1 Connect platform, our results demonstrate that (1) each relationship type captures distinct latent distances between MeSH terms; (2) compared to the widely used indicators proposed by Uzzi et al. (2013), our measures show stronger alignment with peer judgements; and (3) combining all three distance metrics yields more effective identification of novel papers than using any single perspective alone.

cs.DL

Constructing the Field of Philanthropic and Nonprofit Studies: Evidence from Citation Networks

Philanthropic and nonprofit studies (PNPS) has grown rapidly as an interdisciplinary field, yet its intellectual structure and boundaries remain only partially visible. This study maps the field by combining journal- and keyword-based retrieval with citation network analysis, natural language processing, and large language model-assisted cluster labeling. Using 60,917 Web of Science articles, we identify major topical communities, examine their structural connections, and assess how well mainstream PNPS journals represent the broader landscape of related scholarship. The results show that PNPS is a loosely connected field dominated by three distinct "power centers": nonprofit organizations, social movements, and voluntary action. The analysis also reveals a broader disciplinary footprint than commonly recognized, extending beyond the social sciences and humanities into areas such as biomedicine and technology. These findings clarify the organization of PNPS and highlight opportunities for stronger integration across research domains.

cs.SI

MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relatively short texts. We introduce MultiGhostBench, a multilingual benchmark comprising 928 books generated by five recent LLMs across six languages and three scripts, with an average length of approximately 59K words per book. The benchmark supports evaluation under domain, author, and language shifts. Evaluation of representative AA methods shows that no single method consistently performs best across settings, and performance generally degrades under distribution shifts. Transformer-based detectors can retain generator-related information across languages, although transfer effectiveness varies by language pair, whereas statistical and fingerprint-based detectors are more language-dependent. We envision MultiGhostBench as a valuable resource for the development and evaluation of robust AA methods. The dataset and code can be found at https://github.com/GrecoMT/MultiGhostBench.

cs.CL

A Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts

In this work, we target Handwritten Text Recognition (HTR) in low-resource scenarios, which arise from underrepresented languages, rare scripts, and degraded visual conditions typical of historical documents. We introduce SCAM (Sahidic Coptic Ancient Manuscripts), a new line-level dataset built from digitized ancient manuscripts written in the extinct Sahidic Coptic dialect. The dataset reflects a realistic and challenging setting, as it combines heterogeneous acquisition conditions across libraries with typical manuscript degradations such as ink fading, bleed-through, and material deterioration. In addition to visual complexity, SCAM poses significant linguistic challenges due to the scarcity of resources for Sahidic Coptic, its uncommon alphabet, and dialect-specific diacritics. To support research in low-resource HTR, we benchmark several state-of-the-art approaches based on different paradigms, highlighting their limitations and strengths in this setting. Our results underline the gap between current HTR performance on well-resourced modern scripts and historically grounded, low-resource scenarios, thus providing a reference point for future developments.

cs.CV

Qualified Cross-References as a Verification Method: The Normative Environment of the EU AI Act

Legal cross-references are commonly represented as links between instruments or provisions. In a curated legal knowledge base, a link must also identify the legal character of the interaction, its supporting provisions and conditions, and remain consistent from either instrument. This paper presents a provision-level model and a protocol for qualified cross-references, developed through a bilingual corpus of fourteen instruments surrounding Regulation (EU) 2024/1689 (the AI Act). The model distinguishes direct textual reference, bounded presumption of conformity, substantive interaction without textual reference, mediated intersection, and institutional analogy; applicative interaction and definitional overlap are independent dimensions. Its methodological contribution is bidirectional inversion: a relationship documented from act A towards act B is reconstructed from B's perspective against both instruments. This verification tests provisions, qualification, direction, and conditions before deciding how to render the relationship from either side. Applied during construction, the protocol surfaced six incorrect article references, three inaccurate legal qualifications, and one divergence between two published descriptions of the same interaction. A reference to Regulation (EU) 2019/881 illustrates why qualification matters: the AI Act's bounded cybersecurity presumption differs from other product legislation's uses of the same certification framework. The contribution is a map of one regulatory environment and an explicitly specified method for making curated cross-reference knowledge bases inspectable and internally testable. A later, separately scoped corpus audit is discussed among the limitations, without pooling its findings with the original measurements; limited access to historical inputs constrains independent reproduction.

cs.CY

Lot Machine: Multimodal Lot Extraction from Auction Catalogs

For provenance research and art market studies, auction catalogs are an essential resource to trace specific objects over time and space. While historical auction catalogs follow established domain conventions, their internal formatting remains highly variable, and their large-scale analysis is currently restricted by the lack of machine-readable representations of the auction lots. We propose a pipeline to automatically extract structured lot-level metadata from German Sales, a large database of historical auction and sales catalogs from the 19th and 20th centuries. Using a manually annotated test set of representative catalog pages, we evaluate Vision-Language Models (VLMs) under varying prompt strategies and constrained decoding frameworks. To reflect the practical constraints faced by cultural heritage institutions, including budget, compute resources, and data privacy requirements, we benchmark the methods across different deployment modes ranging from commercial providers to locally hosted, quantized models. We find that commercial endpoints establish the performance ceiling, while institutional gateways offer a viable, privacy-preserving alternative. Local deployments remain feasible, but strictly require enforcing the output structure during generation to guarantee a valid JSON format. While varying degrees of human-in-the-loop correction are still necessary, this work demonstrates that a VLM-based pipeline can successfully unlock historical auction catalogs for large-scale automated analysis.

cs.CV

The zbMATH Open Knowledge Graph: Tracing Centuries of Mathematical Research

We present the zbMATH Open Knowledge Graph, a large-scale RDF knowledge graph (KG) covering more than 250 years of mathematical scholarship. Unlike existing scholarly knowledge graphs that primarily capture bibliographic metadata and citation structures, the zbMATH Open KG integrates expert-curated semantic content, including reviews, keywords, subject classifications, software references, and disambiguated authorship. This combination of domain-specific representation of mathematical knowledge and extensive temporal coverage supports analyses that require fine-grained exploration of mathematical concepts, research fields, and scholarly relationships over time. The resulting graph comprises 34 million entities and 168 million RDF triples represented using established Semantic Web vocabularies, supporting interoperability and FAIR data principles. We further demonstrate its capabilities through query-driven historically grounded scholarly exploration use cases, illustrating how the knowledge graph can surface relationships and patterns that may be difficult to identify from bibliographic and citation information alone. The zbMATH Open KG provides an open semantic infrastructure for studying the development of mathematical knowledge and tracing scholarly connections across centuries of scholarship.

cs.DL

Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation

Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We introduce a masked-citation task to compare human and LLM-generated citation behavior. For each citation context, an LLM generates a replacement citation sentence, producing a counterfactual corpus directly comparable to human citation. We analyze what, whom, and how models cite, using an LLM-as-a-judge to classify citation intent and a 20-million-edge coauthorship network to measure social distance between cited authors. Across six popular LLMs and 1,746 top NLP conference papers (63k+ contexts, 132k+ citations), three patterns emerge: (1) Compared with human citation, LLMs cite significantly less critically; (2) LLMs over-cite popular and older papers, a tendency amplified for contrasting citations where human writing more often draws on recent, niche work; (3) Whereas humans often cite within their close social network, especially for supporting citations, LLMs tend to draw on more socially distant authors. Together, these differences are double-edged: LLM citation reaches beyond a scholar's close collaborators while being less critical and amplifying visibility bias, reshaping the rhetoric and reach of scientific citation.

cs.DL

Beyond Human-Likeness: Mapping the Scientific Critique Profiles of LLMs and Human Reviewers

Large language models (LLMs) are increasingly discussed as tools for peer review, but their value is often assessed through human-likeness, perceived usefulness, or textual overlap with reviewer comments. This study shifts attention from whether LLMs resemble human reviewers to what functions of scientific critique they perform. Using ICLR 2025 peer-review data, we compare human reviews with LLM reviews generated under baseline and expert prompts. We operationalize scientific critique through two review acts, weakness critique and scientific questioning, and annotate point-level review text using five theory-guided frameworks: Anderson's knowledge types, Toulmin's argumentation model, Graesser's question depth, SOLO cognitive complexity, and Hattie's feedback functions. The results reveal a differentiated critique profile. Human reviews placed greater emphasis on scientific framing and revision guidance, more often identifying higher-order weaknesses and asking questions oriented toward improvement. LLM reviews showed higher rates of explanatory depth, integrative reasoning, and explicit argument structuring. Expert prompting did not make LLM critique uniformly more human-like; it partially narrowed some gaps but mainly amplified LLM-specific tendencies toward integration and formal argumentation. These findings show that LLM-assisted peer review changes the functional composition of review text, making it important to distinguish LLM-amplified critique from areas requiring human prioritization and accountable judgement.

cs.DL

Guiding LLM Peer Reviewers: The Impact of Score Anchors on Review Evidence and Accuracy

Large language models (LLMs) are increasingly used for research quality evaluation, with prior work exploring their scoring accuracy and the plausibility of review rationales. However, less is known about whether external score guidance changes the evidence presented in the generated review as well as the final score. This study uses 98 Allied Health Professions research outputs submitted for internal REF-style assessment, with specialist human review reports and adjudicated 1-4 reference scores. No-guidance baseline reviews are compared with oracle-guided reviews, where the supplied score is set to the rounded human reference score; extracted evaluation points are used to compare human and LLM evidence use. Using this design, oracle guidance improves scoring accuracy, with score-following checks showing that models do not simply copy the supplied score. Corrected score mismatches are associated with changes in the generated review frame, showing that the score signal can steer review rationales. This effect is direction-dependent: LLM reviews cover human strength or upgrade points more reliably than human weakness or downgrade points, with the weakest alignment for expert downgrade evidence. The results show that score-guided review generation can be evaluated at the level of review evidence, as well as the final score.

cs.DL

Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature

Tracing connections between historical texts is an important part of intertextual research, enabling scholars to reconstruct the virtual library of a writer and identify the sources influencing their creative process. These intertextual links manifest in diverse forms, ranging from direct verbatim quotations to subtle allusions and paraphrases disguised by morphological variation. Language models offer a promising path forward due to their capability of capturing semantic similarity beyond lexical overlap. However, the development of new methods for this task is held back by the scarcity of standardized benchmarks and easy-to-use datasets. We address this gap by introducing Loci Similes, a benchmark for Latin intertextuality detection comprising a curated dataset of ~176k text segments and 1,490 expert-verified parallels, including 945 labeled references from an existing dataset. Using this data, we establish baselines for retrieval and classification of intertextualities with both lexical methods and pretrained encoder language models.

cs.IR

Sustaining Dryad: Reconfiguring relationships and 'thinking like a business' to maintain open data infrastructures

Open data infrastructures underpin how open research is practiced across communities. While critical, the longevity of such infrastructures is far from certain, as they grapple with challenges related to technologies, dominant ideologies, and continued funding streams. This paper draws on a mixed-methods study of Dryad, a well-known open data infrastructure, to deeply explore one of these challenges: the quest for stable funding. Our analysis shows how Dryad has carefully reconfigured different forms of relationships and revenue models throughout its history to work towards financial sustainability. We identify four types of relationships with customers, collaborators, and competitors that have been critical to Dryad's financial evolution: reinforcing, forging, positioning, and excluding relationships. We argue that while implementing strategic, business thinking is a critical strategy, it is also one which shapes other factors important in sustainability: interpretations of value(s), community, and governance. We conclude by highlighting emerging tensions that provide insight for other open data infrastructures working to become financially sustainable. As a whole, our analysis focuses not just on financial mechanisms for funding open data infrastructures (although those emerge) but on the relationships which enable them.

cs.DL

China's Shrinking Home Bias and Rising Disruptive Impact: Evidence from a Global Citation Network Analysis

China has become the world's largest producer of scientific publications, yet concerns persist that this growth is inflated by excessive domestic citation practices. In this study, we analyze a citation network of over 45 million publications from Web of Science (1980-2025) to investigate China's home citation bias and research impact. Using a network reshuffling null model to control for the structural effect of publication volume, we find that China's home citation bias is less pronounced than commonly assumed and has been steadily declining over the past two decades. Chinese researchers do not exhibit a significantly stronger home citation preference than other major countries, indicating increasing internationalization rather than insularity. Furthermore, using the persistent disruption framework, we show that Chinese papers are converging toward American papers in their capacity to produce paradigm-shifting work. These findings challenge prevailing narratives about Chinese scientific home bias and suggest that China's advances in research impact.

physics.soc-ph

SoniMet - A tool for sonifying and visualizing the performance of single researchers

For centuries, the scientific community has predominantly relied on visual tools to communicate empirical results and complex datasets. While visual representations dominate bibliometric analyses, the human auditory system possesses sensitive capacities for processing complex temporal information, distinguishing intricate patterns, and tracking parallel data streams. Data sonification translates data relations into acoustic signals, offering an alternative method for data exploration, pattern recognition, and scientific communication. This paper applies this concept to the field of bibliometrics through metrics sonification-the auditory translation of bibliometric information-and introduces SoniMet (Sonifying Metrics), a web-based tool designed to visualize and sonify the publication and citation data of individual scientists (see https://sonimet.kennebec.co.uk). SoniMet connects to the OpenAlex database to retrieve bibliometric records and displays them on an interactive chronological timeline. The tool employs direct parameter mapping to translate citation impact indicators into non-speech sound: field-weighted citation impact and citation counts determine the pitch and volume of a synthesized note and are mapped to the acoustic echo strength. Although SoniMet expands the methodological toolkit for research evaluation, current limitations include its restriction to individual scholar profiles and the challenge of integrating transient audio files into traditional, text-based scientific publishing workflows. Future empirical user studies are necessary to systematically evaluate the analytical utility and cognitive benefits of metrics sonification compared to established visual methods.

cs.DL
Compare source metadata on this page
WorkPublishedSource identifierSource
Traceable Scholarship: Page Anchors and Ariadne's Thread for Humanistic Inquiry in the Age of Generative AI2026-09-052607.20916arxiv
SurveyAgent-HKA: A multi-agent framework for scientific survey generation with LLMs and human knowledge augmentation2026-09-052609.05938arxiv
The Generative AI Gold Rush in Theoretical and Computational Research2026-09-042609.04872arxiv
Westlake Scholar: AI-Enhanced Scholarly Discovery over an Institutional Repository2026-09-042609.05072arxiv
Measuring the Novelty of Biomedical Papers Using the Latent Distances between Knowledge Units2026-09-042609.05175arxiv
Constructing the Field of Philanthropic and Nonprofit Studies: Evidence from Citation Networks2026-09-032609.03291arxiv
MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts2026-09-022609.02379arxiv
A Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts2026-09-012606.15987arxiv
Qualified Cross-References as a Verification Method: The Normative Environment of the EU AI Act2026-09-012608.19194arxiv
Lot Machine: Multimodal Lot Extraction from Auction Catalogs2026-09-012608.30510arxiv
The zbMATH Open Knowledge Graph: Tracing Centuries of Mathematical Research2026-09-012609.00969arxiv
Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation2026-09-012609.01432arxiv
Beyond Human-Likeness: Mapping the Scientific Critique Profiles of LLMs and Human Reviewers2026-09-012609.01895arxiv
Guiding LLM Peer Reviewers: The Impact of Score Anchors on Review Evidence and Accuracy2026-09-012609.01905arxiv
Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature2026-08-312601.07533arxiv
Sustaining Dryad: Reconfiguring relationships and 'thinking like a business' to maintain open data infrastructures2026-08-312604.27580arxiv
China's Shrinking Home Bias and Rising Disruptive Impact: Evidence from a Global Citation Network Analysis2026-08-312608.28139arxiv
SoniMet - A tool for sonifying and visualizing the performance of single researchers2026-08-312608.30274arxiv

These are bibliographic comparisons, not experimental rankings. Follow the original document for methods and conditions.