Search arXivSearch

SEARCH · Search arXiv

Results for “cs.DL”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

5,200 recordsLinked to original sources

Are Widely Known Findings Easier to Retract?

Failures of retraction are common in science. Why do they occur? And what determines whether a retraction is successful? We use data from citation records and Altmetrics to test proposed answers to these questions. LaCroix et al. employ network models to argue the social spread of information helps explain failures of retraction. One prediction is that widely known results, surprisingly, should be easier to retract, since their retraction is more relevant. Our results support this conclusion. We find highly cited papers show more significant reductions in citation after retraction and garner more attention to their retractions as they occur.

cs.DL

Filling holes in science draws collective attention, but most higher-order holes remain unexplored

Much scientific discovery involves filling holes between ideas and arguments that unleash techno-scientific advance. Representing knowledge as high-dimensional concept embeddings, we use persistent homology to detect holes of increasing order, from gaps between disconnected ideas to higher-order cavities, and identify the research works that fill them. We find two empirical asymmetries. Researchers who fill anticipated holes are poised to draw collective attention by staging outsized novelty and foresight, indicating that bridging holes anticipates where science will converge, most strongly in empirical fields and least in formal and design fields. Yet as knowledge grows, higher-order holes explode while the fraction science fills collapses, leaving most higher-order combinations unexplored. These results call for a richer science of holes, and mark a frontier where contemporary AI might help fill the high-dimensional gaps human science opens.

cs.CY

Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature

Tracing connections between historical texts is an important part of intertextual research, enabling scholars to reconstruct the virtual library of a writer and identify the sources influencing their creative process. These intertextual links manifest in diverse forms, ranging from direct verbatim quotations to subtle allusions and paraphrases disguised by morphological variation. Language models offer a promising path forward due to their capability of capturing semantic similarity beyond lexical overlap. However, the development of new methods for this task is held back by the scarcity of standardized benchmarks and easy-to-use datasets. We address this gap by introducing Loci Similes, a benchmark for Latin intertextuality detection comprising a curated dataset of ~176k text segments and 1,490 expert-verified parallels, including 945 labeled references from an existing dataset. Using this data, we establish baselines for retrieval and classification of intertextualities with both lexical methods and pretrained encoder language models.

cs.IR

China's Shrinking Home Bias and Rising Disruptive Impact: Evidence from a Global Citation Network Analysis

China has become the world's largest producer of scientific publications, yet concerns persist that this growth is inflated by excessive domestic citation practices. In this study, we analyze a citation network of over 45 million publications from Web of Science (1980-2025) to investigate China's home citation bias and research impact. Using a network reshuffling null model to control for the structural effect of publication volume, we find that China's home citation bias is less pronounced than commonly assumed and has been steadily declining over the past two decades. Chinese researchers do not exhibit a significantly stronger home citation preference than other major countries, indicating increasing internationalization rather than insularity. Furthermore, using the persistent disruption framework, we show that Chinese papers are converging toward American papers in their capacity to produce paradigm-shifting work. These findings challenge prevailing narratives about Chinese scientific home bias and suggest that China's advances in research impact.

physics.soc-ph

MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relatively short texts. We introduce MultiGhostBench, a multilingual benchmark comprising 928 books generated by five recent LLMs across six languages and three scripts, with an average length of approximately 59K words per book. The benchmark supports evaluation under domain, author, and language shifts. Evaluation of representative AA methods shows that no single method consistently performs best across settings, and performance generally degrades under distribution shifts. Transformer-based detectors can retain generator-related information across languages, although transfer effectiveness varies by language pair, whereas statistical and fingerprint-based detectors are more language-dependent. We envision MultiGhostBench as a valuable resource for the development and evaluation of robust AA methods. The dataset and code can be found at https://github.com/GrecoMT/MultiGhostBench.

cs.CL

The zbMATH Open Knowledge Graph: Tracing Centuries of Mathematical Research

We present the zbMATH Open Knowledge Graph, a large-scale RDF knowledge graph (KG) covering more than 250 years of mathematical scholarship. Unlike existing scholarly knowledge graphs that primarily capture bibliographic metadata and citation structures, the zbMATH Open KG integrates expert-curated semantic content, including reviews, keywords, subject classifications, software references, and disambiguated authorship. This combination of domain-specific representation of mathematical knowledge and extensive temporal coverage supports analyses that require fine-grained exploration of mathematical concepts, research fields, and scholarly relationships over time. The resulting graph comprises 34 million entities and 168 million RDF triples represented using established Semantic Web vocabularies, supporting interoperability and FAIR data principles. We further demonstrate its capabilities through query-driven historically grounded scholarly exploration use cases, illustrating how the knowledge graph can surface relationships and patterns that may be difficult to identify from bibliographic and citation information alone. The zbMATH Open KG provides an open semantic infrastructure for studying the development of mathematical knowledge and tracing scholarly connections across centuries of scholarship.

cs.DL

A Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts

In this work, we target Handwritten Text Recognition (HTR) in low-resource scenarios, which arise from underrepresented languages, rare scripts, and degraded visual conditions typical of historical documents. We introduce SCAM (Sahidic Coptic Ancient Manuscripts), a new line-level dataset built from digitized ancient manuscripts written in the extinct Sahidic Coptic dialect. The dataset reflects a realistic and challenging setting, as it combines heterogeneous acquisition conditions across libraries with typical manuscript degradations such as ink fading, bleed-through, and material deterioration. In addition to visual complexity, SCAM poses significant linguistic challenges due to the scarcity of resources for Sahidic Coptic, its uncommon alphabet, and dialect-specific diacritics. To support research in low-resource HTR, we benchmark several state-of-the-art approaches based on different paradigms, highlighting their limitations and strengths in this setting. Our results underline the gap between current HTR performance on well-resourced modern scripts and historically grounded, low-resource scenarios, thus providing a reference point for future developments.

cs.CV

Sustaining Dryad: Reconfiguring relationships and 'thinking like a business' to maintain open data infrastructures

Open data infrastructures underpin how open research is practiced across communities. While critical, the longevity of such infrastructures is far from certain, as they grapple with challenges related to technologies, dominant ideologies, and continued funding streams. This paper draws on a mixed-methods study of Dryad, a well-known open data infrastructure, to deeply explore one of these challenges: the quest for stable funding. Our analysis shows how Dryad has carefully reconfigured different forms of relationships and revenue models throughout its history to work towards financial sustainability. We identify four types of relationships with customers, collaborators, and competitors that have been critical to Dryad's financial evolution: reinforcing, forging, positioning, and excluding relationships. We argue that while implementing strategic, business thinking is a critical strategy, it is also one which shapes other factors important in sustainability: interpretations of value(s), community, and governance. We conclude by highlighting emerging tensions that provide insight for other open data infrastructures working to become financially sustainable. As a whole, our analysis focuses not just on financial mechanisms for funding open data infrastructures (although those emerge) but on the relationships which enable them.

cs.DL

LLMs Interpret, Embeddings Organize, Graphs Emerge: Agent-Driven Compilation of Scientific Knowledge

Sustained scientific work requires a knowledge substrate that carries interpretation across tasks and preserves paths to source evidence. We call this process \emph{scientific knowledge compilation} and implement it in ASKS, the \emph{Agent-Driven Scientific Knowledge System}. For each source, an LLM produces a readable Wiki view and machine-facing semantics. Deterministic checks convert the latter into a document-local GraphDelta, and embedding geometry together with explicit graph rules integrates the proposed changes into persistent state. Each ingest is an inspectable state transition over accumulated knowledge, with compiled Wiki and graph views linked to the preserved source record. We examine this process by chronologically compiling 56 published papers from one research program. Branch survival, cross-paper support, lineage, coverage, and churn yield a source-traceable author research portrait centered on tensor-network methods, with branches into quantum many-body research, tensor-network machine learning, and quantum-AI-oriented directions. In this run, higher-level Hub organization remains stable and low-churn. Canonical-node growth is predominantly additive. Graph-level measurements and navigation paths retain links to the source records from which they were compiled.

cs.AI

Guiding LLM Peer Reviewers: The Impact of Score Anchors on Review Evidence and Accuracy

Large language models (LLMs) are increasingly used for research quality evaluation, with prior work exploring their scoring accuracy and the plausibility of review rationales. However, less is known about whether external score guidance changes the evidence presented in the generated review as well as the final score. This study uses 98 Allied Health Professions research outputs submitted for internal REF-style assessment, with specialist human review reports and adjudicated 1-4 reference scores. No-guidance baseline reviews are compared with oracle-guided reviews, where the supplied score is set to the rounded human reference score; extracted evaluation points are used to compare human and LLM evidence use. Using this design, oracle guidance improves scoring accuracy, with score-following checks showing that models do not simply copy the supplied score. Corrected score mismatches are associated with changes in the generated review frame, showing that the score signal can steer review rationales. This effect is direction-dependent: LLM reviews cover human strength or upgrade points more reliably than human weakness or downgrade points, with the weakest alignment for expert downgrade evidence. The results show that score-guided review generation can be evaluated at the level of review evidence, as well as the final score.

cs.DL

Tracing high-profile attention to questionable research as a case for funder due diligence

When questionable research has entered the scientific record, it stands the same chance as other literature of shaping national and international policy, clinical guidance, and research and development activities. This study uses a known authorship-for-sale network identified in August 2022 to examine whether questionable research influences policy, patents, and clinical guidelines, and whether its authors continue to secure funding and publish after exposure. Across nearly 2,000 publications in our dataset, 57 were cited in policy documents, 12 in clinical guidelines, and 480 in patents. Nearly a quarter of these publications are linked to one or more grants: funding that could otherwise have supported rigorous, ethical science. Among the 278 authors we traced, publishing and funding continued well after the PNN's public exposure in 2022: over 90% continued publishing, and 23% were linked to a grant. As our findings show, the participants of paper mills and authorship-for-sale networks can still inform policy and clinical practice and support their authors' career progression even after the questionable practices behind them are exposed. We present a case for funders to consider authorship and network structure as part of a holistic due diligence process, alongside metrics such as citation counts and attention data.

cs.DL

Advisor career stage and PhD advisee outcomes

PhD advisors are central to doctoral training, but their influence may vary across career stages. Early-, mid-, and late-career advisors may differ in research activity, mentoring capacity, professional networks and access to resources. However, little is known about how PhD advisor career stage is associated with PhD student development outcomes. Drawing on multiple large-scale datasets comprising 250,838 advisor-advisee pairs from 312 U.S. PhD-granting institutions, we examine the relationship between advisor career stage and PhD advisee outcomes in knowledge production, collaboration networks and academic career placement. We find that early-career PhD advisors are associated with advisees' higher research productivity and citation performance, more opportunities to engage in direct and intensive research collaboration, and greater likelihood of securing a faculty position. Mid- and late-career faculty, by contrast, appear to have advantages in providing network capital which students can inherit after graduation, training PhD advisees to produce disruptive research, and supporting them in securing faculty positions at top institutions. This study contributes to a more comprehensive understanding of the reproduction of scientific talent by revealing the role of advisor career stage in shaping this process. These findings have implications for doctoral applicants' decision-making and for institutional policymaking on PhD training, faculty support and faculty evaluation.

cs.CY

Beyond Human-Likeness: Mapping the Scientific Critique Profiles of LLMs and Human Reviewers

Large language models (LLMs) are increasingly discussed as tools for peer review, but their value is often assessed through human-likeness, perceived usefulness, or textual overlap with reviewer comments. This study shifts attention from whether LLMs resemble human reviewers to what functions of scientific critique they perform. Using ICLR 2025 peer-review data, we compare human reviews with LLM reviews generated under baseline and expert prompts. We operationalize scientific critique through two review acts, weakness critique and scientific questioning, and annotate point-level review text using five theory-guided frameworks: Anderson's knowledge types, Toulmin's argumentation model, Graesser's question depth, SOLO cognitive complexity, and Hattie's feedback functions. The results reveal a differentiated critique profile. Human reviews placed greater emphasis on scientific framing and revision guidance, more often identifying higher-order weaknesses and asking questions oriented toward improvement. LLM reviews showed higher rates of explanatory depth, integrative reasoning, and explicit argument structuring. Expert prompting did not make LLM critique uniformly more human-like; it partially narrowed some gaps but mainly amplified LLM-specific tendencies toward integration and formal argumentation. These findings show that LLM-assisted peer review changes the functional composition of review text, making it important to distinguish LLM-amplified critique from areas requiring human prioritization and accountable judgement.

cs.DL

Knowledge Distillation for mmWave Beam Prediction Using Sub-6 GHz Channels

Beamforming in millimeter-wave (mmWave) high-mobility environments typically incurs substantial training overhead. While prior studies suggest that sub-6 GHz channels can be exploited to predict optimal mmWave beams, existing methods depend on large deep learning (DL) models with prohibitive computational and memory requirements. In this paper, we propose a computationally efficient framework for sub-6 GHz channel-mmWave beam mapping based on the knowledge distillation (KD) technique. We develop two compact student DL architectures based on individual and relational distillation strategies, which retain only a few hidden layers yet closely mimic the performance of large teacher DL models. Extensive simulations demonstrate that the proposed student models achieve the teacher's beam prediction accuracy and spectral efficiency while reducing trainable parameters and computational complexity by 99%.

eess.SP

Visible but Not Yet Curatable: Characterizing the Curatability of Compact and Derived Open LLM Artifacts

Open Large Language Model (LLM) research increasingly produces compact and derived artifacts, such as adapters, quantized checkpoints, merged models, and distilled variants, that are distributed across papers, model hubs, model cards, code repositories, and release statements. Although these artifacts are publicly visible, digital libraries often lack sufficient evidence to identify, preserve, and cite them as coherent scholarly objects. We introduce a framework that conceptualizes curatability as a record-level property of distributed scholarly records and operationalizes it through four evidence dimensions: artifact identity, scholarly linkage, upstream evidence, and release assets. Guided by this framework, we conduct the first collection-scale characterization of open LLM curatability using a May 2026 snapshot of 191,375 public Hugging Face repositories and a core corpus of 2,214 scholarly papers. Our results reveal a pronounced visibility-to-curatability funnel. While 90.7% of paper records contain at least one useful curation signal, only 18.1% combine usable upstream evidence with concrete release evidence, and only 6.1% provide sufficiently coordinated evidence to support high-curatability records. Based on these findings, we derive a minimal seven-field curatable record and complementary responsibilities for model hubs, scholarly indexes, and digital libraries, providing practical guidance for improving the preservation and bibliographic control of open LLM artifacts.

cs.HC

Lot Machine: Multimodal Lot Extraction from Auction Catalogs

For provenance research and art market studies, auction catalogs are an essential resource to trace specific objects over time and space. While historical auction catalogs follow established domain conventions, their internal formatting remains highly variable, and their large-scale analysis is currently restricted by the lack of machine-readable representations of the auction lots. We propose a pipeline to automatically extract structured lot-level metadata from German Sales, a large database of historical auction and sales catalogs from the 19th and 20th centuries. Using a manually annotated test set of representative catalog pages, we evaluate Vision-Language Models (VLMs) under varying prompt strategies and constrained decoding frameworks. To reflect the practical constraints faced by cultural heritage institutions, including budget, compute resources, and data privacy requirements, we benchmark the methods across different deployment modes ranging from commercial providers to locally hosted, quantized models. We find that commercial endpoints establish the performance ceiling, while institutional gateways offer a viable, privacy-preserving alternative. Local deployments remain feasible, but strictly require enforcing the output structure during generation to guarantee a valid JSON format. While varying degrees of human-in-the-loop correction are still necessary, this work demonstrates that a VLM-based pipeline can successfully unlock historical auction catalogs for large-scale automated analysis.

cs.CV

Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation

Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We introduce a masked-citation task to compare human and LLM-generated citation behavior. For each citation context, an LLM generates a replacement citation sentence, producing a counterfactual corpus directly comparable to human citation. We analyze what, whom, and how models cite, using an LLM-as-a-judge to classify citation intent and a 20-million-edge coauthorship network to measure social distance between cited authors. Across six popular LLMs and 1,746 top NLP conference papers (63k+ contexts, 132k+ citations), three patterns emerge: (1) Compared with human citation, LLMs cite significantly less critically; (2) LLMs over-cite popular and older papers, a tendency amplified for contrasting citations where human writing more often draws on recent, niche work; (3) Whereas humans often cite within their close social network, especially for supporting citations, LLMs tend to draw on more socially distant authors. Together, these differences are double-edged: LLM citation reaches beyond a scholar's close collaborators while being less critical and amplifying visibility bias, reshaping the rhetoric and reach of scientific citation.

cs.DL

SoniMet - A tool for sonifying and visualizing the performance of single researchers

For centuries, the scientific community has predominantly relied on visual tools to communicate empirical results and complex datasets. While visual representations dominate bibliometric analyses, the human auditory system possesses sensitive capacities for processing complex temporal information, distinguishing intricate patterns, and tracking parallel data streams. Data sonification translates data relations into acoustic signals, offering an alternative method for data exploration, pattern recognition, and scientific communication. This paper applies this concept to the field of bibliometrics through metrics sonification-the auditory translation of bibliometric information-and introduces SoniMet (Sonifying Metrics), a web-based tool designed to visualize and sonify the publication and citation data of individual scientists (see https://sonimet.kennebec.co.uk). SoniMet connects to the OpenAlex database to retrieve bibliometric records and displays them on an interactive chronological timeline. The tool employs direct parameter mapping to translate citation impact indicators into non-speech sound: field-weighted citation impact and citation counts determine the pitch and volume of a synthesized note and are mapped to the acoustic echo strength. Although SoniMet expands the methodological toolkit for research evaluation, current limitations include its restriction to individual scholar profiles and the challenge of integrating transient audio files into traditional, text-based scientific publishing workflows. Future empirical user studies are necessary to systematically evaluate the analytical utility and cognitive benefits of metrics sonification compared to established visual methods.

cs.DL