Search arXivSearch

arXiv · 2607.09409

Two Vocabularies, One Phenomenon: Metadata Bias in AI Evidence Synthesis on Fertility Decline

Abstract

Declining fertility is one of the defining policy questions of the next decade, and increasingly, what policymakers know about it is shaped by AI-synthesising the evidence base. But ask such a tool about reproduction and the answer depends on the word you use. The same phenomenon, framed clinically (e.g, infertility, IVF) or socially (e.g., childlessness, fertility intentions), is catalogued with radically different completeness. And the catalogue as much, if not more than the underlying scholarship, is what AI synthesis begins with (Bolaños et al., 2024). That databases under-index the social sciences, books, and grey literature is well established (Visser et al., 2021). What is new here is holding the topic fixed and asking whether metadata gaps act as a hidden policy filter on a single contested issue: the determinants of (in)fertility. We use two OpenAlex queries on the same phenomenon: a clinical basket (infertility, subfertility, ART, IVF, fecundity; n=101,645) and a social basket (childlessness, social infertility, fertility intentions, reproductive decision-making; n=3,646). We compare them on metadata completeness, open access, output type, and institutional provenance. The social framing is consistently less machine-legible: output skewed to books and dissertations, authorship university -- rather than healthcare-based. Open access rates are essentially equal (43.1% vs 41.3%), so the gap is in indexing depth, not paywalls, suggesting simple OA mandates will not fix it. On this same mixed literature, even before any coverage bias enters the picture, LLM tools already miss more than they catch when asked to extract hypotheses and claims (Uprety et al., 2025); the bias documented here compounds an already-imperfect extraction stage.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Boyana Buyuklieva, Veronica-Nicolle Hera. 2026-07-10. Two Vocabularies, One Phenomenon: Metadata Bias in AI Evidence Synthesis on Fertility Decline. https://arxiv.org/abs/2607.09409

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The illusion of neutrality in metric-based research evaluation

Venue prestige and citation counts are two widely used, albeit imperfect, signals of research quality. When the two signals conflict, evaluators must decide how much weight to assign each. Yet, it remains unknown how researchers across disciplines trade off these two signals when evaluating research outcome. To fill this gap, we surveyed 869 researchers using paired choices between hypothetical departmental hiring rules that assigned different weights to venue prestige and citation counts, asking which would produce better science. Among 795 respondents whose choices were largely internally consistent, choices placed nearly equal aggregate weight on the two signals. This apparent balance concealed substantial individual heterogeneity: nearly two in five respondents occupied the most venue-heavy or citation-heavy intervals. Moreover, respondents on average chose more citation-heavy rules than they believed their departments used in hiring. By revealing the subjective judgments that arise when research indicators conflict, our findings reinforce calls for greater caution when quantitative indicators are used to evaluate research.

cs.DL

Scoring Grant Applications with Large Language Models

Purpose: Assessing grant applications is time-consuming and difficult, adding to the overall burden of academic peer review. Whilst funders are exploring whether AI can help, there is no published research into the accuracy of Large Language Models (LLMs) for scoring contemporary grants. Design/methodology/approach: This study investigates whether six open-weight LLMs (Gemma 3 1B/4B/12B/27B, DeepSeek R1 32B, Qwen 3 32B) can give useful scores for 2267 recent UK Economic and Social Research Council (ESRC), and Engineering and Physical Sciences Research Council (EPSRC) grant applications, comparing them with scores from the original reviewers and funding panel members. Findings: Although the LLM scores are individually inaccurate, when averaged and converted to ranks they correlate positively with expert average scores. The best performing LLM, Gemma 3 27B (10 iterations with varied prompts), had moderate rank correlations with average reviewer scores (mean rho=0.26). Gemma 3 27B's average correlation with individual reviewers was 0.19, which is lower than the inter-reviewer mean correlation of 0.24, suggesting that it scores are slightly weaker than individual reviewer scores. Gemma 3 27B had weak rank correlations with average panel member scores (mean rho=0.17), with lower average correlations with individual panellists (mean rho=0.14), which is substantially lower than the inter-panellist correlation (mean rho=0.38). Whilst the correlations seem too weak to replace expert review at the final panel stage, LLM scores might help with the initial reviewing state, such as by helping identify the weakest proposals for fast-track desk rejections, to replace one human reviewer, or for triangulation to check for bias.

cs.DL

Open Science, Closed Models: How Funding Shapes AI in Science

How funding shapes AI engagement in science is poorly understood despite its structural importance. We analyze 104,226 scientific papers (2018-2025), linking funding acknowledgments to how each paper engages with foundation models: whether it extends a model (fine-tunes or builds on it), uses one without modification, or references models only peripherally. Three findings emerge. First, funding source is associated with the character of engagement: public funding is associated with higher open-weight model engage- ment; private-only funding selectively enables extension with no robust openness association; mixed funding shows nominally the highest open-weight engagement rates in four of the seven largest disciplines. Second, papers acknowledging industry cloud credits are less likely to use open-weight models, consistent with credit programs steering research toward closed models. Third, industry collaboration carries an independent extension premium consistent with internal corporate resources flowing through coauthorship channels invisible to acknowledgment-based measures. This nexus concentrates in Computer Science and Global North collaborations; Global South research is largely excluded. As public funding contracts and industry compute provision expands, scientific AI work shifts toward closed proprietary infrastructure.

cs.DL