Search arXiv⌕ Search

arXiv · 2609.28946

Digital Medicines Information at National Scale: Search Behaviour and System Performance of MediVerify Across 1.5 Million Queries

Abstract

National-scale digital health platforms generate usage data that reveal population healthcare needs. MediVerify, Sri Lanka's online medicines information platform providing access to National Medicines Regulatory Authority (NMRA)-approved medicines, recorded over 1.49 million queries in its first year. Large-scale medicine search behaviour in low- and middle-income countries (LMICs) remains poorly characterised. Objectives: To characterise medicine information-seeking behaviour, identify mismatches between public demand and essential medicines policy, and evaluate technical performance. Methods: We retrospectively analysed 1,497,304 anonymised queries submitted between July 2024-2025. Queries were normalised and mapped to 11,933 approved medicines using fuzzy matching based on Levenshtein distance (threshold <=5). Therapeutic categories were assigned using an Anatomical Therapeutic Chemical (ATC)-aligned classification. Query patterns and system performance were analysed, and energy consumption was estimated from processing times. Results: Vitamins/minerals (11.74%), antibiotics (10.57%), anti-diabetes (7.02%), and antihypertensives (6.84%) dominated searches. The top 20 accounted for 37.3% of queries, while approximately 40%-50% of registered medicines were never queried. Zero-result queries (~ 1%) indicated unmet information needs. Frequently searched medicines diverged from essential medicines lists. Median latency was 8ms, with >99% query resolution. Estimated energy consumption was approximately 0.12 kWh per million queries. Conclusions: Large-scale medicine search data provide insights into healthcare information demand and policy alignment. MediVerify demonstrates that a national digital health platform can achieve high utilisation at low computational and environmental cost. Usage analytics could strengthen pharmaceutical policy and digital health infrastructure in LMICs.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Praveen Charuka Athauda-Arachchi, Pandula Mahesh Athauda-Arachchi, Rohini Fernandopulle. 2026-09-24. Digital Medicines Information at National Scale: Search Behaviour and System Performance of MediVerify Across 1.5 Million Queries. https://doi.org/10.3389/fdgth.2026.1855835

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Cross-Section of Stock Returns and AI Exposure

We study 380 trillion tokens of realized AI consumption across more than four hundred LLMs. We build a high-frequency AI factor and show that a long-short strategy based on firms' AI exposure earns significantly positive returns. The average strategy return is larger based on intensive, frontier-oriented AI consumption but smaller based on casual or open-weight usage. Internationally, the return spread is significant in developed countries but insignificant in emerging markets. Examining occupational AI exposure, we find more positive exposure in occupations intensive in nonroutine interactive tasks and more negative exposure in those intensive in nonroutine analytical tasks.

cs.CY↗

What fidelity metrics miss: a structural check on synthetic educational data

Secondary use of educational records is increasingly mediated by platforms that share a differentially private synthetic version of a dataset and validate specific findings against the real data on request. The synthetic version is evaluated by comparing summary statistics of each variable, yet reported confirmation rates suggest that such comparisons do not predict which findings survive. We propose a structural check: the number of connected components of a weekly proximity graph over learners, tracked across a term. Across four annual cohorts of lower-secondary study-habit logs, the synthetic versions reproduced the level of this quantity and the shape of the weekly partition, but its variation across the term was between 2.6 and 4.9 times smaller than in the real data at a common working point, without exception, and those changes fell in different weeks: the synthetic cohorts single out the term's examination weeks and the real cohorts do not. We also show that a routine rule for setting the graph threshold makes naive comparisons between two datasets invalid, and illustrate this with an error of our own. The real curves are also distinguishable from marginal-preserving surrogates of themselves in all four cohorts, where three of the four synthetic ones are not, a comparison that needs no real data; these differences trace to what the generator was given.

cs.CY↗

AI-Moderated Interviews for Market Research and Digital Twins Calibration

AI-moderated interviews are emerging as a scalable market-research method for generating consumer insights and building consumer "digital twins." Yet it remains unclear whether they match human-moderated interviews or improve on simpler, static data collection methods. In a pre-registered, between-subjects study (N = 317) with three industry partners, we compare AI-moderated (N = 139), human-moderated (N = 24), and static interviews (N = 154). AI moderation matches human moderation in depth, covers more themes, and, holding budget constant, recovers significantly more customer needs than human moderation or static interviews. However, participants sound more emotionally engaged when speaking to a live human. We then create digital twins using interview data and evaluate each twin against the participant's own held-out responses to six real-world marketing stimuli. We find that digital twins created from AI-moderated interviews predict consumer responses better than demographics-only personas. However, the additional richness from AI moderation does not translate into better quantitative predictions compared to static interviews. By analyzing open-ended thoughts generated from humans versus their twins, we find that prediction errors are connected both to differences in (self-reported) thinking styles between twins and humans, and to gaps between training and validation data (i.e., asking questions that are too far out of distribution).

cs.CY↗