Search arXivSearch

arXiv · 2201.08443

Diversifying the Genomic Data Science Research Community

Abstract

Over the last 20 years, there has been an explosion of genomic data collected for disease association, functional analyses, and other large-scale discoveries. At the same time, there have been revolutions in cloud computing that enable computational and data science research, while making data accessible to anyone with a web browser and an internet connection. However, students at institutions with limited resources have received relatively little exposure to curricula or professional development opportunities that lead to careers in genomic data science. To broaden participation in genomics research, the scientific community needs to support students, faculty, and administrators at Underserved Institutions (UIs) including Community Colleges, Historically Black Colleges and Universities, Hispanic-Serving Institutions, and Tribal Colleges and Universities in taking advantage of these tools in local educational and research programs. We have formed the Genomic Data Science Community Network (http://www.gdscn.org/) to identify opportunities and support broadening access to cloud-enabled genomic data science. Here, we provide a summary of the priorities for faculty members at UIs, as well as administrators, funders, and R1 researchers to consider as we create a more diverse genomic data science community.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

The Genomic Data Science Community Network, Rosa Alcazar, Maria Alvarez, Rachel Arnold, Mentewab Ayalew, Lyle G. Best, Michael C. Campbell, Kamal Chowdhury, Katherine E. L. Cox, Christina Daulton, Youping Deng, Carla Easter, Karla Fuller, Shazia Tabassum Hakim, Ava M. Hoffman, Natalie Kucher, Andrew Lee, Joslynn Lee, Jeffrey T. Leek, Robert Meller, Loyda B. Méndez, Miguel P. Méndez-González, Stephen Mosher, Michele Nishiguchi, Siddharth Pratap, Tiffany Rolle, Sourav Roy, Rachel Saidi, Michael C. Schatz, Shurjo Sen, James Sniezek, Edu Suarez Martinez, Frederick Tan, Jennifer Vessio, Karriem Watson, Wendy Westbroek, Joseph Wilcox, Xianfa Xie. 2022-06-09. Diversifying the Genomic Data Science Research Community. https://doi.org/10.1101/gr.276496.121

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Flash-Radiomics: A Scalable Hybrid CPU-CUDA Engine for Standardized Scalar Radiomics and Accelerated Spatial Mapping

Background and Objectives: Spatial mapping retains the spatial distribution of radiomic features, but computational cost and fragmented software limit its use. We developed Flash-Radiomics with scalar extraction and spatial mapping, a central processing unit (CPU) backend, a hybrid Compute Unified Device Architecture (CUDA) backend, consistent feature names, and Hierarchical Data Format version 5 (HDF5) storage. Methods: We evaluated Image Biomarker Standardisation Initiative (IBSI) compliance, CPU-CUDA concordance, and end-to-end processing time. Compliance testing included 825 chapter 1 (IBSI-1) tests covering 165 high-consensus features and 323 chapter 2 (IBSI-2) tests with numerical references. Concordance testing included 1,148 scalar pairs and 93 spatial-map pairs. End-to-end processing time was measured five times per input volume of interest (VOI) size. Comparisons included the Medical Image Radiomics Processor (MIRP) and PyRadiomics for 102 shared scalar features and PyRadiomics for 93 shared spatial maps. Results: Both backends passed all 1,148 IBSI tests, and all paired results were concordant. At the largest scalar input, CPU required 76.343 s and hybrid CUDA 81.915 s; CPU was 4.7 times faster than MIRP and 190.7 times faster than PyRadiomics. At the largest spatial input completed by both backends, hybrid CUDA reduced processing time by 68.6% relative to CPU (79.280 versus 252.791 s). At PyRadiomics' largest completed spatial input, hybrid CUDA was 84.9 times faster. Conclusions: Flash-Radiomics unified standardized scalar extraction, spatial mapping, concordant CPU-CUDA results, and HDF5 storage. CPU processing time was similar or shorter for scalar extraction, whereas hybrid CUDA was faster for spatial mapping under the tested conditions.

q-bio.OT

EPI-KAN: A Method For Estimating and Forecasting Time-Dependent COVID-19 Parameters

We introduce EPI-KAN, a novel method for estimating COVID-19 time-varying parameters. EPI-KAN uses historical epidemiological data, Physics-Informed Neural Network (PINN), and the novel Kolmogorov-Arnold Network (KAN). The method harnesses the novel Kolmogorov-Arnold Network (KAN), which is a type of artificial neural network. For the KAN in this paper, we learn activation functions that are represented using Fourier series, hence we abbreviate as KAN-F. In this study, we estimate parameters in the context of an SIRD compartmental differential equations. The time-dependent parameters are the transmission rate $β(t)$, recovery rate $γ(t)$, and mortality rate $μ(t)$. We define three KAN-F functions $\widehatβ$, $\widehatγ$, $\widehatμ$ that model the true parameters $β(t)$, $γ(t)$, $μ(t)$, respectively. We test two model architectures for the KAN-F: the first has 8 input variables consisting of $S$, $I$, $R$, $D$, and their numerical gradients at any time $t$, while the second has 4 input variables excluding the numerical gradients. The objective loss function that has to be minimized is subject to Physics-Informed Neural Network (PINN). Using historical data of COVID-19 from three South-East Asian countries: Indonesia, Singapore, and Malaysia, we are able to estimate $β(t)$, $γ(t)$, and $μ(t)$ on each country with decent accuracy and efficiency. The time period of choice coincides with the period where SARS-CoV-2 Delta variant (B.1.617.2) was dominant. In addition to estimating the rates during the training period, we also predict transmission rates over 30 days during forecast period. We found that the output of KAN-F over the forecast period can give good predictions if we scale the output by a factor of 17\% for Indonesia and 30\% for Singapore and Malaysia.

q-bio.OT

Making Models That Matter: How to Build Trustworthy and Useful Systems Biology Models

Computational models supporting mechanistic understanding of (complex) biological systems, systems behaviour prediction, and experimental design are becoming more and more embedded in research on complex biological systems. Reuse and refinement of models, rather than continuous reinvention, is becoming increasingly important as models' demands on computational infrastructure increase. However published models - despite the variety of efforts taken so far - are frequently difficult to reproduce or reuse, substantially limiting their scientific value. Here we address the requirements for model reusability in the light of the field-specific CURE framework (Credible, Understandable, Reproducible, Extensible) and the more general FAIR principles (Findable, Accessible, Interoperable, Reusable). Considering published guidance we identify broad agreement on requirements for findability, accessibility, and interoperability, but continued lack of clarity and consensus around reusability. Focusing on the scientific quality and usability of computational models we discuss six key practices underpinning model sharing and re-use. Mapping the FAIR and CURE principles onto the model lifecycle we propose ten recommendations for building and sharing systems biology models that are both FAIR- and CURE-compliant.

q-bio.OT