Search arXivSearch

arXiv · 2003.07681

The Data Science Fire Next Time: Innovative strategies for mentoring in data science

Abstract

As data mining research and applications continue to expand in to a variety of fields such as medicine, finance, security, etc., the need for talented and diverse individuals is clearly felt. This is particularly the case as Big Data initiatives have taken off in the federal, private and academic sectors, providing a wealth of opportunities, nationally and internationally. The Broadening Participation in Data Mining (BPDM) workshop was created more than 7 years ago with the goal of fostering mentorship, guidance, and connections for minority and underrepresented groups in the data science and machine learning community, while also enriching technical aptitude and exposure for a group of talented students. To date it has impacted the lives of more than 330 underrepresented trainees in data science. We provide a venue to connect talented students with innovative researchers in industry, academia, professional societies, and government. Our mission is to facilitate meaningful, lasting relationships between BPDM participants to ultimately increase diversity in data mining. This most recent workshop took place at Howard University in Washington, DC in February 2019. Here we report on the mentoring strategies that we undertook at the 2019 BPDM and how those were received.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Latifa Jackson, Heriberto Acosta Maestre. 2020-03-01. The Data Science Fire Next Time: Innovative strategies for mentoring in data science. https://arxiv.org/abs/2003.07681

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Data-driven modeling in the introductory physics laboratory: Scaling analysis and data collapse in the specific heat of water experiment

In introductory physics laboratories, a central instructional goal is to help students construct and evaluate mathematical models from empirical data rather than applying given formulas. We present a data-driven redesign of the classic specific heat of water experiment that emphasizes scaling analysis and data collapse as tools for model construction. The activity combines structured experimental and analytical guidance with instructor-mediated questioning, while thermodynamic theory is deliberately postponed. Students collect temperature-time data under various experimental conditions, producing multiple data sets that initially appear unrelated. Through successive rescaling, students reduce the dimensionality of the variable space and achieve data collapse onto a single master curve, from which they formulate an empirical model relating energy input, mass, and temperature change. The analysis highlights a limitation of multiplicative scaling: the additive contribution of the calorimeter cannot be eliminated, leading to a structural non-identifiability of the subsystem contributions. To clarify the domain of validity of the model, a thermodynamic description is introduced a posteriori as a boundary-setting framework for interpreting the empirical model. In this sense, the central contribution of this work is to use data-driven modeling both to construct models and to reveal their intrinsic limitations. The experiment provides an accessible example of how scaling, data collapse, and theoretical reasoning can be integrated in an introductory laboratory.

physics.ed-ph

Missing Data on Physics Exams: Demographic Patterns, Course-Level Predictions, and Implications for Equity

In a previous quantitative retrospective study we showed that different demographic groups of students leave different numbers of problems blank on physics exams, leading to inequities in course outcomes. In that work we argued that there were good reasons to treat these blanks as missing data, rather than indicators of a lack of understanding. In this paper, we refine this analysis and show more detailed breakdowns of uncollected test item responses by race/ethnicity and first generation college student status, coming to the same conclusion: test item responses are uncollected for students with different ethnic and racial backgrounds at different rates, and these patterns are not exclusive to low-performing students. We also correct an error from our previous work, finding here that there is no significant gender difference in uncollected test item responses. Finally, we provide a more robust analysis of course level data illustrating that blanks are a variable controlled at the course level rather than the student level, providing more evidence for the use of a course deficit model (rather than a student deficit model) when examining equity disparities, and also suggesting that there are plausible means for instructors to minimize uncollected test item responses, and therefore reduce or even eliminate the bias associated with this missing data. We provide a couple suggestions for faculty who want to minimize the impacts of blanks.

physics.ed-ph

A Framework for Characterizing Learning Contributions Across the Initial Achievement Spectrum

Conceptual assessments are widely used in physics education research to evaluate changes in student understanding, yet class average measures can obscure how those changes are distributed across students with different levels of initial achievement. We introduce a Learning Contribution Framework that characterizes this distribution through the Learning Contribution Curve (LCC) and Learning Contribution Profile (LCP). For a general contribution measure G, the LCC represents cumulative contribution across students ranked by initial achievement, whereas the LCP describes the local mean contribution relative to the population mean of the individual G values. We apply the framework to the individual Hake normalized gain and develop a statistical model linking LCC and LCP behavior to the joint structure of pretest and posttest scores. Under the central assumption that the conditional mean of the subsequent score is linear in the initial score, we identify a score structure parameter $β$ and a critical score structure parameter $β_c$. Whether $β$ is greater than, less than, or equal to $β_c$ determines whether the expected LCP increases, decreases, or remains constant across initial achievement. Given the pretest distribution, the model further yields analytical predictions for the complete LCP and LCC that closely reproduce simulation results. Application to classroom concept-assessment data illustrates how the framework reveals local and cumulative patterns of learning contribution that are not evident from an overall class average revealed by the traditional Hake's normalized gain.

physics.ed-ph