Search arXivSearch

arXiv · 2403.05458

Evaluating AI and Human Authorship Quality in Academic Writing through Physics Essays

Abstract

This study evaluates $n = 300$ short-form physics essay submissions, equally divided between student work submitted before the introduction of ChatGPT and those generated by OpenAI's GPT-4. In blinded evaluations conducted by five independent markers who were unaware of the origin of the essays, we observed no statistically significant differences in scores between essays authored by humans and those produced by AI (p-value $= 0.107$, $α$ = 0.05). Additionally, when the markers subsequently attempted to identify the authorship of the essays on a 4-point Likert scale - from `Definitely AI' to `Definitely Human' - their performance was only marginally better than random chance. This outcome not only underscores the convergence of AI and human authorship quality but also highlights the difficulty of discerning AI-generated content solely through human judgment. Furthermore, the effectiveness of five commercially available software tools for identifying essay authorship was evaluated. Among these, ZeroGPT was the most accurate, achieving a 98% accuracy rate and a precision score of 1.0 when its classifications were reduced to binary outcomes. This result is a source of potential optimism for maintaining assessment integrity. Finally, we propose that texts with $\leq 50\%$ AI-generated content should be considered the upper limit for classification as human-authored, a boundary inclusive of a future with ubiquitous AI assistance whilst also respecting human-authorship.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Will Yeadon, Elise Agra, Oto-obong Inyang, Paul Mackay, Arin Mizouri. 2024-03-08. Evaluating AI and Human Authorship Quality in Academic Writing through Physics Essays. https://arxiv.org/abs/2403.05458

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Force Concept Inventory Across Continents: Testing Q-Matrix Transferability and Cross-Cultural Differences in Mechanics Reasoning

The Force Concept Inventory (FCI) is one of the most widely used research-based assessments in physics education, yet the assumption that its underlying cognitive structure is transferable across educational contexts remains largely untested. This study investigates the transferability of FCI Q-matrices using the Generalized Deterministic Inputs, Noisy "And" Gate (G-DINA) cognitive diagnostic applied to two large cohorts: students from the Learning About STEM Student Outcomes (LASSO) online system database in the United States (N = 4,750) and introductory physics students at the University of Johannesburg, South Africa (N = 1,016). Rather than treating the analysis as a local model calibration exercise, we frame the problem as one of cross-context cognitive invariance. Differential Item Functioning (DIF) analyses revealed substantial cross-cultural differences, with 14 of 30 items exhibiting high DIF after controlling for latent skill mastery. These differences were concentrated in force dynamics and contact-force reasoning and remained invariant under alternative Q-matrix specifications. The findings suggest that observed differences reflect genuine variations in students' conceptual reasoning rather than psychometric artifacts, highlighting the importance of validating Q-matrix structures before deploying cognitive diagnostic and adaptive assessments across diverse non-local educational settings.

physics.ed-ph

Inference uncertainty about an aircraft crash

Problem-based learning benefits from situations taken from real life, which usually stimulate student interest. In this paper, we examine the shooting down of the Rwandan president's aircraft on April 6th, 1994. We discuss methods to infer information about the location from which the missile was launched, its trajectory and type, where the aircraft was struck and its trajectory during the fall. To this end, we developed a physics-based analysis based on expert reports, witness statements, and other publicly available information, as interpreted by our calculations. The analysis is designed to be understandable using undergraduate-level physics. The uncertainty of each result is discussed and propagated to ensure a proper assessment of the hypotheses and a traceability of their consequences. Such approach encourages the students to exercise their critical mind and teaches inference methods that are routinely used in physics research. In addition, it illustrates the importance and limits of scientific expertise during a judiciary process.

physics.ed-ph

Skepticism vs. Convenience: Physics Students' Perceptions and Use of Large Language Models Before and After Instruction

The recent emergence of large language models (LLMs) has produced research focusing on the ability of these tools to solve physics problems, evaluate student work, or otherwise impact the problem-solving process of students. However, studies exploring how physics students perceive LLMs (in terms of capabilities, educational impacts, usage, and role in problem solving) remain limited. This study evaluates the first-year physics students' perceptions toward LLMs and further explores how these perceptions change after practicing problem-solving with and without LLMs and engaging in a reflective lesson on the functioning and educational impacts of LLMs. We find that student opinions toward LLMs vary, with generally favorable perceptions of their capabilities but greater skepticism regarding their value for learning. Despite this skepticism, a majority of students self-report regularly using LLMs to obtain help, commonly reporting deadlines and convenience as motivating factors. Following the lesson, students expressed greater skepticism toward LLMs in several areas, with 88% of students believing that LLMs can leave them with a false sense of confidence about their understanding, up from 58% before the lesson.

physics.ed-ph