Search arXivSearch

arXiv · 2412.10019

Performance of ChatGPT on tasks involving physics visual representations: the case of the Brief Electricity and Magnetism Assessment

Abstract

Artificial intelligence-based chatbots are increasingly influencing physics education due to their ability to interpret and respond to textual and visual inputs. This study evaluates the performance of two large multimodal model-based chatbots, ChatGPT-4 and ChatGPT-4o on the Brief Electricity and Magnetism Assessment (BEMA), a conceptual physics inventory rich in visual representations such as vector fields, circuit diagrams, and graphs. Quantitative analysis shows that ChatGPT-4o outperforms both ChatGPT-4 and a large sample of university students, and demonstrates improvements in ChatGPT-4o's vision interpretation ability over its predecessor ChatGPT-4. However, qualitative analysis of ChatGPT-4o's responses reveals persistent challenges. We identified three types of difficulties in the chatbot's responses to tasks on BEMA: (1) difficulties with visual interpretation, (2) difficulties in providing correct physics laws or rules, and (3) difficulties with spatial coordination and application of physics representations. Spatial reasoning tasks, particularly those requiring the use of the right-hand rule, proved especially problematic. These findings highlight that the most broadly used large multimodal model-based chatbot, ChatGPT-4o, still exhibits significant difficulties in engaging with physics tasks involving visual representations. While the chatbot shows potential for educational applications, including personalized tutoring and accessibility support for students who are blind or have low vision, its limitations necessitate caution. On the other hand, our findings can also be leveraged to design assessments that are difficult for chatbots to solve.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Giulia Polverini, Jakob Melin, Elias Onerud, Bor Gregorcic. 2025-05-15. Performance of ChatGPT on tasks involving physics visual representations: the case of the Brief Electricity and Magnetism Assessment. https://doi.org/10.1103/physrevphyseducres.21.010154

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Force Concept Inventory Across Continents: Testing Q-Matrix Transferability and Cross-Cultural Differences in Mechanics Reasoning

The Force Concept Inventory (FCI) is one of the most widely used research-based assessments in physics education, yet the assumption that its underlying cognitive structure is transferable across educational contexts remains largely untested. This study investigates the transferability of FCI Q-matrices using the Generalized Deterministic Inputs, Noisy "And" Gate (G-DINA) cognitive diagnostic applied to two large cohorts: students from the Learning About STEM Student Outcomes (LASSO) online system database in the United States (N = 4,750) and introductory physics students at the University of Johannesburg, South Africa (N = 1,016). Rather than treating the analysis as a local model calibration exercise, we frame the problem as one of cross-context cognitive invariance. Differential Item Functioning (DIF) analyses revealed substantial cross-cultural differences, with 14 of 30 items exhibiting high DIF after controlling for latent skill mastery. These differences were concentrated in force dynamics and contact-force reasoning and remained invariant under alternative Q-matrix specifications. The findings suggest that observed differences reflect genuine variations in students' conceptual reasoning rather than psychometric artifacts, highlighting the importance of validating Q-matrix structures before deploying cognitive diagnostic and adaptive assessments across diverse non-local educational settings.

physics.ed-ph

Inference uncertainty about an aircraft crash

Problem-based learning benefits from situations taken from real life, which usually stimulate student interest. In this paper, we examine the shooting down of the Rwandan president's aircraft on April 6th, 1994. We discuss methods to infer information about the location from which the missile was launched, its trajectory and type, where the aircraft was struck and its trajectory during the fall. To this end, we developed a physics-based analysis based on expert reports, witness statements, and other publicly available information, as interpreted by our calculations. The analysis is designed to be understandable using undergraduate-level physics. The uncertainty of each result is discussed and propagated to ensure a proper assessment of the hypotheses and a traceability of their consequences. Such approach encourages the students to exercise their critical mind and teaches inference methods that are routinely used in physics research. In addition, it illustrates the importance and limits of scientific expertise during a judiciary process.

physics.ed-ph

Skepticism vs. Convenience: Physics Students' Perceptions and Use of Large Language Models Before and After Instruction

The recent emergence of large language models (LLMs) has produced research focusing on the ability of these tools to solve physics problems, evaluate student work, or otherwise impact the problem-solving process of students. However, studies exploring how physics students perceive LLMs (in terms of capabilities, educational impacts, usage, and role in problem solving) remain limited. This study evaluates the first-year physics students' perceptions toward LLMs and further explores how these perceptions change after practicing problem-solving with and without LLMs and engaging in a reflective lesson on the functioning and educational impacts of LLMs. We find that student opinions toward LLMs vary, with generally favorable perceptions of their capabilities but greater skepticism regarding their value for learning. Despite this skepticism, a majority of students self-report regularly using LLMs to obtain help, commonly reporting deadlines and convenience as motivating factors. Following the lesson, students expressed greater skepticism toward LLMs in several areas, with 88% of students believing that LLMs can leave them with a false sense of confidence about their understanding, up from 58% before the lesson.

physics.ed-ph