LLM-Based Educational Simulation: Evaluating Temporal Student Persona Stability Across ADHD Profiles
Large language model (LLM)-based student simulation offers a scalable alternative for educational research, teacher training, and learner practice. However, its validity depends on whether LLMs maintain stable personas across and within interactions. We test this using a dual-assessment framework measuring self-reported characteristics and observer-rated behavioral expressions. Across three ICD-11-informed ADHD-related intensity conditions and a default condition, five LLMs, and three prompt designs, we quantify between-conversation (Exp. I; N=$4,962$) and within-conversation stability (Exp. II; N=$3,952$). Experiment I shows that self-reports and observer ratings are more stable at high than moderate intensities across the tested models, prompts, and independent runs. Experiment II shows that self-reports remain stable throughout extended interactions, but observer-rated behavior drifts in unscripted dialog for high- and moderate-intensity personas. Scripted interactions with recurring task-relevant prompts eliminate this drift almost entirely (up to 97\% reduction). Structured interaction design helps simulated learners maintain behavioral stability in sustained, path-dependent interactions.