Search arXiv⌕ Search

arXiv subjects

Daeun Hwangbo

Publications and source records attributed to Daeun Hwangbo.

3 recordsLinked to original sources

Modeling Transition Dynamics and Network Structure in Cross-National Process Data: A Hierarchical Multi-State Survival Framework

Process data from computer-based assessments record the sequence and timing of actions through which respondents solve a task, providing information about both the pace and structure of problem-solving behavior. Modeling such processes across countries is challenging because country-by-response-group cells are often small and unbalanced and the observed transition supports can differ substantially across countries. We propose a hierarchical framework that integrates a Bayesian multi-state survival model with a network-based representation of transition structure. Partial pooling across countries yields country-specific covariate and key-action effects, transition speed, and estimates of between-country heterogeneity. Posterior transition probability networks are embedded in a common latent space using a directed graph auto-encoder adapted to heterogeneous supports, and 1-Wasserstein distances between node-role distributions are evaluated across posterior draws to characterize global network structure while propagating estimation uncertainty. We apply the framework to two problem-solving items from the Programme for the International Assessment of Adult Competencies across 14 countries. The results reveal cross-country heterogeneity in transition speed, systematic response-group differences in global network organization, item-dependent variation in within-group dispersion across countries, and local differences in routing around shared intermediate actions.

stat.AP↗

A Representation-Learning Item Response Model for Identifying Behaviorally Important Actions in PIAAC Process Data

Problem-solving log process data from computer-based assessments provide detailed information about how respondents approach and complete tasks. However, the resulting action sequences are complex and noisy, making it difficult to identify specific behaviors associated with successful performance. This paper proposes a representation-learning item response modeling (IRT) framework for identifying behaviorally important actions while accounting for respondent proficiency and item-level differences. Raw log sequences and timing information are first transformed into action representations that incorporate the hierarchical structure of action labels and the sequential and temporal context in which each action occurs. These respondent-specific representations are then entered as covariates in an extended IRT model, with spike-and-slab priors used to identify action-item combinations associated with response accuracy. The framework therefore evaluates actions contextually rather than as simple occurrence indicators and provides posterior uncertainty for their associations with performance. We apply the approach to problem-solving process data from the OECD Programme for the International Assessment of Adult Competencies (PIAAC). The analysis identifies a sparse set of actions associated with successful and unsuccessful performance and reveals differences across items in where behavioral information occurs within the problem-solving process.

stat.AP↗

Analyzing Process Data from Computer-Based Assessments: A Tutorial on Preprocessing, Feature Extraction, and Model-Based Inference

Computer-based assessments routinely generate detailed interaction logs -- commonly referred to as process data -- that record every action a respondent performs during task completion, yet systematic preprocessing guidance, integrated analytical workflows, and cross-method consistency checks remain scarce in the literature. This paper provides a unified, end-to-end analytical framework for analyzing process data from large-scale assessments -- covering the full pipeline from raw log preprocessing to model-based inference -- using the Programme for the International Assessment of Adult Competencies (PIAAC) Problem Solving in Technology-Rich Environments (PS-TRE) domain as an illustrative example. We first present a systematic preprocessing pipeline -- including timestamp correction, duplicate removal, action block consolidation, and LLM-assisted standardization -- that transforms raw event-level logs into analysis-ready action sequences. We then review and demonstrate two complementary families of analytical methods. The first consists of feature-based methods and their downstream applications, including descriptive process indicators, n-gram analysis with TF--IDF weighting, multidimensional scaling, and process data-informed differential item functioning (DIF) analysis. The second consists of model-based approaches, namely hidden Markov models and the subtask identification procedure. Empirical illustrations using the United States sample illustrate that n-gram-based behavioral clusters carry differential diagnostic information primarily among incorrect respondents, that multidimentionsl scaling-derived features comprehensively reconstruct observed behavioral variables, and that process-informed DIF analyses can identify and mitigate construct-irrelevant sources of group differences. Reproducible R code implementations are provided for all major techniques.

stat.AP↗