Search arXiv⌕ Search

arXiv · 2609.31750

Dimension-Specific Imbalance and an Adaptive Hybrid Label Strategy for Multi-Task Affective State Recognition in Classroom Video

Abstract

Recognition of student affective states from classroom video is constrained by a class imbalance problem whose true nature is, we argue, under-analyzed. We show that on DAiSEE -- the de facto benchmark for this task -- class imbalance is fundamentally a dimension-specific phenomenon: raw imbalance ratios reach as high as approximately 79:1 (Frustration), and the structure of imbalance -- not only its magnitude -- varies across dimensions, so uniform algorithmic treatments (aggressive class reweighting, focal loss, uniform binary simplification) fail in dimension-dependent ways. We propose the Adaptive Hybrid Label Strategy (AHLS), which assigns four-level classification to dimensions with manageable imbalance and binary classification to severely long-tailed dimensions, coupled with a null-class placeholder mechanism that stabilizes multi-task optimization by compressing the maximum effective inverse-frequency weight ratio from as high as ~79$\times$ to at most ~8.5$\times$ across the four dimensions. Validated on a lightweight FERShuffleNetV2 + LTCN architecture (0.39 M parameters, 0.07 G FLOPs), the strategy attains the highest mean Macro-F1 of 47.70% across all four affective dimensions among 11 compared methods, including five state-of-the-art deep models (up to 167$\times$ larger) and five traditional machine-learning baselines. A cross-model behavioral analysis on DAiSEE indicates that, in this setting, classification strategy may influence minority-state recognition more strongly than further backbone scaling alone. We also report evidence of an annotation-density limitation of DAiSEE on minority affective states, which suggests that benchmark design may bound future progress as much as algorithmic refinement. Index Terms Affective computing, classroom video analysis, class imbalance, multi-task learning, lightweight deep learning, DAiSEE, engagement recognition.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xiangqian Li. 2026-09-23. Dimension-Specific Imbalance and an Adaptive Hybrid Label Strategy for Multi-Task Affective State Recognition in Classroom Video. https://arxiv.org/abs/2609.31750

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Self-eXplainable AI for Medical Image Analysis: A Survey and New Outlooks

The increasing demand for transparent and reliable models, particularly in high-stakes decision-making areas such as medical image analysis, has led to the emergence of eXplainable Artificial Intelligence (XAI). Post-hoc XAI techniques, which aim to explain black-box models after training, have raised concerns about their fidelity to model predictions. In contrast, Self-eXplainable AI (S-XAI) offers a compelling alternative by incorporating explainability directly into the training process of deep learning models. This approach allows models to generate inherent explanations that are closely aligned with their internal decision-making processes, enhancing transparency and supporting the trustworthiness, robustness, and accountability of AI systems in real-world medical applications. To facilitate the development of S-XAI methods for medical image analysis, this survey presents a comprehensive review across various image modalities and clinical applications. It covers more than 200 papers from three novel perspectives: 1) input explainability through the integration of explainable feature engineering and knowledge graph, 2) model explainability via attention-based learning, concept-based learning, and prototype-based learning, and 3) output explainability by providing textual explanations and visual counterfactual explanations. This paper also outlines desirable characteristics of S-XAI models for assessing explanation quality, while discussing major challenges and future research directions in developing S-XAI for medical image analysis.

cs.CV↗

Imitating Radiological Scrolling: A Global-Local Attention Model for 3D Chest CT Volumes Multi-Label Anomaly Classification

The rapid increase in the number of Computed Tomography (CT) scan examinations has created an urgent need for automated tools, such as organ segmentation, anomaly classification, and report generation, to assist radiologists with their growing workload. Multi-label classification of Three-Dimensional (3D) CT scans is a challenging task due to the volumetric nature of the data and the variety of anomalies to be detected. Existing deep learning methods based on Convolutional Neural Networks (CNNs) struggle to capture long-range dependencies effectively, while Vision Transformers require extensive pre-training, posing challenges for practical use. Additionally, these existing methods do not explicitly model the radiologist's navigational behavior while scrolling through CT scan slices, which requires both global context understanding and local detail awareness. In this study, we present CT-Scroll, a novel global-local attention model specifically designed to emulate the scrolling behavior of radiologists during the analysis of 3D CT scans. Our approach is evaluated on two public datasets, demonstrating its efficacy through comprehensive experiments and an ablation study that highlights the contribution of each model component.

cs.CV↗

Noise-adapted Neural Operator for Robust Non-Line-of-Sight Imaging

This work has been submitted for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible. Computational imaging, especially non-line-of-sight (NLOS) imaging, the extraction of information from obscured or hidden scenes is achieved through the utilization of indirect light signals resulting from multiple reflections or scattering. The inherently weak nature of these signals, coupled with their susceptibility to noise, necessitates the integration of physical processes to ensure accurate reconstruction. This paper presents a parameterized inverse problem framework tailored for large-scale linear problems in 3D imaging reconstruction. Initially, a noise estimation module is employed to adaptively assess the noise levels present in transient data. Subsequently, a parameterized neural operator is developed to approximate the inverse mapping, facilitating end-to-end rapid image reconstruction. Our 3D image reconstruction framework, grounded in operator learning, is constructed through deep algorithm unfolding, which not only provides commendable model interpretability but also enables dynamic adaptation to varying noise levels in the acquired data, thereby ensuring consistently robust and accurate reconstruction outcomes. Furthermore, we introduce a novel method for the fusion of global and local spatiotemporal data features. By integrating structural and detailed information, this method significantly enhances both accuracy and robustness. Comprehensive numerical experiments conducted on both simulated and real datasets substantiate the efficacy of the proposed method. It demonstrates remarkable performance with fast scanning data and sparse illumination point data, offering a viable solution for NLOS imaging in complex scenarios.

cs.CV↗