Search arXivSearch

arXiv subjects

Subham Kumar

Publications and source records attributed to Subham Kumar.

8 recordsLinked to original sources

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown. In this study, we first conduct the systematic audit of ASR performance on real-world psychiatric interview data spanning Kannada, Hindi and Indian English, comparing eight state-of-the-art models including IndicWhisper, WhisperLargeV3, Sarvam, GoogleS2T, Gemma3n, OmniLingual, Vaani, and Gemini. Our results reveal substantial variability across models and languages, with some systems performing competitively in Indian English but failing in regional speech. We further fine-tune two of the best performing opensource models, i.e., Gemma3n and OmniLingual, using various methods. With this, we uncover systematic performance gaps tied to speaker role and gender, raising concerns about equitable deployment in clinical settings, which are further mitigated by fairness-aware fine-tuning. To this end, we propose SamaVaani, a unified debiasing technique that simultaneously improves ASR performance and improves fairness across demographic groups.

cs.CL

Design and Performance Evaluation of Secure RF and WiFi-Based Communication in Drone Swarms via Testbed Implementation

Unmanned aerial vehicle (UAV) swarms rely on distributed coordination and cooperative communication to support scalable operations, extended coverage, and applications such as surveillance and real-time data exchange. Wireless technologies such as radio frequency (RF) and WiFi are widely used for UAV-to-UAV and UAV-to-ground control station (GCS) communication but introduce significant security challenges. MAVLink, the predominant communication protocol in UAV systems, provides message integrity and authentication but lacks built-in encryption, leaving telemetry traffic vulnerable to eavesdropping. In our previous work, we proposed MAVShield, a lightweight encryption framework for MAVLink communications. In this paper, MAVShield, AES-CTR, Speck-CTR, ChaCha20, and Rabbit are integrated into four custom-built UAVs to establish secure communication links over RF and WiFi channels. Their performance is evaluated through flight experiments using a UAV swarm testbed. Encrypted telemetry data enable autonomous formation control and collision avoidance during flight. For collision avoidance, we develop a modified artificial potential field (APF) algorithm that computes attractive and repulsive forces directly in geodetic coordinates, eliminating Cartesian transformations and reducing trajectory oscillations while avoiding local-minimum trapping. CPU utilization, memory consumption, and packet delivery ratio (PDR) are measured for each encryption scheme. Results show that MAVShield achieves performance comparable to unencrypted communication while outperforming AES-CTR, Speck-CTR, ChaCha20, and Rabbit in overall efficiency. Algebraic cryptanalysis and Wireshark-based traffic analysis demonstrate resistance to key-recovery attacks and protection of telemetry confidentiality. The results indicate that MAVShield is an efficient and secure solution for UAV swarm communication.

cs.CR

ASR Under the Stethoscope: Evaluating Biases in Clinical Speech Recognition across Indian Languages

Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare contexts remains largely unknown. In this study, we conduct the first systematic audit of ASR performance on real world clinical interview data spanning Kannada, Hindi, and Indian English, comparing leading models including Indic Whisper, Whisper, Sarvam, Google speech to text, Gemma3n, Omnilingual, Vaani, and Gemini. We evaluate transcription accuracy across languages, speakers, and demographic subgroups, with a particular focus on error patterns affecting patients vs. clinicians and gender based or intersectional disparities. Our results reveal substantial variability across models and languages, with some systems performing competitively on Indian English but failing on code mixed or vernacular speech. We also uncover systematic performance gaps tied to speaker role and gender, raising concerns about equitable deployment in clinical settings. By providing a comprehensive multilingual benchmark and fairness analysis, our work highlights the need for culturally and demographically inclusive ASR development for healthcare ecosystem in India.

cs.CL

GRACE: Graph Neural Networks for Locus-of-Care Prediction under Extreme Class Imbalance

Determining the appropriate locus of care for addiction patients is one of the most critical clinical decisions that affects patient treatment outcomes and effective use of resources. With a lack of sufficient specialized treatment resources, such as inpatient beds or staff, there is an unmet need to develop an automated framework for the same. Current decision-making approaches suffer from severe class imbalances in addiction datasets. To address this limitation, we propose a novel graph neural network (GRACE) framework that formalizes locus of care prediction as a structured learning problem. In addition, we propose a new approach of obtaining an unbiased meta-graph to train a GNN to overcome the class imbalance problem. Experimental results with real-world data show an improvement of 11-35% in terms of the F1 score of the minority class over competitive baselines. Further, if we jointly finetune the base embedding fed into GRACE as input together with the rest of the GNN component of GRACE, there is a remarkable boost of 15.8% in performance.

cs.LG

Investigating the fission-like fragments in the $^{12}$C + $^{208}$Pb system at E$^{\star}$ $\approx$ 31.8--45.4 MeV

In this work, the cross-sections of 25 fission-like fragments within the mass range 76$\leq$A$\leq$141, expected to be populated via fission of moderately excited compound nucleus produced as a result of complete and/or incomplete fusion in $^{12}$C+$^{208}$Pb reaction at E$_{\rm lab}$ = 81.9 and 75.8 MeV, have been measured using activation technique followed by offline $\gamma$-ray spectroscopy. The yields of different fission-like fragments have been analyzed to generate isotopic and isobaric yield distributions. The value of the mass dispersion parameter, $\sigma^2_A$, is found to be 2.93 and 2.65 for Antimony (Sb) isotope at excitation energy E$^{\star}$ = 45.4 and 39.6 MeV, and 1.24 for Indium (In) isotope at E$^{\star}$ = 45.4 MeV. The charge dispersion parameter $\sigma_Z$ for Sb is estimated to be 0.769 and 0.714 at E$^{\star}$ = 45.4 and 39.6 MeV, respectively. For In isotopes, the value of $\sigma_Z$ is estimated to be 0.430 at E$^{\star}$ = 45.4 MeV. The value of mass and charge dispersion parameters for Sb and In isotopes have been found to be in good agreement with the values reported in the literature for similar systems. The mass distribution of fission-like fragments is found to be fitted with a Gaussian function, except for a few data points, indicating their population via compound nucleus fission. Further, the mass variance ($\sigma^2_M$) displays linear increment with an increase in excitation energy. Two medically important isotopes, $^{99m}$Tc and $^{111}$In, are populated in this system, suggesting a potential formation route.

nucl-ex

Incomplete fusion in $^{193}$Ir($^{12}$C, x)$^{205}$Bi reaction at $E_{lab}$ $\approx$ 5-7 AMeV

Low-energy heavy-ion induced reactions often involve incomplete fusion, but the dependence of ICF on various entrance-channel parameters remains unclear. In this work, we measure channel-by-channel production cross-sections of different evaporation residues populated via complete and/or incomplete fusion in $^{12}$C+$^{193}$Ir system at $E_{lab}$ $\approx$ 64--84 MeV ($\approx$ 5--7 AMeV) using the stacked-foil activation technique followed by offline $\gamma$-spectroscopy. Experimentally measured excitation functions have been analyzed in the framework of the statistical model code PACE4 using different values of the level-density parameter ($a$ = A/9-A/15 MeV${^{-1}}$). In the analysis of excitation functions, the $xn$ and $pxn$ channels (after correcting with their precursor contributions) have been explained fairly well with $a$ = A/13 MeV${^{-1}}$; however, almost all $\alpha$-emitting channels showed substantial enhancement over PACE4 predictions, which has been attributed to incomplete fusion. The incomplete fusion fraction ($F_{ICF}$) increases linearly with energy from 12\% to 18\% at 64 and 84 MeV, respectively. For better insights into the onset and strength of ICF, the variations of $F_{ICF}$ have been studied as a function of different entrance-channel parameters, which are found to increase with mass asymmetry, Coulomb factor, and neutron skin thickness. Further analysis of the data suggests the onset of ICF below the critical angular momentum ($\ell<\ell_{crit}$). Projectile breakup-driven incomplete fusion is found to suppress complete fusion by $\approx12\%$ and $\approx6\%$ w.r.t. the universal fusion function and the improved fusion function, respectively. These findings highlight the critical role of projectile structure at 5--7 AMeV energies, with implications for high-spin spectroscopy and reaction modeling.

nucl-ex

newsSweeper at SemEval-2020 Task 11: Context-Aware Rich Feature Representations For Propaganda Classification

This paper describes our submissions to SemEval 2020 Task 11: Detection of Propaganda Techniques in News Articles for each of the two subtasks of Span Identification and Technique Classification. We make use of pre-trained BERT language model enhanced with tagging techniques developed for the task of Named Entity Recognition (NER), to develop a system for identifying propaganda spans in the text. For the second subtask, we incorporate contextual features in a pre-trained RoBERTa model for the classification of propaganda techniques. We were ranked 5th in the propaganda technique classification subtask.

cs.CL

Can I teach a robot to replicate a line art

Line art is arguably one of the fundamental and versatile modes of expression. We propose a pipeline for a robot to look at a grayscale line art and redraw it. The key novel elements of our pipeline are: a) we propose a novel task of mimicking line drawings, b) to solve the pipeline we modify the Quick-draw dataset to obtain supervised training for converting a line drawing into a series of strokes c) we propose a multi-stage segmentation and graph interpretation pipeline for solving the problem. The resultant method has also been deployed on a CNC plotter as well as a robotic arm. We have trained several variations of the proposed methods and evaluate these on a dataset obtained from Quick-draw. Through the best methods we observe an accuracy of around 98% for this task, which is a significant improvement over the baseline architecture we adapted from. This therefore allows for deployment of the method on robots for replicating line art in a reliable manner. We also show that while the rule-based vectorization methods do suffice for simple drawings, it fails for more complicated sketches, unlike our method which generalizes well to more complicated distributions.

cs.CV