Search arXiv⌕ Search

arXiv · 2204.04802

On the pragmatism of using binary classifiers over data intensive neural network classifiers for detection of COVID-19 from voice

Abstract

Lately, there has been a global effort by multiple research groups to detect COVID-19 from voice. Different researchers use different kinds of information from the voice signal to achieve this. Various types of phonated sounds and the sound of cough and breath have all been used with varying degree of success in automated voice-based COVID-19 detection apps. In this paper, we show that detecting COVID-19 from voice does not require custom-made non-standard features or complicated neural network classifiers rather it can be successfully done with just standard features and simple binary classifiers. In fact, we show that the latter is not only more accurate and interpretable but also more computationally efficient in that they can be run locally on small devices. We demonstrate this on a human-curated dataset of over 1000 subjects, collected and calibrated in clinical settings.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ankit Shah, Hira Dhamyal, Yang Gao, Daniel Arancibia, Mario Arancibia, Bhiksha Raj, Rita Singh. 2022-10-25. On the pragmatism of using binary classifiers over data intensive neural network classifiers for detection of COVID-19 from voice. https://arxiv.org/abs/2204.04802

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Building on the Joint-Embedding Predictive Architecture (JEPA) paradigm, a recent self-supervised learning framework that predicts latent representations of masked regions in high-level feature spaces, we propose Audio-JEPA (Audio Joint-Embedding Predictive Architecture), tailored specifically for audio data. Audio-JEPA uses a simple Vision Transformer backbone to predict latent representations of masked spectrogram patches rather than reconstructing raw audio. We pre-train on unlabeled AudioSet clips (10s, 32kHz) with random patch masking on mel-spectrograms. We evaluate on the X-ARES suite covering speech, music, and environmental sound tasks. Although our implementation is a straightforward translation of the original model to audio, the results still show comparable performance to wav2vec 2.0 and data2vec while using less than one-fifth of their training data and with no hyper-parameter tuning. All code and pretrained checkpoints are available on GitHub.

cs.SD↗

Channel-Preserving Representation Alignment for EEG-to-Music Reconstruction

Reconstructing naturalistic music from EEG requires extracting music-specific information from weak signals distributed across electrodes. Existing spectrogram regression and diffusion approaches provide audio synthesis mechanisms, but learning the EEG-to-music mapping faces an information--estimation tradeoff. Combining electrode measurements into fewer features can ease estimation from limited recordings while obscuring distinctions between music segments. We propose channel-preserving representation alignment, separating electrode information retention from prediction regularization. Per-electrode tokenization preserves spatial identity for learned aggregation, while multi-view self-distillation and channel dropout encourage consistency across electrode subsets during pretraining and music alignment. The aligned representation conditions a pretrained audio generator. Our theory explains these complementary choices. A finite-sample regression analysis identifies when retained task information outweighs estimation cost. For nonlinear alignment, an exact decomposition reveals a prediction-consistency penalty in the masked contrastive objective, with margin-dependent bounds connecting prediction variability to ranking errors across electrode subsets. On NMED-T/H, our method improves within-subject 50-way identification from 0.388 to 0.485 and generated-audio CLAP similarity from 0.635 to 0.683 compared with previous work.

cs.SD↗

Patient-level validation of a foundation-model pipeline for pediatric lung sounds: physician-annotated adventitious events are recognized, disease groups are not reliably predicted

Background: Lung-sound classifiers are usually evaluated on recording- or event-level splits, although each child contributes many recordings. We evaluated a foundation-model pipeline (PulmoVec) with the patient as the unit of partitioning and asked whether disease group can be predicted beyond age and sex. Methods: We analyzed 19693 physician-annotated respiratory events from 736 children in the public SPRSound database. A frozen Health Acoustic Representations (HeAR) encoder with low-rank adapters was trained for screening (normal versus adventitious), sound pattern (normal, crackles, wheeze/rhonchi) and disease group (pneumonia, bronchial disease, normal/other); event probabilities were stacked with age, sex and auscultation site. The primary analysis was nested patient-grouped cross-validation; comparators were the majority class, annotated event duration and demographics. Results: The area under the receiver operating characteristic curve (AUC) was 0.95 (95% CI 0.94 to 0.96) for both screening and sound pattern, against 0.78 and 0.75 for event duration alone; the positive predictive value for adventitious events was 0.72. Disease-group prediction reached an AUC of 0.58 (95% CI 0.54 to 0.62), and its accuracy was below the majority-class rate at event level (0.44 versus 0.62) and at patient level (0.46 versus 0.60). Conclusions: Recognition of physician-annotated adventitious events holds under patient-level validation and is not explained by event duration or demographics, although its positive predictive value would limit unaided use. Prediction of disease group, a clinical diagnosis, from lung-sound events alone was not demonstrated. Lung-sound studies should partition data by patient and report non-acoustic comparators.

cs.SD↗