Search arXiv⌕ Search

arXiv · 2609.37845

Acoustic Honeybee Queen-State Detection Under Unseen Conditions

Abstract

Honeybee queen loss is a major threat to colony health, yet queen-status assessment remains largely manual and disruptive. Acoustic monitoring offers a non-invasive alternative by enabling continuous analysis of hive sounds. In this paper, we benchmark conventional and learned acoustic representations for automated detection of queen absence, comparing task-specific convolutional neural networks with pretrained audio transformers. Experiments are performed on 5,129 audio recordings from 3,285 hives across 47 apiaries collected between October 2024 and August 2026, evaluated using hive- and apiary-independent splits. Results show that models relying on modulation spectrograms achieve the best performance, reaching an Area Under the Receiver Operating Characteristic (AUROC) of 0.81 and Area Under the Precision-Recall Curve (AUPRC) of 0.37 on unseen hives; performance decreases under unseen-apiary evaluation. Overall, our results highlight the promise of modulation-based audio representations for non-invasive queen-status monitoring and highlight the challenge in cross-apiary model generalization.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mahsa Abdollahi, Nico Coallier, Maxime Fraser Franco, Tiago H. Falk. 2026-09-29. Acoustic Honeybee Queen-State Detection Under Unseen Conditions. https://arxiv.org/abs/2609.37845

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

NVV-Locator: From Transcript Tags to Acoustic Boundaries for Fine-Grained Nonverbal Vocalization Grounding

Human speech includes nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, which convey affective and interactional information. Existing approaches typically represent NVVs as transcript-level tags, providing limited supervision for their waveform-time boundaries. We present NVV-Locator for fine-grained NVV temporal grounding. We first unify 26 NVV categories across public resources and construct large-scale timestamp-supervised training data through dual-LLM verification, transcript-guided forced alignment, and energy-based boundary refinement. We further introduce NVV-TimeBench, an expert-refined benchmark with 667 utterances and 1,094 events. NVV-Locator uses a non-autoregressive slot-filling architecture to jointly predict lexical timestamps, NVV categories, and event boundaries. On NVV-TimeBench, it achieves 71.0% Micro F1, 70.2% Macro F1, 80.4% Macro mIoU, and 59.6 ms Macro mMAE, outperforming the evaluated large audio model counterparts. Evaluation on an external corpus further demonstrates the cross-corpus generalization of NVV-Locator.

eess.AS↗

A barrier or a booster? Familiarity effects on Mandarin emotion prosody recognition using AI-powered voice cloning

Emotion prosody perception requires simultaneous processing of acoustic cues and speaker identity. While listeners effortlessly decode natural speech, AI synthetic voices introduce cognitive complexities due to subtle acoustic atypicalities. It remains unclear how these synthetic features interact with a listener's prior social knowledge and memory of a familiar speaker. This study investigated how speech sources (human vs. AI) and speaker familiarity affect emotion recognition accuracy and cognitive load. A within-subject task with Mandarin-speaking adults evaluated behavioral (accuracy, reaction time) and physiological data (heart rate variability). Results showed that human voices yielded significantly higher accuracy and faster processing times than AI voices, while HRV did not significantly differentiate between conditions. These findings show that decoding synthetic speech is gated by top-down social cognition, highlighting limitations in current AI synthesis technologies.

eess.AS↗

Improving Predicted MOS Scores, Not Perceived Quality: Multi-Predictor Test-Time Optimization of Enhanced Speech

Non-intrusive MOS predictors are widely used instead of subjective listening tests to evaluate and rank speech enhancement (SE) systems. If they accurately reflect perceived quality, raising their scores should lead to higher-quality speech. We present the first comprehensive analysis of test-time optimization for the SE task, which directly modifies the enhanced signal to raise the average of multiple MOS predictor scores. On seven systems from the URGENT 2026 challenge, we find that 1)~all the optimized predicted scores increase while reference-based metrics remain nearly unchanged, 2)~a non-optimized predicted score does not increase, and 3)~a MUSHRA listening test shows no improvement in perceived quality. These findings reveal a risk that such optimization can distort evaluations, e.g., biasing comparisons of SE systems regardless of their perceived quality. We believe these findings can inform future evaluation practices: they suggest that predictors used for optimization should not be used for evaluation, and that challenges should keep the predictors used for ranking undisclosed.

eess.AS↗