Search arXiv⌕ Search

arXiv subjects

Arnoldas Jasonas

Publications and source records attributed to Arnoldas Jasonas.

3 recordsLinked to original sources

COSED: Setting the Bar for Open-Vocabulary Sound Event Detection

Open-vocabulary Sound Event Detection detects and temporally localizes acoustic events described by arbitrary text queries. Progress in this emerging field is hard to assess: recent methods report on disjoint task subsets under incompatible protocols without a benchmark spanning the acoustic domains and query types the task presents. We establish a comprehensive benchmark by assembling six temporally-annotated tasks: four with fixed class vocabularies over domestic, urban and mixed indoor/outdoor scenes, plus two free-text grounding tasks. We evaluate five recent methods on identical data and metrics under a label-space zero-shot criterion. Our benchmark demonstrates that no prior method is competitive across all six tasks. We then introduce COSED, which surpasses prior work on five out of six tasks while staying on par with the best method on the sixth, with margins of 12-33% on three of them. COSED is the only system in our comparison competitive on every task, and so generalizes across acoustic domains and query types better than prior work. We also provide a leave-one-out ablation study that isolates the sources of the performance benefits: scoping negatives to their corpus of origin (25.8%), combining closed- and open-world supervision (16.8%), and improving temporal processing (16.4%).

eess.AS↗

Sound Event Detection with Boundary-Aware Optimization and Inference

Temporal detection problems appear in many fields including time-series estimation, activity recognition and sound event detection (SED). In this work, we propose a new approach to temporal event modeling by explicitly modeling event onsets and offsets, and by introducing boundary-aware optimization and inference strategies that substantially enhance temporal event detection. The presented methodology incorporates new temporal modeling layers - Recurrent Event Detection (RED) and Event Proposal Network (EPN) - which, together with tailored loss functions, enable more effective and precise temporal event detection. We evaluate the proposed method in the SED domain using a subset of the temporally-strongly annotated portion of AudioSet. Experimental results show that our approach not only outperforms traditional frame-wise SED models with state-of-the-art post-processing, but also removes the need for post-processing hyperparameter tuning, and scales to achieve new state-of-the-art performance across all AudioSet Strong classes.

eess.AS↗

More Than A Shortcut: A Hyperbolic Approach To Early-Exit Networks

Deploying accurate event detection on resource-constrained devices is challenged by the trade-off between performance and computational cost. While Early-Exit (EE) networks offer a solution through adaptive computation, they often fail to enforce a coherent hierarchical structure, limiting the reliability of their early predictions. To address this, we propose Hyperbolic Early-Exit networks (HypEE), a novel framework that learns EE representations in the hyperbolic space. Our core contribution is a hierarchical training objective with a novel entailment loss, which enforces a partial-ordering constraint to ensure that deeper network layers geometrically refine the representations of shallower ones. Experiments on multiple audio event detection tasks and backbone architectures show that HypEE significantly outperforms standard Euclidean EE baselines, especially at the earliest, most computationally-critical exits. The learned geometry also provides a principled measure of uncertainty, enabling a novel triggering mechanism that makes the overall system both more efficient and more accurate than a conventional EE and standard backbone models without early-exits.

cs.SD↗