arXiv · 2609.29203
Few-Shot Calibration for Sim-to-Real Single-Channel Speaker Distance Estimation
Abstract
Speaker distance estimators are trained almost exclusively on simulated room acoustics, because real recordings annotated with the true talker-to-microphone distance are scarce. We show that models trained this way transfer poorly. On three real corpora we evaluate, simply predicting the average distance of the corpus is more accurate than any learned model. Then, we ask how few labelled real utterances are needed to make a frozen, synthetic-trained estimator useful, and study post-hoc calibration maps that rescale its output without gradients or retraining. An analysis of the achievable error shows that what the calibration is not limited by the absolute accuracy of the estimator, but how well it orders utterances by distance, since a constant bias or a wrong output scale is removed exactly by the calibration itself. Balancing this against the cost of estimating each coefficient from few samples yields a criterion that accounts for which map wins on which corpus and at which annotation budget, together with a shrinkage variant that requires no hard decision. Our findings suggest selecting synthetic checkpoints by linear correlation with true distances rather than by absolute error. Code, datasets, and analysis are available at https://github.com/michaelneri/audio-distance-estimation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Michael Neri, Archontis Politis, Tuomas Virtanen. 2026-09-24. Few-Shot Calibration for Sim-to-Real Single-Channel Speaker Distance Estimation. https://arxiv.org/abs/2609.29203
Cite the original work for its findings. Save a collection to share your selection of sources.