arXiv · 2609.26025
MICRO: Multi-Fidelity Active Search for Severe Error Discovery
Abstract
Human feedback can vary in cost and informativeness. Strong feedback can reveal severe errors but is costly, so cheaper quality ratings can help decide which items to annotate. We propose MICRO (Multi-Fidelity Impact Clustered Rollout), an active search framework that allocates a shared budget to these feedback types to maximise confirmed severe error discoveries. MICRO jointly models ratings and annotation losses conditional on item features to steer acquisition. It clusters acquisitions by their predicted impact on severity probabilities to select diverse candidates, then uses rollout to estimate their discovery value. Experiments on WMT20 English-German show that ratings improve both loss reconstruction and severity prediction. MICRO achieves the highest mean discovery count across four budget and rating cost settings, with similar performance to adapted MF-ENS in one and significant gains over all six comparison policies, including two rollout controls, in the other three $(p<.001)$.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Orlando Leone, Niclas Pokel, Pehuén Moure, Yingqiang Gao, Roman Boehringer. 2026-09-22. MICRO: Multi-Fidelity Active Search for Severe Error Discovery. https://arxiv.org/abs/2609.26025
Cite the original work for its findings. Save a collection to share your selection of sources.