arXiv · 2610.07973
Learning a Ranking from Human Feedback in Log-Concave Random Utility Models
Abstract
We study the problem of recovering the ranking of a fixed set of items according to their unknown numerical utilities. At each interaction with the environment, a learner presents the item set to a human and receives comparative feedback of two types. Under full-ranking feedback, each interaction reveals a noisy ranking of all items, whereas under winner-only feedback, it reveals only the item ranked first. In both settings, we model human feedback using a random utility model with log-concave noise and study the number of observations needed to recover an $ε$-accurate ranking with high probability. This novel criterion tolerates ordering errors only between items whose utilities differ by less than $ε$. For both feedback types, we establish worst-case sample-complexity lower bounds and develop algorithms that match these bounds up to logarithmic factors. Neither algorithm requires knowledge of the noise distribution, while only requiring an upper bound on its variance. Our results show that the ranking problem under winner-only feedback is intrinsically harder by exposing the sample complexity dependence on the minimum winning probability across the item set.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Diego Alovisetti, Marco Mussi, Alberto Maria Metelli. 2026-10-06. Learning a Ranking from Human Feedback in Log-Concave Random Utility Models. https://arxiv.org/abs/2610.07973
Cite the original work for its findings. Save a collection to share your selection of sources.