TY - RPRT TI - Subspace Inference Enables Efficient Active Reward Learning from Preferences AU - Yutai Zhou AU - Erdem Bıyık PY - 2026 UR - https://arxiv.org/abs/2609.04066 ID - 2609.04066 ER -