Search arXivSearch

arXiv · 2608.28199

Fine-Tuning Autobidders with Group Relative Policy Optimization

Abstract

Automated bidding (autobidding) is a core component of modern online advertising systems. Within this component, advertisers delegate sequential bid decisions to algorithms that must maximize campaign value while adhering to constraints such as a limited budget and a target cost-per-click (CPC). One of the approaches to resolve the autobidding problem is to formulate it as a Markov decision process and use reinforcement learning (RL) to train a bid generation function. The standard RL framework is actor-critic, which consists of an actor network that generates actions and a critic network that estimates the value of those actions. In our setting, the action is typically a bid or related pacing multiplier, and the value is the expected return from the auction given the bid. However, the alternating training of actor-critic RL models leads to instability and reduced robustness to noise. To address these issues, we adapt the Group Relative Policy Optimization (GRPO) framework to the autobidding setting. This framework is a \emph{critic-free} policy-gradient method originally developed for large language model post-training, where the ground-truth target is unknown. The autobidding setting shares this property, since the optimal bid is unknown in advance. Moreover, GRPO in the LLM domain is used to fine-tune the pre-trained model, and we use the same technique to enhance the performance of the strong heuristic baseline. We empirically compare Autobidding GRPO with actor-critic models, simple heuristics, and controller-based methods on the BAT, iPinYou, and AuctionNet benchmarks. Extensive experiments show that Autobidding GRPO consistently outperforms baselines in clicks and is the best or second-best method in conversion volume.

Explore related subjects

Keep this discovery

BibTeXRIS

Anton Safin, Alexandra Khirianova, Andrey Pudovikov, Aleksandr Katrutsa, Egor Samosvat. 2026-08-28. Fine-Tuning Autobidders with Group Relative Policy Optimization. https://arxiv.org/abs/2608.28199

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per-scene optimization, achieving strong generalization. However, enforcing explicit multi-view geometric consistency, e.g., through bundle adjustment, is computationally costly and is thus not imposed during VFM pretraining, so such inconsistency can arise. To address this, implicit self-consistency derived from model outputs (e.g., pointmaps, features), though enforced at test-time in prior work, delivers inherently limited performance gain, especially on scenes where the pretrained VFM is highly inaccurate. In contrast to this implicit signal, we propose Self-Geometry, a plug-and-play test-time adaptation pipeline that directly imposes explicit multi-view geometric constraints using 2D pixel correspondences as pseudo ground-truth. Our proposed Self-Geometry consists of Geometric Disentanglement Optimization, which combines Multi-View Consistency and Epipolar Consistency losses with Gradient Disentanglement to prevent gradient conflict; Frame Angular-Neighbor, a view sampler based on SO(3) geodesic distances for lightly imposing these constraints; and Lightweight TTA, which adapts VFMs via LoRA. Our method achieves consistent improvements in both pose and geometry estimation across six VFMs (VGGT, $π^3$, DA3-Giant/Large/Base/Small) and four benchmarks (7Scenes, ETH3D, ScanNet++, HiRoom).

cs.CV

Peer Oversight in Collective Decision Making

This article introduces peer $k$-oversight, a property of sequential collective decision mechanisms requiring at least $k$ agents to be responsible for every harmful outcome. It is shown that whenever $k$-oversight can be achieved by redistributing control over the decisions in a mechanism, it can be achieved using just $k$ agents. A polynomial-time algorithm is also presented that determines whether such a redistribution exists and, when it does, constructs one. These results establish peer oversight as a tractable design principle for multiagent decision-making systems.

cs.GT