Search arXiv⌕ Search

arXiv · 2609.36787

Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions

Abstract

Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable policy-gradient algorithm for large sequential games with continuous or mixed discrete and continuous actions. It combines magnetic mirror descent with a mixture of Gaussians reparametrization, trained via self-play. We show that it approximates equilibrium in games where gradient descent fails. In sequential games, it outperforms neural fictitious self-play and matches or outperforms the final strategies of policy space response oracles with 3.5--5.5$\times$ fewer samples. In heads-up no-limit Texas hold'em, it performs on par with Slumbot.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ondřej Kubíček, Viliam Lisý, Tuomas Sandholm. 2026-09-29. Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions. https://arxiv.org/abs/2609.36787

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

REGREACT: Self-Correcting Multi-Agent Pipelines for Structured Regulatory Information Extraction

Extracting structured, machine-readable compliance criteria from regulatory documents remains an open challenge. Single-pass language models hallucinate structural elements, lose hierarchical relationships, and fail to resolve inter-document dependencies. We introduce RegReAct, a self-correcting multi-agent framework that decomposes regulatory information extraction into seven specialized stages, each with an Observe--Diagnose--Repair (ODR) loop that validates outputs against the source, correcting not only model hallucinations but also cross-reference errors in the regulations themselves. To ensure structural accuracy, RegReAct constructs a typed criterion graph; to ensure completeness, it resolves external dependencies by retrieving, summarizing, and embedding referenced legal content inline, producing self-contained outputs. Applying RegReAct to three EU Taxonomy Delegated Acts, we construct a dataset comprising 242 activities with over 4,800 hierarchical criteria, thresholds, and enriched source summaries. Evaluation against a GPT-4o single-pass baseline shows that RegReAct outperforms it across all structural and semantic metrics.

cs.MA↗

Hierarchical Multiagent Reinforcement Learning for Multi-Group Tax Game

Taxation is a fundamental instrument of economic policy and has long been studied through economic models. However, most existing taxation models focus on interactions within a single economic group, typically with one government and multiple households. Such a setting overlooks interactions among independent economic groups, where governments may compete for economic resources through fiscal policies. To capture these interactions, we formulate taxation as a hierarchical multi-group tax game. Within each group, a government sets tax policies while households respond, forming a hierarchical government--household interaction. Across groups, governments compete through fiscal policies, coupling the economic dynamics of distinct groups. This coupled structure poses challenges for standard multi-agent reinforcement learning (MARL) methods. To this end, we propose Multi-Group PPO (MGPPO), a bilevel MARL algorithm designed for hierarchical multi-group interactions. MGPPO incorporates Hierarchical Sampling to coordinate learning across agent levels and Curriculum Learning to improve training stability. We further develop a multi-group taxation simulation environment grounded in classical economic models, supporting the evaluation of fiscal policies under inter-group competition. Simulations show that in the two-group settings, MGPPO reduces the income and wealth Gini coefficients by 5.0\% and 4.9\%, respectively, and increases mean GDP by 4.1\% compared with IPPO.

cs.MA↗

Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems

Large Language Model (LLM)-based automatic Multi-Agent Systems (MAS) generation has become a crucial frontier for tackling complex tasks. However, existing methods face a dilemma between model capability and experience retention. Inference-time MAS leverages frozen frontier LLMs but repeats identical searches without learning from past experience. Conversely, Training-time MAS internalizes experience via gradient updates but is constrained by the low capability ceiling of smaller models, and is hard to scale to large frontier LLMs. To bridge this gap, we propose Skill-MAS, a novel third path that decouples experience retention from parametric updates by conceptualizing the high-level orchestration capability as an evolvable Meta-Skill. Skill-MAS refines this architectural knowledge through a closed optimization loop: (1) Multi-Trajectory Rollout samples a behavioral distribution for each task under the current Meta-Skill; and (2) Selective Reflection adaptively selects priority tasks and applies hierarchical contrastive analysis to distill systemic experience into generalizable, strategy-level principles. Extensive experiments across four complex benchmarks and four distinct LLMs demonstrate that Skill-MAS not only achieves remarkable performance gains but also maintains a favorable cost-performance trade-off. Further analysis reveals that the evolved Meta-Skills are highly robust and exhibit strong transferability across unseen tasks and different LLMs.

cs.MA↗