arXiv · 2609.36478
Learning to Harvest Without Collapse in a Regenerative Commons: A Lagrangian Framework
Abstract
The tragedy of the commons poses a multi-agent safety problem: reward-seeking agents can deplete a shared resource, and cooperation among its users does not itself specify how much must be preserved. We make preservation an explicit requirement by formulating a regenerative commons as a constrained Markov game or a constrained multi-agent MDP with a designer-specified depletion budget. We develop a nonstationary Lagrangian framework that constructs a policy sequence from solutions of unconstrained games or cooperative control problems. Extending earlier time-average constructions, we introduce average-epoch solution concepts for reset episodes with discounted rewards and terminal costs. We prove a reward-independent feasibility certificate, cooperative feasibility and approximate optimality against feasible policy mixtures, and an extension to unbiased sampled costs. For self-interested agents, a constrained Nash certificate quantifies the price-dispersion term introduced by deviations that redistribute budget across epochs. Under the stated assumptions on solver accuracy and multiplier updates, these results give constrained policy-sequence guarantees using solutions of unconstrained problems. Experiments with constrained IPPO and MAPPO in a Gordon-Schaefer fishery examine how depletion budgets shape stock retention, harvest rewards, and price adaptation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jose Tupayachi, Xueping Li, Soham Das. 2026-09-29. Learning to Harvest Without Collapse in a Regenerative Commons: A Lagrangian Framework. https://arxiv.org/abs/2609.36478
Cite the original work for its findings. Save a collection to share your selection of sources.