Search arXivSearch

arXiv subjects

Maryam Kamgarpour

Publications and source records attributed to Maryam Kamgarpour.

2 recordsLinked to original sources

Provably Safe Sim-to-Real Transfer

To mitigate the sample complexity of real-world reinforcement learning (RL), a common practice is to first train a policy in a simulator, where samples are cheap, and then deploy the learned policy in the real world with the hope that it generalizes effectively. Such direct sim-to-real transfer is not guaranteed to succeed: simulator-trained policies can be suboptimal in the real world due to sim-to-real mismatch. Correcting this mismatch requires collecting data from the real system, but in many applications, such as robotics and healthcare, this data-collection process is itself subject to safety constraints. This gives rise to the problem of safe sim-to-real transfer: how can an agent exploit an imperfect simulator while ensuring safe real-world data collection and learning a near-optimal feasible policy for the target system? We address this problem by formulating safe sim-to-real transfer within the framework of reward-free safe RL. We design a computationally efficient algorithm that exploits simulator information to provably reduce real-world interaction while ensuring safe exploration and enabling the computation of a near-optimal feasible policy for any potential reward function. Our real-world sample complexity bound characterizes the benefit of using the simulator in terms of the sim-to-real mismatch.

cs.LG

Refundable Deposits: How to Restore Cooperation in Finitely Repeated Games

While infinitely repeated games admit a rich set of Nash equilibria, finitely repeated games typically have a much smaller and often inefficient one. We show how to enlarge this set using deposits: in each period a player may place a refundable sum with a neutral intermediary, returned when the game ends and forfeited following a deviation. Paying these deposits is voluntary and incentive compatible at every stage, so no commitment by the players is assumed, the only commitment required being that of the intermediary to a refund rule fixed before play begins. The mechanism sustains payoff profiles more efficient than those of the standard equilibria, without altering the underlying game and without transfers between players. We demonstrate it on the prisoner's dilemma, a congestion game, and a public goods game, all settings where cooperation cannot emerge in the standard finitely repeated version. We also apply it to a dynamic common-pool resource, suggesting that the construction extends beyond repeated stage-games.

cs.GT