Search arXivSearch

arXiv · 2607.21512

Group boarding for airplanes: benchmarking static policies and optimizing dynamic assignment with deep reinforcement learning

Abstract

Improving boarding efficiency reduces airplane turnaround time and improves passenger experience. Airlines typically assign passengers to a few sequential boarding groups using static seat-based rules. Yet arrivals, seat choices, and luggage are sequential and random, and a static rule ignores the seats earlier passengers have already taken. We propose the first dynamic formulation of boarding group assignment. As each passenger checks in, we observe earlier passengers' seats and groups, the current passenger's seat, and optional luggage information, then assign a group while keeping companions together. We formulate dynamic group assignment as a Markov decision process and solve it with reinforcement learning (RL). The policy uses a convolutional neural network to encode the checked-in seat-assignment state and is trained by proximal policy optimization. The reward balances total boarding time and average individual boarding time. We benchmark the proposed RL policy against three companion-compatible static policies (back-to-front, modified Steffen, and alternating block) in an in-house simulator covering six single- and double-aisle layouts. Back-to-front with optimized group sizes achieves the shortest total boarding time and average individual boarding time among the static benchmarks across all layouts. The dynamic RL policy further outperforms it on both metrics in every layout. On a representative case, the RL policy outperforms the optimal back-to-front by up to 9.8\% in total boarding time and 22.8\% in average individual time. Sweeping the reward weight yields an approximate Pareto frontier for operator choice. Trained policies remain robust under out-of-distribution operating conditions, including varying load factors, companion sizes, and luggage loads.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Minyu Shen, Weihua Gu, Junqi Ma, Boqian Song, Li Zhen, Gang Kou. 2026-07-23. Group boarding for airplanes: benchmarking static policies and optimizing dynamic assignment with deep reinforcement learning. https://arxiv.org/abs/2607.21512

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The time interpretation of expected utility theory

Ergodicity economics is a new branch of economic theory that notes the conceptual difference between time averages and expectation values, which coincide only for ergodic observables. It postulates that individual agents maximise the time average growth rate of wealth, known widely as growth optimality. This contrasts with the dominant behavioural model in economics, expected utility theory, in which agents maximise expectation values of changes in psychologically transformed wealth. Historically, growth optimality was explored for additive and multiplicative gambles. Here we apply it to a general class of wealth dynamics, extending the range of economic situations where it may be used. Moreover, we show a correspondence between growth optimality and expected utility theory, in which the ergodicity transformation in the former is identified as the utility function in the latter. This correspondence offers a theoretical basis for choosing utility functions and predicts that wealth dynamics are strong determinants of risk preferences.

econ.GN

Monetary Regimes and Trade before the Classical Gold Standard: Evidence from the Latin Monetary Union

This paper reexamines the trade effects of the Latin Monetary Union (LMU), a 19th century agreement to standardize gold and silver coinage among several European countries. The LMU provides a useful setting for studying whether monetary arrangements fostered trade before the classical gold standard, when gold, silver, bimetallic, and paper regimes coexisted. Because some countries already shared other monetary standards, treating all non-member pairs as a single control group mixes pairs with and without alternative forms of monetary coordination. I classify pairs by standard and estimate the LMU effect relative to pairs without a common standard, bringing the comparison closer to those used in the literature on the gold standard and contemporary currency unions. The results suggest that the LMU increased trade between its members by approximately 30\% during its early years, when bimetallism was still credible. These effects subsequently faded, converging to zero by the end of the 1870s. More broadly, these findings also highlight the importance of accounting for the existing monetary regimes when estimating the trade effects of other international policies.

econ.GN

Access to Live AI Advice and Behavior Under Risk: An Incentivized Experiment

Generative AI has become an everyday advisor, and the systems people consult are live and interactive, not pre-scripted. We ask whether access to such a system changes behavior under risk. In an incentivized experiment (N = 158), participants made lottery choices with an optional decision aid presented as a conventional pre-written tool, a live one-shot AI, or a live interactive AI they could query, with information format held equivalent across conditions. Risk preferences are elicited via DOSE. We find no evidence that access to a live AI advisor changes risk aversion.

econ.GN