Learning Optimal Liquidation with Closing Auctions
We study liquidation when continuous trading is followed by a closing auction. The trader first sells through a limit-order book, then submits signed auction schedules to adjust the remaining inventory. A projected clearing price signal and intermediate auction feedback connect the two phases. We compare deep Q-network (DQN) policies with projected deep deterministic policy gradient (DDPG), twin delayed deep deterministic policy gradient (TD3) and soft actor-critic (SAC) policies, using synthetic rough Heston prices and historical midprice paths within a simulated market. Policies are selected and evaluated by inventory-penalized implementation shortfall, separately from a weighted training objective. They achieve lower inventory-penalized shortfall than the stylized Avellaneda-Stoikov (AS) and time-weighted average price (TWAP) references, and matched synthetic comparisons show that auction access is useful for all four learners. We furthermore find that dense auction credit improves learning; the clearing forecast is informative, but its incremental decision value is learner-dependent.