Martingale deep neural network for very high-dimensional stochastic optimal controls
We propose a martingale deep learning method for very high-dimensional stochastic optimal control problems (SOCPs) via their associated Hamilton--Jacobi--Bellman equations and the verification theory for the optimal control. The method decomposes the problem into two coupled components: a martingale formulation for the value function and an integral optimality condition for the feedback control. For the value function, we extend DeepMartNet by introducing a pilot process for state-space exploration and system processes for martingale property enforcement.This yields a derivative-free formulation that avoids computing PDE solution's derivatives and online simulations of full trajectories, while enabling parallel loss evaluation in both time and space. Also, the martingale condition is imposed in a weak Galerkin form realized with adversarial learning, which avoids direct computation of conditional expectations. The control is learned by minimizing an integral residual of the optimality condition over the explored state space, thereby avoiding state-space point-wise optimization over the control set. Numerical results demonstrate accurate and efficient performance for SOCPs with complex patterns in value function and controls in dimensions up to 10,000.