Search arXiv⌕ Search

arXiv subjects

Sumedh Gupte

Publications and source records attributed to Sumedh Gupte.

4 recordsLinked to original sources

Reinforcement learning with an expectile-based objective

We consider the policy evaluation and control in a finite horizon reinforcement learning (RL) setting under an expectile-based objective. First, we derive the mean-squared error (MSE) and concentration bounds for the classic estimator of expectiles based on independent and identically distributed (i.i.d.) samples. To the best of our knowledge, expectiles have not been analyzed in the non-asymptotic regime and the bounds we derive may be of independent interest. Next, we analyze a Monte Carlo type estimator of expectile of the Markov chain underlying a given policy. We derive upper bounds that hold in expectation as well as with high probability for this estimator. Further, we show the order-optimality of our estimator by deriving a lower bound on expectile-based policy evaluation. For the problem of control, we adopt a policy gradient approach and derive a policy gradient theorem for expectiles. Using this result, we propose a gradient estimator with a $O\left(1/m\right)$ mean-squared error bounds, where $m$ is the number of trajectories. Further, under standard assumptions for policy gradient-type algorithms, we establish smoothness of the expectile-sensitive objective, in turn leading to stationary convergence rate bounds for the overall risk-sensitive policy gradient algorithm that we propose. Finally, we conduct numerical experiments to show the utility of expectiles on popular RL benchmarks.

cs.LG↗

Optimized Certainty Equivalent Risk Minimization Using Samples: Algorithms, Convergence Rates, and Applications

We consider the optimization of the Optimized Certainty Equivalent (OCE) risk, with applications including portfolio optimization in finance, and uncertainty quantification, classification, and regression in machine learning. Our contributions cover popular special cases of OCE, such as entropic risk, mean-variance risk, and smooth variants of Conditional Value-at-Risk. Our treatment sets out the conditions that facilitate the extension of OCE to unbounded r.v.s.. We provide a useful characterization of OCE that links OCE to utility-based shortfall risk (UBSR). Our characterization enables us to form an OCE estimator from the classic sample-average approximation (SAA) of UBSR. We derive mean-squared error (MSE) bounds for our proposed OCE estimator. For OCE optimization, we first derive an expression for the OCE gradient using the characterization linking OCE to UBSR. This expression serves as the basis for a gradient estimator for the OCE. We derive non-asymptotic bounds on the MSE for the proposed OCE gradient estimator. We incorporate the aforementioned gradient estimator into a stochastic gradient (SG) algorithm to optimize OCE and quantify its convergence rate using non-asymptotic bounds that we derive. Finally, we present three experiments that use our OCE optimization algorithm to solve portfolio optimization and uncertainty quantification problems.

stat.ML↗

Gradient-based Stochastic Optimization of Utility-based Shortfall Risk

We consider the problems of estimation and optimization of utility-based shortfall risk (UBSR). We extend UBSR to cover possibly unbounded random variables. We cover prominent risk measures such as entropic risk, expectile risk, Value-at-Risk, and quadratic risk as special cases of the UBSR. In the context of estimation, we derive non-asymptotic bounds on the mean absolute error (MAE) and the mean-squared error (MSE) of the classical sample-average approximation (SAA) estimator for the UBSR. In the context of optimization, we derive an expression for the gradient of UBSR under a smooth parameterization. We propose a gradient estimator for the UBSR and derive non-asymptotic bounds on MAE and MSE for this estimator. We incorporate the aforementioned gradient estimator into a stochastic gradient (SG) optimization algorithm and derive non-asymptotic bounds on the convergence rate of our SG algorithm for optimizing UBSR under three objectives, namely, strongly convex, convex and non-convex. Finally, we conduct experiments on financial applications to demonstrate the performance of our proposed UBSR estimation and optimization algorithms.

cs.CE↗

Optimization of utility-based shortfall risk: A non-asymptotic viewpoint

We consider the problems of estimation and optimization of utility-based shortfall risk (UBSR), which is a popular risk measure in finance. In the context of UBSR estimation, we derive a non-asymptotic bound on the mean-squared error of the classical sample average approximation (SAA) of UBSR. Next, in the context of UBSR optimization, we derive an expression for the UBSR gradient under a smooth parameterization. This expression is a ratio of expectations, both of which involve the UBSR. We use SAA for the numerator as well as denominator in the UBSR gradient expression to arrive at a biased gradient estimator. We derive non-asymptotic bounds on the estimation error, which show that our gradient estimator is asymptotically unbiased. We incorporate the aforementioned gradient estimator into a stochastic gradient (SG) algorithm for UBSR optimization. Finally, we derive non-asymptotic bounds that quantify the rate of convergence of our SG algorithm for UBSR optimization.

cs.LG↗