Search arXivSearch

arXiv · 2509.16821

A model free approach for continuous-time optimal tracking control with unknown user-define cost and constrained control input via advantage function

Abstract

This paper presents a pioneering approach to solving the linear quadratic regulation (LQR) and linear quadratic tracking (LQT) problems with constrained inputs using a novel off-policy continuous-time Q-learning framework. The proposed methodology leverages a novel concept of the Advantage function for linear continuous systems, enabling solutions to be obtained without the need for prior knowledge of the reward matrix weights, state resetting, or assuming the existence of a predefined admissible controller. This framework includes multiple algorithms (Algs) tailored to address these control problems under model-free conditions, without requiring any knowledge about system dynamics. Two distinct implementation methods are explored: the first processes state and input data over a fixed time interval, making it well-suited for LQR problems, while the second method operates over multiple intervals, offering a practical solution for tracking problems with constrained inputs. The convergence of the proposed algorithms is verified theoretically. Finally, the simulation results of the F-16 aircraft system are presented for the two problems to validate the effectiveness of the proposed method.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Duc Cuong Nguyen, Quang Huy Dao, Phuong Nam Dao. 2025-09-20. A model free approach for continuous-time optimal tracking control with unknown user-define cost and constrained control input via advantage function. https://arxiv.org/abs/2509.16821

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Safe learning-based control via function-based uncertainty quantification

Uncertainty quantification is essential when deploying learning-based control methods in safety-critical systems. This is commonly realized by constructing uncertainty tubes that enclose the unknown function of interest, e.g., the reward and constraint functions or the underlying dynamics model, with high probability. However, existing approaches for uncertainty quantification typically rely on restrictive assumptions that encode smoothness properties of the unknown function, such as a known norm in a function space. Moreover, these methods usually struggle with discontinuities. In this paper, we model the unknown function as a random function from which independent and identically distributed realizations can be generated. We then construct uncertainty tubes via the scenario approach that hold with high probability. Our uncertainty tubes rely solely on sampled realizations and can therefore accommodate discontinuities represented by the sampling model. We integrate these uncertainty tubes into a safe Bayesian optimization algorithm with which we safely tune control parameters on a real Furuta pendulum.

eess.SY

Enhanced ShockBurst for Ultra Low-Power On-Demand Sensing

On-demand sensing requires battery-powered Internet-of-Things (IoT) and implantable medical devices to remain in deep sleep and activate wireless communication only when data transmission is required. In such systems, battery lifetime depends strongly on radio active time. This work investigates how communication architecture and physical layer (PHY) configuration influence radio active time by comparing connection-oriented Bluetooth Low Energy (BLE) with connectionless Enhanced ShockBurst (ESB) on identical BLE-compatible hardware. Under identical 2 Mbps PHY configurations, ESB reduces wake-up latency and energy consumption to approximately one-twentieth of BLE by eliminating connection establishment and maintenance overhead. Increasing the ESB PHY rate from 2 to 4 Mbps further shortens packet airtime by approximately 52% and reduces transmission energy by approximately 43%. Finally, a first-in, first-out (FIFO)-triggered implantable loop recorder prototype demonstrates that jointly optimizing communication architecture, PHY configuration, and buffered transmission enables sleep-wake operation and reduces total system power consumption by approximately 60% compared with conventional BLE operation. These results identify minimizing radio active time as a key design principle for ultra-low-power on-demand sensing and provide practical guidance for battery-powered sensing systems.

eess.SY

Receding Horizon Multi-Agent Deceptive Path Planner

Deceptive path planning enables autonomous agents to obscure their true goals from observers by deviating from an expected optimal path. Prior work largely solves full-horizon, end-to-end optimization for single agents, which is expensive to recompute online and difficult to scale or adapt en route. We propose a unified framework for deceptive path planning using a Boltzmann distribution, computing over short-horizon candidate trajectories within a receding-horizon loop. By param- By iterating a user-defined cost that captures deception, resources, and smoothness, and optionally includes coupling terms between agents, the framework yields stochastic policies that balance the tradeoff between optimal paths and deceptive deviation. Policies are updated locally and do not require training. The level of deception and adherence to constraints can be dynamically tuned, enabling online adaptation to changes in goals and constraints such as obstacles. This step-by-step tuning opens the door to new forms of dynamic deception. Simulation studies demonstrate the flexibility of our approach, maintaining deception while adapting to environmental and constraint updates, avoiding the recomputation required by full-horizon methods, and supporting intuitive tuning via a small set of parameters

eess.SY