Search arXiv⌕ Search

arXiv · 2610.02947

FARM: Fundamental Agentic Reward Model For Multi-task Wireless Network Optimization

Abstract

Future wireless networks require learning agents to adapt across heterogeneous channel conditions, traffic patterns, quality-of-service (QoS) requirements, objectives, and operational constraints. Reusing decision knowledge across such tasks is challenging because conventional multi-task and transfer reinforcement learning methods primarily share or transfer policies, coupling transferable knowledge with task-dependent action mappings. This paper proposes FARM (Fundamental Agentic Reward Model for Multi-task Wireless Network Optimization), a reward-space transfer framework that shifts cross-task knowledge reuse from policy space to trajectory-level decision evaluation. FARM introduces an Agentic Reward Model (ARM) that learns a task-conditioned reward prior from heterogeneous source-task trajectories and provides auxiliary guidance for task-specific policy optimization. In Stage I, ARM jointly models task conditions, temporal trajectory dependencies, and objective-dependent reward structures while each source task retains its own controller. In Stage II, the learned reward prior is frozen and reused to guide the adaptation of a target-specific controller for previously unseen tasks, without transferring source-task policies. Experiments on heterogeneous multi-access edge computing (MEC) tasks show that FARM achieves a mean late-stage gain of 29.8% over Single-task SAC on unseen Rate-Latency targets, compared with 16.1% for CRA Transfer, and reaches a 46.2% gain on the moderate-OOD FAR-M case. Further analysis shows that both Mamba and Transformer trajectory encoders support Reward-Space Transfer, while Mamba provides improved robustness as longer history dependencies are introduced.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Feiran You, Changxu Ni, Haozhe Ma, Jun Li, Hongyang Du. 2026-10-02. FARM: Fundamental Agentic Reward Model For Multi-task Wireless Network Optimization. https://arxiv.org/abs/2610.02947

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Policy-Level Recursive Self-Improvement for Embodied AI with a Criticality World Model

Despite advances in policy pretraining, embodied AI systems can plateau during task-specific fine-tuning as uniform scenario collection encounters fewer of the remaining failures. We address this problem with a policy-level recursive self-improvement loop: execution outcomes train a criticality world model, whose risk scores guide scenario collection for specialist training. To account for the resulting shift in scenario frequencies, weighted resampling approximately corrects the collection bias. Because specialized updates can weaken nominal behavior, a risk-based gate selects between the specialist and a frozen nominal policy at inference. We extend this procedure to reinforcement learning and behavior cloning. A sampling analysis characterizes the correction, while a controlled study examines the trade-off between critical and nominal coverage. Across five embodied domains, the composed systems reduce failure rates by 59--67% for locomotion and manipulation and 8--25% for VLA benchmarks relative to their respective baselines.

eess.SY↗

Closed-Loop Refinement and Execution for Learned Driving Planners

Learning-based driving planners are usually trained and evaluated in open loop against logged trajectories. In closed loop, a trajectory with small displacement error can still stall the vehicle, steer it into a conflict with surrounding agents, or be executed with abrupt braking. We introduce Closed-Loop Refinement and Execution (CLRE), a hierarchical receding-horizon control framework designed to mitigate these failure modes while leaving the upstream planner frozen and adding no new learned model. The upper layer treats the nominal trajectory as a reference and solves a finite-horizon optimal control problem that trades route progress against interaction with predicted agents. Solving it from several initializations gives a candidate set, and a prediction-conditioned oriented-bounding-box (OBB) feasibility test retains only candidates whose minimum predicted OBB clearance over the horizon meets a threshold. The lower layer executes the lowest-cost survivor, or a route-centerline backup when none remains, through the tracking controller supplied with the planner, augmented by a range-based speed bound and a saturated proportional braking law. In closed-loop simulation on the 220-route Bench2Drive validation set with VAD as the upstream planner, CLRE raises the driving score from 42.26 to 55.89 and route completion from 55.69 to 71.51, and reduces collision events from 117 to 97.

eess.SY↗

A Dynamic Generalized Kalman Consensus Filter for Switching Sensor Networks

Distributed state estimation is critical for applications such as surveillance, autonomous navigation, and wide-area monitoring, where sensor agents must cooperatively track targets using only local measurements and neighbor-to-neighbor communication. Existing distributed filters have been shown to achieve accurate estimation even under sparse inter-agent communication and limited sensing ranges. However, many of these methods rely on consensus parameters that depend on global properties of the communication graph, such as the maximum degree of the graph, and are therefore sensitive to changes in network topology. This limitation is particularly significant in sensor networks with mobile agents, where communication links change over time. This paper presents a Dynamic Generalized Kalman Consensus Filter for target tracking in sensor networks with switching communication topologies. The proposed algorithm computes information-based consensus weights using only locally available quantities, eliminating the need for global network parameters. Numerical simulations demonstrate that the proposed algorithm maintains estimation accuracy under switching network topologies and outperforms existing distributed filters in the given tracking problem.

eess.SY↗