Search arXivSearch

arXiv · 2604.00031

Decomposable Reward Modeling and Realistic Environment Design for Reinforcement Learning-Based Forex Trading

Abstract

Applying reinforcement learning (RL) to foreign exchange (Forex) trading remains challenging because realistic environments, well-defined reward functions, and expressive action spaces must be satisfied simultaneously, yet many prior studies rely on simplified simulators, single scalar rewards, and restricted action representations, limiting both interpretability and practical relevance. This paper presents a modular RL framework designed to address these limitations through three tightly integrated components: a friction-aware execution engine that enforces strict anti-lookahead semantics, with observations at time t, execution at time t+1, and mark-to-market at time t+1, while incorporating realistic costs such as spread, commission, slippage, rollover financing, and margin-triggered liquidation; a decomposable 11-component reward architecture with fixed weights and per-step diagnostic logging to enable systematic ablation and component-level attribution; and a 10-action discrete interface with legal-action masking that encodes explicit trading primitives while enforcing margin-aware feasibility constraints. Empirical evaluation on EURUSD focuses on learning dynamics rather than generalization and reveals strongly non-monotonic reward interactions, where additional penalties do not reliably improve outcomes; the full reward configuration achieves the highest training Sharpe (0.765) and cumulative return (57.09 percent). The expanded action space increases return but also turnover and reduces Sharpe relative to a conservative 3-action baseline, indicating a return-activity trade-off under a fixed training budget, while scaling-enabled variants consistently reduce drawdown, with the combined configuration achieving the strongest endpoint performance.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nabeel Ahmad Saidd. 2026-03-20. Decomposable Reward Modeling and Realistic Environment Design for Reinforcement Learning-Based Forex Trading. https://arxiv.org/abs/2604.00031

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Virtue of Sparsity in Complexity

Sparsity or complexity? In modern high-dimensional asset pricing, these are often viewed as competing principles: recent empirical evidence favors richer models, while economic intuition has long favored parsimony. We reconcile this tension by distinguishing capacity sparsity-restrictions on effective model capacity-from factor sparsity-the parsimonious structure of priced risks. Revisiting the benchmark empirical design of Didisheim et al. (2025), we combine nonlinear feature expansions with basis pursuit, using column generation and GPU acceleration to scale estimation to 432 million candidate factors. Reaching this scale reveals a reversal in out-of-sample performance: sparse portfolios trail dense ridgeless benchmarks at lower complexity but achieve a higher Sharpe ratio and lower pricing error at the largest candidate set. Capacity expansion and factor sparsity are therefore complements: enlarging the candidate space allows a parsimonious pricing kernel to outperform its dense counterpart.

q-fin.GN

Prediction Markets Beat the Weather Forecast on Tomorrow's High Temperature

The sooner we receive information, and the more accurate it is, the better planning decisions we can make. Every day, prediction markets let anyone bet on tomorrow's high temperature in cities around the world, creating a market-implied forecast built on dispersed information. We use the past five years of market data from the Kalshi exchange for seven American cities to extract, hour by hour, the market-implied forecast. We use this forecast as a measuring instrument to see how much information about the temperature the market makes public before the public forecasting system does. We race it against the leading American and European weather forecasts. In six of the seven cities we study, the market beats the most accurate single public forecast, the National Blend of Models (NBM). Aggregating every city-day, at the end of the market's first hour of trading it beats the best single public product by about 10 percent in root-mean-square error, and holds its lead through the day, overnight, and into the target day. Looking at how the forecasts move over time, we find the National Blend travels four times further toward the market between its postings than the market travels toward the NBM. The market does not react to new weather forecast updates; instead, the forecast slowly publishes information that the market had already shared publicly.

q-fin.GN

Firm Valuation When AI Shapes the Business Model: A Milestone-Based Real-Options Framework for the AI Valuation Uncertainty Problem

Standard valuation methods, including discounted cash flow, the income approach standard IDW S 1 of the Institute of Public Auditors in Germany, and market multiples, compress milestone probabilities, continuation options, and risk shifts into opaque aggregate parameters; none provides a structured protocol for decomposing AI integration into auditable option-level assumptions. We propose an industry-agnostic taxonomy separating AI Integrators from AI Providers. AI Integrators are further classified by their Integration Depth Level, ranging from no integration to AI at the core of the product or process. A milestone-gated real-options overlay decomposes milestone state value into five components, and an Analytic Hierarchy Process-based Success Readiness Index derives per-option probabilities from structured pairwise comparisons for scenario analysis. Applied to an AI-native energy software-as-a-service firm, the framework yields a coherent valuation band traceable to identifiable option-level assumptions. Risk concentrates in later-stage continuation options, matching the structural prediction for AI Providers. The protocol applies across the firm lifecycle, including mergers and acquisitions due diligence. The case is a single-firm demonstration of protocol coherence, not empirical validation; multi-case testing against realised post-exit valuations is left to future research.

q-fin.GN