Search arXivSearch

arXiv · 2007.07207

Applying Dynamic Training-Subset Selection Methods Using Genetic Programming for Forecasting Implied Volatility

Abstract

Volatility is a key variable in option pricing, trading and hedging strategies. The purpose of this paper is to improve the accuracy of forecasting implied volatility using an extension of genetic programming (GP) by means of dynamic training-subset selection methods. These methods manipulate the training data in order to improve the out of sample patterns fitting. When applied with the static subset selection method using a single training data sample, GP could generate forecasting models which are not adapted to some out of sample fitness cases. In order to improve the predictive accuracy of generated GP patterns, dynamic subset selection methods are introduced to the GP algorithm allowing a regular change of the training sample during evolution. Four dynamic training-subset selection methods are proposed based on random, sequential or adaptive subset selection. The latest approach uses an adaptive subset weight measuring the sample difficulty according to the fitness cases errors. Using real data from SP500 index options, these techniques are compared to the static subset selection method. Based on MSE total and percentage of non fitted observations, results show that the dynamic approach improves the forecasting performance of the generated GP models, specially those obtained from the adaptive random training subset selection method applied to the whole set of training samples.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sana Ben Hamida, Wafa Abdelmalek, Fathi Abid. 2020-06-29. Applying Dynamic Training-Subset Selection Methods Using Genetic Programming for Forecasting Implied Volatility. https://arxiv.org/abs/2007.07207

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Prediction Markets Beat the Weather Forecast on Tomorrow's High Temperature

The sooner we receive information, and the more accurate it is, the better planning decisions we can make. Every day, prediction markets let anyone bet on tomorrow's high temperature in cities around the world, creating a market-implied forecast built on dispersed information. We use the past five years of market data from the Kalshi exchange for seven American cities to extract, hour by hour, the market-implied forecast. We use this forecast as a measuring instrument to see how much information about the temperature the market makes public before the public forecasting system does. We race it against the leading American and European weather forecasts. In six of the seven cities we study, the market beats the most accurate single public forecast, the National Blend of Models (NBM). Aggregating every city-day, at the end of the market's first hour of trading it beats the best single public product by about 10 percent in root-mean-square error, and holds its lead through the day, overnight, and into the target day. Looking at how the forecasts move over time, we find the National Blend travels four times further toward the market between its postings than the market travels toward the NBM. The market does not react to new weather forecast updates; instead, the forecast slowly publishes information that the market had already shared publicly.

q-fin.GN

Firm Valuation When AI Shapes the Business Model: A Milestone-Based Real-Options Framework for the AI Valuation Uncertainty Problem

Standard valuation methods, including discounted cash flow, the income approach standard IDW S 1 of the Institute of Public Auditors in Germany, and market multiples, compress milestone probabilities, continuation options, and risk shifts into opaque aggregate parameters; none provides a structured protocol for decomposing AI integration into auditable option-level assumptions. We propose an industry-agnostic taxonomy separating AI Integrators from AI Providers. AI Integrators are further classified by their Integration Depth Level, ranging from no integration to AI at the core of the product or process. A milestone-gated real-options overlay decomposes milestone state value into five components, and an Analytic Hierarchy Process-based Success Readiness Index derives per-option probabilities from structured pairwise comparisons for scenario analysis. Applied to an AI-native energy software-as-a-service firm, the framework yields a coherent valuation band traceable to identifiable option-level assumptions. Risk concentrates in later-stage continuation options, matching the structural prediction for AI Providers. The protocol applies across the firm lifecycle, including mergers and acquisitions due diligence. The case is a single-firm demonstration of protocol coherence, not empirical validation; multi-case testing against realised post-exit valuations is left to future research.

q-fin.GN

Machine Learning Classification and Portfolio Construction: Does the Loss Function Matter?

Classification outperforms regression across matched machine learning models in portfolio construction. A stacking ensemble of gradient boosted trees, random forest, and neural network yields a value-weighted annualized Sharpe ratio of 2.08 for classification and 1.39 for regression. This outperformance strengthens with class granularity and persists across subsamples and after transaction costs. Spanning tests show that classification retains economically large alphas after we control for regression, whereas regression alphas shrink substantially once we control for classification. These results indicate that classification extracts more return information than matched regression. Our diagnostics trace classification's advantage to more precise separation of return deciles.

q-fin.GN