Search arXivSearch

arXiv · 1307.8308

Is it possible to predict long-term success with k-NN? Case Study of four market indices (FTSE100, DAX, HANGSENG, NASDAQ)

Abstract

This case study tests the possibility of prediction for "success" (or "winner") components of four stock & shares market indices in a time period of three years from 02-Jul-2009 to 29-Jun-2012.We compare their performance ain two time frames: initial frame three months at the beginning (02/06/2009-30/09/2009) and the final three month frame (02/04/2012-29/06/2012). To label the components, average price ratio between two time frames in descending order is computed. The average price ratio is defined as the ratio between the mean prices of the beginning and final time period. The "winner" components are referred to the top one third of total components in the same order as average price ratio it means the mean price of final time period is relatively higher than the beginning time period. The "loser" components are referred to the last one third of total components in the same order as they have higher mean prices of beginning time period. We analyse, is there any information about the winner-looser separation in the initial fragments of the daily closing prices log-returns time series. The Leave-One-Out Cross-Validation with k-NN algorithm is applied on the daily log-return of components using a distance and proximity in the experiment. By looking at the error analysis, it shows that for HANGSENG and DAX index, there are clear signs of possibility to evaluate the probability of long-term success. The correlation distance matrix histograms and 2-D/3-D elastic maps generated from ViDaExpert show that the winner components are closer to each other and winner/loser components are separable on elastic maps for HANGSENG and DAX index while for the negative possibility indices, there is no sign of separation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Y. Shi, A. N. Gorban, T. Y. Yang. 2013-07-31. Is it possible to predict long-term success with k-NN? Case Study of four market indices (FTSE100, DAX, HANGSENG, NASDAQ). https://doi.org/10.1088/1742-6596%2F490%2F1%2F012082

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Endogenous Constraint: Hysteresis, Stagflation, and the Structural Inhibition of Monetary Velocity in the Bitcoin Network (2016-2025)

Bitcoin operates as a macroeconomic paradox: it combines a strictly predetermined, inelastic monetary issuance schedule with a stochastic, highly elastic demand for scarce block space. This paper empirically validates the Endogenous Constraint Hypothesis, positing that protocol-level throughput limits generate a non-linear negative feedback loop between network friction and base-layer monetary velocity. Using a verified Transaction Cost Index (TCI) derived from Blockchain.com on-chain data and Hansen's (2000) threshold regression, we identify a definitive structural break at the 90th percentile of friction (TCI ~ 1.63). The analysis reveals a bifurcation in network utility: while the network exhibits robust velocity growth of +15.44% during normal regimes, this collapses to +6.06% during shock regimes, yielding a statistically significant Net Utility Contraction of -9.39% (p = 0.012). Crucially, Instrumental Variable (IV) tests utilizing Hashrate Variation as a supply-side instrument fail to detect a significant relationship in a linear specification (p=0.196), confirming that the velocity constraint is strictly a regime-switching phenomenon rather than a continuous linear function. Furthermore, we document a "Crypto Multiplier" inversion: high friction correlates with a +8.03% increase in capital concentration per entity, suggesting that congestion forces a substitution from active velocity to speculative hoarding.

q-fin.ST

Wasserstein-Barycentric Interaction Fields for Spatial Factor Models: Evidence from Language-Model Representations

Spatial asset-pricing models take the structure of inter-firm interaction as given. We infer that structure from firms' information environments using language-model representations. Each firm is represented as a distribution of news-article embeddings, and a target-anchored Wasserstein barycentric reconstruction selects, for every firm, the weighted combination of other firms whose information footprints jointly reconstruct its own. The resulting directed peer field enters a quadratic exposure-adjustment model in which the spatial coefficient indexes alignment with information peers relative to stand-alone exposure. Using fields built from 2018-2022 news and frozen before 2023-2026 returns, we find that the constructed field organizes cross-sectional return dependence beyond the Fama-French five factors and momentum and raises the held-out mean Gaussian quasi-log score relative to a matched factor-only model. Because factor betas are unchanged, the gain lies in residual covariance. The field outperforms pairwise distance weighting and equal weighting of the same peers, and remains incrementally informative beside persistent news co-mentions under the primary factor-conditioned specification. Linear and quadratic transport generate nearly identical peer-return signals and equivalent held-out predictive performance. The barycentric-proximity ordering persists across alternative embedding models, and a pre-period encoder preserves the held-out advantage under the primary specification. Language-model representations thus serve as a measurement instrument for latent inter-firm information structure in capital markets.

q-fin.ST

Stealing profits: Spread-based temporal hierarchy forecasting for day-ahead electricity markets

Day-ahead electricity price forecasts support trading and storage decisions, but for battery arbitrage predicting intraday price spreads is more relevant than predicting individual hourly prices. Here we show that a temporal hierarchy forecasting (THieF) framework that jointly reconciles forecasts of hourly electricity prices and all intraday price spreads consistently improves performance across two major European electricity markets and three different forecasting architectures. Using five years of out-of-sample data from Germany and Spain, we obtain accuracy improvements of up to 19.7% and profit gains of up to 10.4% relative to unreconciled hourly price forecasts. The gains persist even for a highly accurate pretrained TabPFN foundation model. Our results demonstrate that exploiting coherent relationships between economically relevant forecasting targets can improve both predictive accuracy and decision value, and that better statistical forecasts do not necessarily imply better economic decisions.

q-fin.ST