Search arXiv⌕ Search

arXiv · 2609.35359

From Cointegration to Out-of-Sample Failure: A Pairs-Trading Case Study on PEP-KO

Abstract

This paper examines whether a cointegration-based pairs trading strategy between PepsiCo and The Coca-Cola Company is statistically robust and economically exploitable. We first test for cointegration and estimate the spread's mean-reversion dynamics over 2013-2018, then hold these statistical parameters fixed and optimise a threshold-based trading strategy in-sample over 2018-2023. Robustness is assessed through transaction-cost and parameter sensitivity tests, walk-forward validation, and Adjusted and Deflated Sharpe Ratios. The strategy is then evaluated out-of-sample from 2023 to the present, including an analysis of time-varying hedge ratios using rolling OLS and a Kalman filter. The results show that weakening mean-reversion dynamics in the spread undermine the effectiveness of the strategy out-of-sample.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Davide Graziano. 2026-09-28. From Cointegration to Out-of-Sample Failure: A Pairs-Trading Case Study on PEP-KO. https://arxiv.org/abs/2609.35359

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Variance-Corrected Multi-Asset Equity Simulation with Hybrid Hidden Markov Marginals

Synthetic multi-asset equity data must reproduce each asset's return distribution and its relationship with the market. Reusing a generator fitted to full asset returns creates a problem: adding its draws to a market factor counts market variance twice. We derived a correction that centers and rescales each draw over a fixed horizon before adding the market factor, allowing reuse without fitting a second generator to regression residuals. We tested the correction on 423 non-market assets in a 424-asset United States equity and exchange-traded-fund universe, using hidden Markov generators with heavy-tailed emissions. The corrected paths retained heavy tails and recovered the calibrated market loadings with low error. On 416 complete asset histories held out from 2025, the correction improved the mean Kolmogorov-Smirnov pass rate over naive composition and brought the median ratio of synthetic to observed variance close to one. Its one-day left-tail 99% Value-at-Risk exceedance rate, using thresholds pooled across separate simulated paths, was 1.12%, compared with 1.16% for the residual-fit hidden Markov comparator and a nominal rate of 1%. We also tested a jump-duration mechanism that extended visits to extreme-return states and improved volatility clustering for the broad-market exchange-traded fund with ticker SPY. The multi-asset correction remained effective with jumps enabled, but transferring the SPY jump settings improved temporal fit in training and worsened it in the 2025 holdout. The method provides a way to reuse fitted asset generators for daily-fit comparisons. Its zero-sum residual constraint removes terminal residual uncertainty, and unmodeled dependence among asset-specific residuals can lead to overstated rebalancing returns. Code, cached inputs, result summaries, and instructions for fitting the per-asset models accompany the paper.

q-fin.ST↗

Insider Purchases Far Below the 52-Week High: Decomposing the Disclosure Reaction in Microcap Equities

Purchases reported under transaction code P on SEC Form 4 by insiders of U.S. equities with an estimated filing-date capitalization of USD 30 million to USD 500 million (13,534 lines, 1,192 issuers, 2018-2024) are followed by a first-day abnormal return that rises steeply with the stock's distance below its 52-week high: 4.13% in the quintile farthest below the high against 0.86% nearest it (two-way clustered t = 9.77); random non-event days of the same issuers show 0.14%. Five tests with decision rules fixed in advance characterize the gradient. Most of it is scale: the beaten-down stocks are 3.10 times as volatile, and with a full set of controls the raw gap fails its pre-specified bar (0.94 points, t = 2.11). Per unit of the stock's own volatility the reaction is 2.51 times as large far below the high (t = 7.49), 1.28 to 3.50 on other estimators, though a variance-weighted slope shows none. The gradient is larger than for insider sales by the same issuers and for positive earnings surprises as a class; against the strongest surprises the difference is imprecise. Dropping purchases with a concurrent 8-K leaves the raw gradient intact (t = 7.13), but the per-risk gradient no longer clears the controls (t = 2.52). The reaction runs for two to three sessions; the 29-day drift is imprecise (two-way t = 1.70) and a calendar-time portfolio that skips the first day earns no significant alpha. The analysis quantifies sensitivity to price adjustment and benchmark specification: mixing price bases misassigns run-up buckets, and carrying the estimation-window intercept supplies 57% of the 30-day gradient under that benchmark. A gradient-boosting classifier (test AUC 0.676) is indistinguishable from logistic regression. A separate large-cap extension schedules USD 29,075,559 a year of buyer-cluster flow but fails every matched-comparison gate.

q-fin.ST↗

Cross-Market Alpha: Testing Short-Term Trading Factors in the U.S. Market via Double-Selection LASSO

We test whether 168 short-horizon price-volume signals from the Alpha191 library, originally developed for China's retail-dominated A-share market, contain pricing information for S&P 500 stocks from 2002 to 2022 beyond 153 established U.S. factors. Using the double-selection LASSO of Feng et al. (2020), 17 signals receive significant stochastic discount factor (SDF) loadings in the baseline test-asset design. Their robustness is uneven. Only three signals (a multi-horizon moving-average ratio, a directional-pressure ratio, and a price-gap correlation) remain significant with a finer test-asset grid and under Elastic Net and principal-component control selection; six more pass most checks, and the remaining eight depend on the specification. Robust signals are concentrated in volume-price interaction and short-term mean reversion, whereas volatility-based signals are fragile.

q-fin.ST↗