Search arXivSearch

arXiv · 2608.23053

tse_tick: A Python Library for Parsing and Querying Nikkei NEEDS Tick Data from the Tokyo Stock Exchange

Abstract

Tick-level trade-and-quote data for the Tokyo Stock Exchange is distributed through the Nikkei NEEDS service as thousands of zipped CSV archives spanning four data types with era-dependent schemas and Japanese-language layouts. We present tse_tick, an open-source Python library that converts these raw archives into clean, typed Polars DataFrames and a Hive-partitioned Parquet store queryable through DuckDB. The library offers two access paths sharing one parse-and-clean core: a one-shot reader that returns a ticker- and time-filtered DataFrame directly from raw ZIP files, and a two-stage ingest-then-query pipeline with resume-safe, memory-aware parallel ingestion, part-pruning, and a materialized intraday time key for row-group pruning. The engineering, more than the parsing, is what the library contributes: ingestion runs in per-date atomic units whose completion is recorded by coverage markers rather than inferred from file existence, writes stream in bounded morsels so that peak memory is independent of trading-day size (24.5 GB to 2.4 GB on the worst measured day), a RAM-aware process pool sizes itself to available memory, and part-pruning opens only the archive parts a ticker can occupy. Full English and Japanese column definitions ship for all four types, and a translation layer maps yfinance, Polygon, and ccxt names onto their tse_tick equivalents. In benchmarks on a commodity 16-thread workstation, parsing a representative 4.8-million-row archive part, one of a trading day's nine parts, is 59.8x faster than the original pandas prototype (34.3x against an engine-matched pandas baseline), and a single-ticker time-window query from the store completes roughly 410x faster than a pandas scan of the equivalent CSV. tse_tick is available on PyPI (pip install tse-tick) under the MIT license.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kazumi Li, Masataka Hayashi, Teruo Nakatsuma, Peter Romero. 2026-08-24. tse_tick: A Python Library for Parsing and Querying Nikkei NEEDS Tick Data from the Tokyo Stock Exchange. https://arxiv.org/abs/2608.23053

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Adapting the Actor Model of Concurrency for High-Frequency Trading: Synchronous Message Delivery (fast_send) and a Tick-to-Book Latency Study

The actor model - state isolation, data-race freedom, deadlock resistance, and sequential single-message reasoning - has long been dismissed as unsuitable for high-frequency trading (HFT): actors seem to imply many threads, a mailbox per actor, and a heap-allocated message plus a context switch per interaction, overhead incompatible with a microsecond budget. This paper argues the dismissal is wrong for co-located actors, and supports it both analytically and with a deployed, measured implementation: kaspar-hft, an open-source C++20 framework. Four extensions adapt the model for HFT: fast_send, a synchronous delivery mechanism in which the sending thread runs the receiver's handler inline and returns the reply as a value; actor groups, which co-schedule actors on one thread behind a shared mailbox; per-actor selectable mailbox queues; and a memory pool. fast_send has receiver transparency: the handler cannot tell whether delivery was synchronous or asynchronous, or which thread runs it. A grouped synchronous chain runs on one thread, cutting scheduler context switches from O(N) to O(1), and a thread-local call-chain test catches cyclic invocation before any lock is taken. Microbenchmarks put the synchronous round trip at tens of nanoseconds. On a live CME market-data feed (ES, NQ, ZN futures), socket-to-book latency decomposes into a ~7 microsecond decode-and-book floor plus a per-message slope; the framework's own contribution is under 1% of the floor. The tail is set not by the actor machinery but by the market's non-Poisson, clustered arrival process, characterized in a companion paper. The shared-queue group also yields a production/simulation duality: the same actor code runs unchanged in live trading and deterministic backtest.

q-fin.TR

Post-Rejection Follow-up Sampling: Measuring Outcomes of Rejected Decisions in Algorithmic DEX Trading

Filter-gated algorithmic trading systems on decentralised exchanges reject most candidate tokens they evaluate, yet the observed forward market trajectory of those rejected candidates is rarely measured on the same live venue that produced the rejection. This paper introduces Post-Rejection Follow-up Sampling (PRFS), an observational measurement methodology in which a separate tracking subsystem samples each rejected token's price and liquidity from the same live oracle path used by the rejecting scanner, at a scheduled cadence, from the moment of rejection out to a fixed analytic horizon. The methodology is defined through a formal specification of the observation window and the reason-attribution rule, a reference implementation (prfs v2.0.0) with executable coverage and reason-ledger tests, and a secondary cross-check implementation that re-derives every headline number through a distinct code path. The companion dataset contains 67,000 forward-observation rows across 2,997 rejection events collected on a single Solana automated-market-maker sub-venue during a launch-dynamics window of 8.63 calendar days across 457 unique mints. Under the primary reason-aware three-field event key, 1,455 events (48.55 percent) receive a matched forward observation; under the reason-agnostic two-field diagnostic key, 1,641 events (54.75 percent) match. Mint-level coverage is 100 percent. The 253-row reason-conflict ledger and 8,593 unaligned outcome rows are disclosed, and coverage is shown to be filter-associated (chi-square(6) = 401.29) and source-associated (z = 14.78). No causal or counterfactual claim of a hypothetical acceptance outcome is made; the reported estimand is the observed post-rejection forward return.

q-fin.TR

A Validated Volatility-Volume-Gap Classifier for Regime Identification in MNQ Intraday Data

This paper builds and tests a day-classification system for MNQ (Micro E-Mini Nasdaq 100) futures based on three simultaneously elevated pre-market conditions: absolute overnight gap, absolute first-30-minute return, and first-bar volume relative to a 20-day rolling baseline. The Volatility-Volume-Gap (VVG) classifier is evaluated on 947 trading days of five-minute data from 2021-2025, with all thresholds computed on an expanding window to prevent lookahead bias. The classifier activates on roughly 4.4% of sessions (40 days). Those days exhibit measurably distinct behavior: 77.6% reverse from their intraday peak before the close, mean peak-to-close giveback of 11.73 points, and a 25.6 basis point next-day return spread versus non-classifier days. Year-by-year analysis reveals substantial path heterogeneity -- 2024 classifier days closed at mean +40.74 points while 2025 crashed to -42.48 -- which is the core obstacle for any directional strategy. Eight directional configurations were tested. None passed. Best result: T = 1.46, mean net +7.80 points, 127 OOS trades, reversal entry with OLS regression filter. 2024 broke year stability. Binding constraints are the 40-day sample (roughly 10 per year) and regime-dependent intraday behavior that no fixed rule survives across all test years. The classifier is preserved as a research asset: it identifies a real behavioral phenomenon but cannot generate a deployable directional signal under current constraints.

q-fin.TR