arXiv · 2609.27783
Optimal Weighting of Training Data in Adaptive Coding
Abstract
Adaptive coders are commonly primed with training data, turned into context tables or shipped dictionaries. The training data generally do not match the source being encoded, and the coder must decide how much to trust them. We make that trust a design variable: a Krichevsky--Trofimov (KT) estimator over an $m$-ary alphabet whose training counts are scaled by $ξ\in[0,1]$. For a training sequence of length $\ell$ at Kullback--Leibler divergence $D$ nats per symbol from the message source, the redundancy-minimizing weight is $ξ^*=d/(2\ell D+d)$, $d=m-1$. The effective training length $ξ^*\ell$ follows a harmonic law: any mismatch caps the usable training information at $d/(2D)$ symbols, however much was collected. Three implementable selectors supply the unknown $D$: offline, from the spread of the training set; by a plug-in loop, from the decoded prefix; and by a twice-universal mixture, with no estimation at all. On the ten English texts of the Calgary and Canterbury corpora, weighting removes up to $22.7\%$ of the redundancy of the classical KT code, the $ξ=0$ endpoint, and beats both endpoints on every file. An open-source implementation is provided.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuriy Reznik. 2026-08-17. Optimal Weighting of Training Data in Adaptive Coding. https://arxiv.org/abs/2609.27783
Cite the original work for its findings. Save a collection to share your selection of sources.