Search arXiv⌕ Search

arXiv · 0704.0838

Universal Source Coding for Monotonic and Fast Decaying Monotonic Distributions

Abstract

We study universal compression of sequences generated by monotonic distributions. We show that for a monotonic distribution over an alphabet of size $k$, each probability parameter costs essentially $0.5 \log (n/k^3)$ bits, where $n$ is the coded sequence length, as long as $k = o(n^{1/3})$. Otherwise, for $k = O(n)$, the total average sequence redundancy is $O(n^{1/3+ε})$ bits overall. We then show that there exists a sub-class of monotonic distributions over infinite alphabets for which redundancy of $O(n^{1/3+ε})$ bits overall is still achievable. This class contains fast decaying distributions, including many distributions over the integers and geometric distributions. For some slower decays, including other distributions over the integers, redundancy of $o(n)$ bits overall is achievable, where a method to compute specific redundancy rates for such distributions is derived. The results are specifically true for finite entropy monotonic distributions. Finally, we study individual sequence redundancy behavior assuming a sequence is governed by a monotonic distribution. We show that for sequences whose empirical distributions are monotonic, individual redundancy bounds similar to those in the average case can be obtained. However, even if the monotonicity in the empirical distribution is violated, diminishing per symbol individual sequence redundancies with respect to the monotonic maximum likelihood description length may still be achievable.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Gil I. Shamir. 2007-04-06. Universal Source Coding for Monotonic and Fast Decaying Monotonic Distributions. https://arxiv.org/abs/0704.0838

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Orbit Reduction and Learned Run Distributions for Finite-Blocklength Binary Deletion Channels

In DNA data storage, racetrack memories, and packet networks, data are written as many short strands of fixed length. The receiver knows where each strand begins and ends, and deletions occur only inside a strand. The relevant limit for such systems is the block capacity $C_N(d)$, the largest mutual information between a length-$N$ input and the output of a binary deletion channel with deletion probability $d$. Because the boundaries are known, $C_N(d)/N$ is never smaller than the classical capacity $C(d)$. Computing $C_N(d)$ is hard because the input takes $2^N$ values. For $0\le d<1$, we show that the optimal input is unique, gives positive probability to every string, and is unchanged by complementing or reversing the strings. We then introduce the optimized run distribution (ORD), which assigns one probability to each number of runs; it is optimal for $N\le3$ and is solved with a numerical optimality certificate up to $N=16$. For longer strands, an exact recursion for the output distribution gives unbiased rate estimates, and an empirical Bernstein inequality turns them into confidence intervals. Within the statistical precision, a one-parameter Markov input performs as well as the best inputs found. For strands of 100-200 bits, known boundaries increase the rate by up to 0.026 bits per symbol beyond the best certified upper bound on the capacity of the unsegmented channel. Via Fano's inequality, the block capacity also bounds the rate of any code that uses a single strand; in the tabulated case, this bound is tighter than the best known finite-length bound for strands longer than about 90 bits. Conversely, block rates yield lower bounds on $C(d)$; at $d=0.1$, the bound is within $1.1\times10^{-4}$ bits/use of the best certified lower bound. Finally, InfoNCE estimates, even with the optimal critic, empirically lose most of their ability to rank inputs as the strand length grows.

cs.IT↗

Improved Lower Bounds on the Capacity of the Binary Deletion Channel via a Learning Approach to Run-Length Inputs

The best constructive lower bounds on the capacity of the binary deletion channel come from random codes with independent run lengths, yet the two strongest such bounds, due to Drinea and Mitzenmacher and to Venkataramanan, Tatikonda, and Ramchandran, have been evaluated mainly for geometric or low-parameter run-length laws. We let a learning algorithm choose the run-length law freely, which raises three challenges. First, the Drinea-Mitzenmacher functional is an infinite sum over the ways deletions merge runs; we show that it depends on the law only through its mean, its full-deletion probability, and bilinear forms in the law and its renewal weights, so that gradients of truncations are exact and every truncation can only lower the bound. Second, for non-Markov inputs the output is no longer Markov and the Venkataramanan-Tatikonda-Ramchandran analysis breaks down; we show that the residual length of the current input run turns the output into a hidden Markov chain, which extends the bound to every finite-support law. Third, its correction term counts only output runs formed from three input runs; we prove a larger correction that accounts for every output run formed by several input runs, which improves the published bound even for truncated geometric laws. Certified by interval arithmetic, the new bounds exceed all previously published deterministic lower bounds at every tabulated deletion probability, by up to $6.86\times10^{-3}$ bits per channel use and 6.8%. At large deletion probability the learned laws concentrate on run-length clusters with survivor counts about two standard deviations apart, like a pulse-amplitude constellation.

cs.IT↗

Complementary Waveform Coordination for Doppler-Resilient MIMO Radar

Complementary waveform libraries yield impulse-like aggregate delay responses at zero Doppler, but pulse-dependent Doppler phases degrade their sidelobe cancellation. This paper develops a linear-algebraic formulation for coordinating a (D)-ary transmission schedule and slow-time receive weights for paraunitary MIMO waveform libraries. Using a modal decomposition of the CPI ambiguity matrix, prescribed-order Doppler-nulling conditions for the (D-1) nonzero modes are expressed as homogeneous linear constraints on the receive weights for a fixed schedule. This representation gives the SNR-optimal fixed-schedule weights by subspace projection and reduces the remaining design to a discrete schedule search. Degrees-of-freedom and modal-growth analyses relate the design to the CPI length, null order, and waveform dimension. Numerical examples compare the resulting designs with established binary constructions and demonstrate Doppler suppression across the full ambiguity matrix of a four-waveform configuration.

cs.IT↗