Search arXivSearch

arXiv · math/0103007

Source Coding, Large Deviations, and Approximate Pattern Matching

Abstract

We present a development of parts of rate-distortion theory and pattern- matching algorithms for lossy data compression, centered around a lossy version of the Asymptotic Equipartition Property (AEP). This treatment closely parallels the corresponding development in lossless compression, a point of view that was advanced in an important paper of Wyner and Ziv in 1989. In the lossless case we review how the AEP underlies the analysis of the Lempel-Ziv algorithm by viewing it as a random code and reducing it to the idealized Shannon code. This also provides information about the redundancy of the Lempel-Ziv algorithm and about the asymptotic behavior of several relevant quantities. In the lossy case we give various versions of the statement of the generalized AEP and we outline the general methodology of its proof via large deviations. Its relationship with Barron's generalized AEP is also discussed. The lossy AEP is applied to: (i) prove strengthened versions of Shannon's source coding theorem and universal coding theorems; (ii) characterize the performance of mismatched codebooks; (iii) analyze the performance of pattern- matching algorithms for lossy compression; (iv) determine the first order asymptotics of waiting times (with distortion) between stationary processes; (v) characterize the best achievable rate of weighted codebooks as an optimal sphere-covering exponent. We then present a refinement to the lossy AEP and use it to: (i) prove second order coding theorems; (ii) characterize which sources are easier to compress; (iii) determine the second order asymptotics of waiting times; (iv) determine the precise asymptotic behavior of longest match-lengths. Extensions to random fields are also given.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

A. Dembo, I. Kontoyiannis. 2001-03-02. Source Coding, Large Deviations, and Approximate Pattern Matching. https://arxiv.org/abs/math/0103007

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Bounded weak solutions to cross-diffusion semiconductor model with electron-hole scattering

Semiconductor model is a system of parabolic partial differential equations with cross-diffusion phenomenon. Previous results showed that a weak solution exists and is not bounded in general. So semiconductor model was categorized as a cross-diffusion system without bounded weak solutions. In this work, we show that once the initial value is bounded, there exists a weak solution that is also bounded. The entropy method is a major tool in global existence analysis of cross-diffusion systems. We notice that traditional entropies in volume-filling cases may not provide required positive semi-definiteness result for the existence proof. In this situation, a transformation of variables technique has been applied. The product between Hessian matrix of the entropy and replacement diffusion matrix is positive semi-definite, then we apply the entropy method to show semiconductor model has a bounded weak solution.

math.PR

Global existence and uniqueness analysis of cross-diffusion multispecies chemotaxis system with volume-filling

The system of multispecies chemotaxis equations is a cross-diffusion system with volume-filling. In this work, we show that a weak solution of the two species chemotaxis system exists. The entropy method is a major tool in existence analysis of cross-diffusion systems. Previous investigations indicate that traditional entropies in volume-filling cases may not be able to provide required gradient estimates. In this situation, we upgrade existing matrix computation methods to derive gradient estimates. Due to the cross-diffusion phenomenon, the uniqueness of the weak solution to a cross-diffusion system is very difficult to prove in general. In this work, we apply the distance functional to show that when parameters of the chemotaxis system are identical, the weak solution is unique.

math.PR

Self-normalized scaled quadratic variation

The concept of a scaled quadratic variation was originally introduced by E. Gladyshev in 1961 for processes with Gaussian increments. Using certain deterministic scaling, arrived at from the covariance of the process, Gladyshev showed that the sum of scaled square increments along the dyadic partition sequence converges almost surely to a finite limit. In this paper, we propose a pathwise counterpart in which the deterministic normalization is replaced by a self-normalizing factor built from the $p$-th variation of the path along a given sequence of partitions. The resulting quantity requires no probabilistic assumption and no knowledge of a covariance structure, and its scale is both path-dependent and sensitive to the partition sequence. Under a mild regularity condition on the limiting $p$-th variation, we show that the self-normalized and the classical deterministic normalizations are comparable, and for fractional Brownian motion the two agree up to a multiplicative constant. We establish a switching behaviour in the index, and prove that for $p \ge 2$ the self-normalized scaled quadratic variation obeys a smooth-transformation formula under $C^2$ maps; at $p=2$ this recovers the known transformation rule for quadratic variation. Since only squared increments are scaled, the construction polarizes, yielding a matrix-valued scaled quadratic variation for every $p \geq 1$ for $\mathbb R^d$ valued paths. We conclude with examples beyond the Gaussian setting.

math.PR