Search arXiv⌕ Search

arXiv · 0710.1203

Semantic distillation: a method for clustering objects by their contextual specificity

Abstract

Techniques for data-mining, latent semantic analysis, contextual search of databases, etc. have long ago been developed by computer scientists working on information retrieval (IR). Experimental scientists, from all disciplines, having to analyse large collections of raw experimental data (astronomical, physical, biological, etc.) have developed powerful methods for their statistical analysis and for clustering, categorising, and classifying objects. Finally, physicists have developed a theory of quantum measurement, unifying the logical, algebraic, and probabilistic aspects of queries into a single formalism. The purpose of this paper is twofold: first to show that when formulated at an abstract level, problems from IR, from statistical data analysis, and from physical measurement theories are very similar and hence can profitably be cross-fertilised, and, secondly, to propose a novel method of fuzzy hierarchical clustering, termed \textit{semantic distillation} -- strongly inspired from the theory of quantum measurement --, we developed to analyse raw data coming from various types of experiments on DNA arrays. We illustrate the method by analysing DNA arrays experiments and clustering the genes of the array according to their specificity.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Thomas Sierocinski, Anthony Le Béchec, Nathalie Théret, Dimitri Petritis. 2007-10-06. Semantic distillation: a method for clustering objects by their contextual specificity. https://arxiv.org/abs/0710.1203

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Selection in the Zero-Noise Limit for High-Dimensional Diffusions with Measurable Drift

This paper investigates the zero-noise limit of high-dimensional small-noise diffusion processes governed by the stochastic differential equation (SDE) \begin{equation*} \mathrm{d}X_{t}^{\varepsilon }=b(X_{t}^{\varepsilon })\,\mathrm{d}t+\varepsilon \,\mathrm{d}W_{t},\quad X_{0}^{\varepsilon }=0,\quad \varepsilon >0, \end{equation*} where the drift coefficient $b$ is assumed to be measurable and bounded. Under this condition, the associated deterministic ordinary differential equation (ODE) $\dot{x}_{t}=b(x_{t})$ may admit multiple Filippov solutions, i.e., solutions in the sense of differential inclusions, due to the lack of Lipschitz continuity or uniqueness criteria. However, the introduction of non-degenerate additive noise restores well-posedness: the perturbed system admits a unique strong (pathwise) solution for each $\varepsilon >0$. Our analysis is based on the radial-spherical decomposition. We prove that, under certain assumptions, the radius dominates a one-dimensional comparison process and escapes to infinity at a linear rate, and the extrinsic angular martingale converges. Under an explicit separable gradient structure, expressed in terms of a radial potential $V(r,ω)=\langle b(rω),ω\rangle $, the angle converges, for every fixed noise, almost surely to the set of global angular maximizers of $V$, and its distance from that set tends to zero in probability as $\varepsilon\rightarrow0$. The limiting measure is compact, carried by immediate-departure Filippov paths, and, under a Lipschitz radial-graph condition, purely singular with Hausdorff dimension at most $d-1$. A continuous-mapping criterion for exit-time convergence is also established.

math.PR↗

Asymptotic optimality of dynamic first-fit packing on the half-axis

We revisit a classical problem in dynamic storage allocation. Items arrive in a linear storage medium, modeled as a half-axis, at a Poisson rate $r$ and depart after an independent exponentially distributed unit mean service time. The arriving item sizes (lengths) are assumed to be independent and identically distributed (i.i.d.) from a common distribution $H$. A widely employed algorithm for allocating the items is the "first-fit" discipline, namely, each arriving item is placed in the left-most vacant interval large enough to accommodate it. In a seminal 1985 paper, Coffman, Kadota, and Shepp ([6]) proved that in the special case of unit length items (i.e. degenerate $H$), as $r$ tends towards infinity, the first-fit algorithm is asymptotically optimal in the following sense: the steady-state ratio of expected "empty space" (gaps between items) to expected occupied space tends towards $0$. In a sequel to [6], Coffman, Kadota, and Shepp ([5]) conjectured that the first-fit discipline is also asymptotically optimal for non-degenerate $H$. In this paper we provide the first proof of first-fit asymptotic optimality for non-degenerate distributions $H$ of item sizes. Our main result is for the case when $H$ is concentrated on countably many positive real sizes forming an increasing sequence that is either finite or goes to infinity, with the average item size being finite. We prove that under the first-fit discipline, as $r$ tends towards infinity, the steady-state packing configuration (scaled down by $r$) converges in distribution to the limiting packing configuration with smaller items on the left, larger items on the right, and with no gaps between. In particular, this proves asymptotic optimality of first-fit in the following sense: if $P$ is the expected occupied space, then in steady-state the empty space (scaled down by $r$) in $[0,P]$ vanishes.

math.PR↗

Annealed almost periodic entropy

This work studies certain notions of entropy that can be associated to (i) a representation of a separable, unital C*-algebra $\mathfrak{A}$ and (ii) an auxiliary random sequence $(π_n)_{n\ge 1}$ of finite-dimensional representations of $\mathfrak{A}$. This continues a previous research program into the properties of these entropy notions when each $π_n$ is deterministic, which uncovered a range of analogies with entropy in ergodic theory and also with non-commutative generalizations of Szegő's limit theorems. We associate two new notions of entropy to data as in (i) and (ii) above: `annealed' AP entropy, which is roughly a kind of first-moment average of deterministic AP entropies; and `zeroth-order' AP entropy, which controls the large deviations probabilities that certain positive definite functions appear in the representations $π_n$ at all. After developing some of this general theory, we then focus on the special case in which $\mathfrak{A}$ is the group C*-algebra of a finitely-generated free group and each $π_n$ is generated by choosing a tuple of $n$-by-$n$ unitary matrices independently at random from Haar measure. In that case, explicit formulas can be derived for some of our notions of entropy, and new large deviations principles in random matrix theory are obtained as a consequence.

math.PR↗