Search arXivSearch

arXiv · 2209.10031

The exact probability law for the approximated similarity from the Minhashing method

Abstract

We propose a probabilistic setting in which we study the probability law of the Rajaraman and Ullman \textit{RU} algorithm and a modified version of it denoted by \textit{RUM}. These algorithms aim at estimating the similarity index between huge texts in the context of the web. We give a foundation of this method by showing, in the ideal case of carefully chosen probability laws, the exact similarity is the mathematical expectation of the random similarity provided by the algorithm. Some extensions are given. \noindent \textbf{Résumé.} Nous proposons un cadre probabilistique dans lequel nous étudions la loi de probabilité de l'algorithme de Rajaraman et Ullman \textit{RU} ainsi qu'une version modifiée de cet algorithme notée \textit{RUM}. Ces alogrithmes visent à estimer l'indice de la similarité entre des textes de grandes tailles dans le contexte du Web. Nous donnons une base de validité de cette méthode en montrant que pour des lois de probabilités minutieusement choisies, la similarité exacte est l'espérance mathématique de la similarité aléatoire donnée par l'algorithme \textit{RUM}. Des généralisations sont abordées.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Soumaila Dembele, Gane Samb Lo. 2022-09-25. The exact probability law for the approximated similarity from the Minhashing method. https://doi.org/10.16929/as%2F2017.1199.100

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The extremal process of a cascading family of branching Brownian motion

We study the asymptotic behaviour of the extremal process of a cascading family of branching Brownian motions. This is a particle system on the real line such that each particle has a type in addition to his position. Particles of type $1$ move on the real line according to Brownian motions and branch at rate $1$ into two children of type $1$. Furthermore, at rate $α$, they give birth to children too of type $2$. Particles of type $2$ move according to standard Brownian motion and branch at rate $1$, but cannot give birth to descendants of type $1$. We obtain the asymptotic behaviour of the extremal process of particles of type $2$.

math.PR

Breuer-Major Theorems for Hilbert Space-Valued Random Variables

Let $\{X_k\}_{k\in\mathbb Z}$ be a stationary Gaussian process with values in a separable Hilbert space $\mathcal H_1$, and let $G:\mathcal H_1\to\mathcal H_2$ be a measurable map into another separable Hilbert space $\mathcal H_2$. We derive a central limit theorem for the centered normalized partial sums of the Hilbert space-valued subordinated process $\{G[X_k]\}_{k\in\mathbb Z}$. Our result holds under either of two sets of sufficient conditions, formulated in terms of the transformation $G$ and the temporal and cross-sectional dependence structure of $\{X_k\}_{k\in\mathbb Z}$. These conditions coincide in finite dimensions but lead to genuinely different phenomena in the infinite-dimensional setting. The proof relies on the recently developed Fourth Moment Theorem on Hilbert spaces, leveraging tools from the infinite-dimensional Malliavin-Stein framework. We also provide continuous-time and quantitative versions of the central limit theorem. In a series of examples, we recover and strengthen limit theorems for a wide array of statistics relevant in functional data analysis, and present, as an application of our result, a novel limit theorem in the framework of neural operators.

math.PR

Controlled rough SDEs, pathwise stochastic control and dynamic programming principles

We study stochastic optimal control of rough stochastic differential equations (RSDEs). This is in the spirit of the pathwise control problem (Lions--Souganidis 1998, Buckdahn--Ma 2007; also Davis--Burstein 1992), with renewed interest and recent works drawing motivation from filtering, SPDEs, and reinforcement learning. Results include regularity of rough value functions, validity of a rough dynamic programming principles and new rough stability results for HJB equations, removing excessive regularity demands previously imposed by flow transformation methods. Measurable selection is used to relate RSDEs to "doubly stochastic" SDEs under conditioning. In contrast to previous works, Brownian statistics for the to-be-conditioned-on noise are not required, aligned with the "pathwise" intuition that these should not matter upon conditioning. Depending on the chosen class of admissible controls, the involved processes may also be anticipating. The resulting stochastic value functions coincide in great generality for different classes of controls. RSDE theory offers a powerful and unified perspective on this problem class.

math.PR