Search arXiv⌕ Search

arXiv · 0710.1525

Efficient Optimally Lazy Algorithms for Minimal-Interval Semantics

Abstract

Minimal-interval semantics associates with each query over a document a set of intervals, called witnesses, that are incomparable with respect to inclusion (i.e., they form an antichain): witnesses define the minimal regions of the document satisfying the query. Minimal-interval semantics makes it easy to define and compute several sophisticated proximity operators, provides snippets for user presentation, and can be used to rank documents. In this paper we provide algorithms for computing conjunction and disjunction that are linear in the number of intervals and logarithmic in the number of operands; for additional operators, such as ordered conjunction and Brouwerian difference, we provide linear algorithms. In all cases, space is linear in the number of operands. More importantly, we define a formal notion of optimal laziness, and either prove it, or prove its impossibility, for each algorithm. We cast our results in a general framework of antichains of intervals on total orders, making our algorithms directly applicable to other domains.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sebastiano Vigna, Paolo Boldi. 2016-08-11. Efficient Optimally Lazy Algorithms for Minimal-Interval Semantics. https://arxiv.org/abs/0710.1525

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Tractable Gap-Constraint Languages for Complex Event Recognition

For strings $u, D \in Σ^*$, a subsequence embedding of $u$ in $D$ is a function $e \colon \{1, 2, \ldots, |u|\} \to \{1,2,\ldots,|D|\}$ with $e(i) < e(i+1)$ for every $i \in \{1,2,\ldots,|u|-1\}$ and the $i$-th symbol of $u$ equals the $e(i)$-th symbol of $D$. A gap-constraint for $u$ is a triple $(i,j,L)$ with $1 \leq i < j \leq |u|$ and $L$ is a regular language over $Σ$. An embedding $e$ satisfies a gap-constraint $(i, j, L)$ if the factor of $D$ strictly between positions $e(i)$ and $e(j)$ is a word from $L$. We investigate the subsequence matching problem with gap-constraints, which is relevant in the context of complex event recognition (CER): given $u, D \in Σ^*$ and a set $C$ of gap-constraints, find an embedding of $u$ in $D$ that satisfies all gap-constraints from $C$. In general, subsequence matching with gap constraints is NP-complete and the only known tractable variants restrict the interval structure of the gap-constraints. In this work, we show that we can solve subsequence matching with gap-constraints with an arbitrary interval structure rather efficiently (in fact, optimally under SETH) in time $O(|D| (|u|+|C|))$ if the gap-constraint languages satisfy a property which we dub left-convexity: whenever $u v w \in L$ and $v \in L$, then also $uv \in L$. Left-convex languages are sufficiently expressive to model interesting real-world scenarios considered in CER, e.g., length constraints $L = \{w \mid a \leq |w| \leq b\}$ for $a, b \in \mathbb{N}$. We also show how our algorithm can be used in order to efficiently enumerate all satisfying embeddings, which is particularly relevant for possible applications in CER. Finally, we show how non-left-convex languages can lead to intractability, i.e., if in addition to length constraints we allow $\{aa,ε\}$ as the only non-left-convex constraint language, then the problem is NP-complete again.

cs.DS↗

A Gossiping Protocol for Sparse Ad-Hoc Radio Networks

We study the problem of gossiping (all-to-all information exchange) in ad-hoc radio networks. Such a network is represented by a strongly-connected directed graph with \(n\) vertices, whose topology is initially unknown to the protocol. In 2004, Gasieniec, Radzik, and Xin gave a \(\tilde O(n^{4/3})\)-time deterministic protocol for this problem, and closing the gap between their upper bound and the \(\tildeΩ(n)\) lower bound on the time complexity of gossiping remains a central open problem. We develop a deterministic protocol for gossiping in ad-hoc radio networks that achieves running time \(\tilde O((mn)^{3/5})\) for directed graphs with at most \(m\) edges. Our protocol improves on the \(\tilde O(n^{4/3})\) bound when \(m = O(n^c)\), for \(c < 11/9\). We also present a \(\tilde O(Δ^{1/2} n)\)-time gossiping protocol for \(Δ\)-regular graphs.

cs.DS↗

Generating cyclic pivot Gray codes for well-ordered $k$-degenerate graphs in constant amortized time

A graph $G$ is $k$-degenerate if there exists an ordering $v_1, v_2, \dots, v_n$ of its vertices such that each vertex $v_i$ has at most $k$ neighbors $v_j$ in $G$ with $j < i$. A well-ordered $k$-degenerate graph is a labeled graph on the vertex set $\{1, 2, \dots, n\}$ in which every vertex $i$ has at most $k$ neighbors among $1, 2, \dots, i-1$. We present the first simple algorithms that generate, rank, and unrank cyclic pivot Gray codes for well-ordered $k$-degenerate graphs, where consecutive graphs differ by the addition, removal, or pivoting of a single edge. Our algorithm generates each well-ordered $k$-degenerate graph in constant amortized time per graph, using $O(n^2)$ space, while ranking and unranking take $O(n^2)$ time and $O(n^2)$ space.

cs.DS↗