Search arXivSearch

arXiv · 1511.05770

On LR(k)-parsers of polynomial size

Abstract

Usually, a parser for an $LR(k)$-grammar $G$ is a deterministic pushdown transducer which produces backwards the unique rightmost derivation for a given input string $x \in L(G)$. The best known upper bound for the size of such a parser is $O(2^{|G||Σ|^k+k\log |Σ| + \log |G|})$ where $|G|$ and $|Σ|$ are the sizes of the grammar $G$ and the terminal alphabet $Σ$, respectively. If we add to a parser the possibility to manipulate a directed graph of size $O(|G|n)$ where $n$ is the length of the input then we obtain an extended parser. The graph is used for an efficient parallel simulation of all potential leftmost derivations of the current right sentential form such that the unique rightmost derivation of the input can be computed. Given an arbitrary $LR(k)$-grammar $G$, we show how to construct an extended parser of $O(|G| + \#LA |N|2^k k \log k)$ size where $|N|$ is the number of nonterminal symbols and $\#LA$ is the number of relevant lookaheads with respect to the grammar $G$. As the usual parser, this extended parser uses only tables as data structure. Using some ingenious data structures and increasing the parsing time by a small constant factor, the size of the extended parser can be reduced to $O(|G| + \#LA|N|k^2)$. The parsing time is $O(ld(input) + k|G|n)$ where $ld(input)$ is the length of the derivation of the input. Moreover, we have constructed a one pass parser.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Norbert Blum. 2015-11-18. On LR(k)-parsers of polynomial size. https://arxiv.org/abs/1511.05770

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Myhill-Nerode Theorem for Generalized Automata, with Applications to Pattern Matching and Compression

The model of generalized automata, introduced by Eilenberg in 1974, allows representing a regular language more concisely than conventional automata by allowing edges to be labeled not only with characters, but also strings. Giammarresi and Montalbano introduced a notion of determinism for generalized automata [STACS 1995]. While generalized deterministic automata retain many properties of conventional deterministic automata, the uniqueness of a minimal generalized deterministic automaton is lost. In the first part of the paper, we show that the lack of uniqueness can be explained by introducing a set $ \mathcal{W(A)} $ associated with a generalized automaton $ \mathcal{A} $. In this way, we derive for the first time a full Myhill-Nerode theorem for generalized automata, which contains the textbook Myhill-Nerode theorem for conventional automata as a degenerate case. In the second part of the paper, we show that the set $ \mathcal{W(A)} $ leads to applications for pattern matching and data compression. We show that a Wheeler generalized automata can be stored using $ \mathfrak{e} \log σ(1 + o(1)) + O(e) $ bits so that pattern matching queries can be solved in $ O(m \log \log σ) $ time, where $ \mathfrak{e} $ is the total length of all edge labels, $ e $ is the number of edges, $ σ$ is the size of the alphabet and $ m $ is the length of the pattern.

cs.FL

Quadratic Word Equations with a Linear Side: Polynomial Nielsen Graph Diameter and NP-Completeness

The satisfiability problem for word equations asks whether variables can be replaced by words so that the two sides become equal. For regular word equations, in which each variable occurs at most once on each side, satisfiability is NP-complete. For general quadratic word equations, in which each variable occurs at most twice in total, satisfiability is NP-hard, but its membership in NP remains open. We consider an intermediate class: quadratic word equations with a linear side, where each variable occurs at most once on one designated side. We show that the Nielsen graph of an equation $U=V$ in this class, with total length $N=|U|+|V|$, has diameter $O(N^{12})$, measured over reachable pairs of vertices. Together with the known NP-hardness for regular word equations, this result establishes NP-completeness of satisfiability for this class. We also show that each strongly connected component is isomorphic to the length-preserving reachability graph of a regular equation, and that the condensation graph has depth at most $|U|+2|V|$.

cs.FL

Contributions to the hierarchy of probabilistic languages

We reconsider the theory of probabilistic formal languages generated by n-gram models and by probabilistic context-free grammars (PCFGs). The expected hierarchy of probabilistic grammars is established by proving that every probabilistic language generated by an n-gram model is also generated by some PCFG, while some probabilistic languages generated by PCFGs cannot be generated by any $n$-gram model. We introduce the notion of fully connected PCFGs, namely PCFGs in Chomsky normal form where every production rule only involving non-terminals has non-zero probability. Our main result shows that any probabilistic language generated by an $n$-gram model differs from any probabilistic language generated by a fully connected PCFG. Therefore, the class of probabilistic languages generated by $n$-gram models is not a subset of the class generated by fully connected PCFGs.

cs.FL