Search arXivSearch

arXiv · 2209.08383

Practical LR Parser Generation

Abstract

Parsing is a fundamental building block in modern compilers, and for industrial programming languages, it is a surprisingly involved task. There are known approaches to generate parsers automatically, but the prevailing consensus is that automatic parser generation is not practical for real programming languages: LR/LALR parsers are considered to be far too restrictive in the grammars they support, and LR parsers are often considered too inefficient in practice. As a result, virtually all modern languages use recursive-descent parsers written by hand, a lengthy and error-prone process that dramatically increases the barrier to new programming language development. In this work we demonstrate that, contrary to the prevailing consensus, we can have the best of both worlds: for a very general, practical class of grammars -- a strict superset of Knuth's canonical LR -- we can generate parsers automatically, and the resulting parser code, as well as the generation procedure itself, is highly efficient. This advance relies on several new ideas, including novel automata optimization procedures; a new grammar transformation ("CPS"); per-symbol attributes; recursive-descent actions; and an extension of canonical LR parsing, which we refer to as XLR, which endows shift/reduce parsers with the power of bounded nondeterministic choice. With these ingredients, we can automatically generate efficient parsers for virtually all programming languages that are intuitively easy to parse -- a claim we support experimentally, by implementing the new algorithms in a new software tool called langcc, and running them on syntax specifications for Golang 1.17.8 and Python 3.9.12. The tool handles both languages automatically, and the generated code, when run on standard codebases, is 1.2x faster than the corresponding hand-written parser for Golang, and 4.3x faster than the CPython parser, respectively.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Joe Zimmerman. 2022-09-17. Practical LR Parser Generation. https://arxiv.org/abs/2209.08383

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Adding Reconfiguration to Zielonka's Asynchronous Automata

We study an extension of Zielonka's (fixed) asynchronous automata called reconfigurable asynchronous automata where processes can dynamically change who they communicate with. We show that reconfigurable asynchronous automata are not more expressive than fixed asynchronous automata by giving translations from one to the other. However, going from reconfigurable to fixed comes at the cost of disseminating communication (and knowledge) to all processes in the system. We then show that this is unavoidable by describing a language accepted by a reconfigurable automaton such that in every equivalent fixed automaton, every process must either be aware of all communication or be irrelevant.

cs.FL

Certificates for short extending words in a finite automaton

Let $\mathcal A$ be a complete deterministic finite automaton on a state set $Q$ of size $n$ with $k$ letters, and for a proper nonempty subset $S$ of $Q$ let $\mathrm{minext}(S)$ be the length of a shortest word $u$ with $|Su^{-1}|>|S|$, where $Su^{-1}=\{q: q\cdot u\in S\}$. To each state $q$ attach the integer $β^{\ast}_q=\sum_{t=1}^{n-1}k^{\,n-1-t}(\mathrm{indeg}_t(q)-k^{t})$, where $\mathrm{indeg}_t(q)$ counts the pairs $(p,u)$ with $|u|=t$ and $p\cdot u=q$, and let $B(S)=\sum_{q\in S}β^{\ast}_q$. On every synchronizing automaton, $B(S)\ge0$ implies $\mathrm{minext}(S)\le n-1$, so, as $B(Q)=0$, one of $S$ and $Q\setminus S$ extends within $n-1$; when $B(S)>0$ no hypothesis is needed. Kari's Eulerian extension lemma is the case $β^{\ast}=0$, and $β^{\ast}$, like every member of the family $\sum_{t=1}^{n-1}c_tσ_t$, $c_t>0$, vanishes identically if and only if the automaton is Eulerian, where $σ_t(S)=\sum_{q\in S}(\mathrm{indeg}_t(q)-k^{t})$. On strongly connected automata $σ_t(S)/k^{t}$ has Cesàro limit $n\,e(S)/e(Q)-|S|$ for Friedman's weight $e$; that limit certifies singletons but no larger subset in general. The hypothesis $B(S)\ge0$ cannot be relaxed by one integer unit, nor can the constant $n-1$ be improved. A second-moment test on the sizes $|Su^{-1}|$ certifies 60 to 95 percent of the subsets with $B(S)<0$ at $n\le7$. Along non-Eulerian automata whose words of length $n-1$ merge a fraction of the state pairs bounded below, with $\max_q\mathrm{indeg}_{n-1}(q)=o(nk^{n-1})$, it certifies all but a vanishing share of them. The functional $B$ certifies half of the subsets outside $\{B=0\}$. At each subset size coprime to $n$ ($n\ge4$) some synchronizing Eulerian binary automaton attains the constant $n-1$; whether only there is open. No reset bound follows: Černý's automata have subsets not extending within $n-1$.

cs.FL

Quadratic Word Equations with a Linear Side: Polynomial Nielsen Graph Diameter and NP-Completeness

The satisfiability problem for word equations asks whether variables can be replaced by words so that the two sides become equal. For regular word equations, in which each variable occurs at most once on each side, satisfiability is NP-complete. For general quadratic word equations, in which each variable occurs at most twice in total, satisfiability is NP-hard, but its membership in NP remains open. We consider an intermediate class: quadratic word equations with a linear side, where each variable occurs at most once on one designated side. We show that the Nielsen graph of an equation $U=V$ in this class, with total length $N=|U|+|V|$, has diameter $O(N^{12})$, measured over reachable pairs of vertices. Together with the known NP-hardness for regular word equations, this result establishes NP-completeness of satisfiability for this class.

cs.FL