Search arXivSearch

arXiv · 2606.10059

Compiling Rewrite Rules to Finite-State Transducers with the Worsening Trick

Abstract

Finite-state transducers (FSTs) are essential for modeling string rewriting in computational linguistics and natural language processing (NLP), particularly for phonological and morphological rewrite rules. Compiling general rewrite rules of the form $A \to B / L \, \_ \, R$, where $A$, $B$, $L$, and $R$ are arbitrary regular languages, is complex due to overlapping matches and context constraints. Traditional methods, such as those by Kaplan and Kay or Karttunen, rely on intricate transducer compositions with auxiliary markers. This paper presents a compact compilation scheme based on the "worsening trick'': generate all legal rewrite candidates, then filter candidates that are worse than another candidate for the same input. Implemented as the built-in rewrite compiler in PyFoma, the construction supports multiple contexts, arbitrary transductions, markup, directed rewriting, weights, and parallel rewriting. The resulting formulas are short and uniform, and where semantics coincide, they reproduce the same rule transducers as earlier approaches while remaining easier to extend. The implementation has been validated against foma on both a substantial collection of rewrite grammars and an automated regression suite covering the major rewrite modalities, with the resulting transducers matching exactly apart from state numbering.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mans Hulden, Michael Ginn. 2026-06-08. Compiling Rewrite Rules to Finite-State Transducers with the Worsening Trick. https://arxiv.org/abs/2606.10059

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Regular Expressions with Backreferences on Multiple Context-Free Languages, and the Closed-Star Condition

Backreference is a well-known practical extension of regular expressions and is supported by the regular expression engines in the standard libraries of most modern programming languages, such as Java, Python, JavaScript and more. A difficulty of backreference is non-regularity: backreference strictly enhances the expressive power of regular expressions to the point that regular expressions with backreferences (rewbs) can describe non-regular (in fact, even non-context-free) languages. In this paper, we investigate the expressive power of rewbs by comparing rewbs to multiple context-free languages (MCFL) and parallel multiple context-free languages (PMCFL). First, we prove that the language class of rewbs is a proper subclass of unary-PMCFLs, which coincide with the EDT0L languages. Our result strictly improves the known (non-trivial) upper bound of rewbs, because the best-known bound was the intersection of the class of nondeterministic logspace languages and that of indexed languages, and, as we shall show in this paper, the class of EDT0L languages is a proper subclass of the intersection. Additionally, we show that, however, the language class of rewbs is not contained in that of MCFLs even when restricted to rewbs with only one capturing group and no captured references. Therefore, in general, the parallelism seems essential for rewbs. Backed by these results, we define a novel syntactic condition on rewbs that we call closed-star and observe that it provides an upper bound on the number of times a rewb references the same captured string. The closed-star condition allows dispensing with the parallelism: we prove that the language class of closed-star rewbs falls inside the class of unary-MCFLs, which is equivalent to that of EDT0L systems of finite index. Furthermore, we show that the language class of closed-star rewbs also falls inside the class of nonerasing stack languages.

cs.FL

Compressed Subsequence Checking is PSPACE-complete

It is shown that the (scattered) subsequence problem for two words represented by straight-line programs is PSPACE-complete, even over a binary alphabet. The lower bound is obtained by a polynomial-time reduction from quantified subset sum.

cs.FL

On the Kanazawa--Salvati Conjecture

The language $\mathrm{MIX}$ consists of all words over a three-letter alphabet that have an equal number of occurrences of each letter. It is also the word problem of $\mathbb{Z}^2$ with respect to a suitable choice of generators. The Kanazawa--Salvati conjecture states that $\mathrm{MIX}$ is not a well-nested multiple context-free language. Every well-nested multiple context-free language is an indexed language. We reduce the conjecture to an explicit combinatorial problem about tuples of words, which is easier to state than the original formulation in terms of arbitrary well-nested multiple context-free grammars. More generally, for every surjective monoid homomorphism $ψ\colon Σ^* \to \mathbb{Z}^d$, we define a family of well-nested multiple context-free grammars $G_ψ[r]$ for $r \geq 1$, each of which generates a sublanguage of $ψ^{-1}(\mathbf{0})$. We prove that every well-nested multiple context-free sublanguage of $ψ^{-1}(\mathbf{0})$ is contained in $L(G_ψ[r])$ for some $r \geq 1$. Using this family, we prove that the four-letter analogue $\mathrm{MIX}_4$, which is a word problem of $\mathbb{Z}^3$, is not a well-nested multiple context-free language. The proof reduces this claim to a result of Bishop--Elder--Evetts--Gallot--Levine stating that the two-letter analogue $\mathrm{MIX}_2$ is not generated by any non-branching multiple context-free grammar. The Kanazawa--Salvati conjecture itself remains open.

cs.FL