Search arXivSearch

arXiv · 2504.14370

Density Measures for Language Generation

Abstract

The recent successes of large language models (LLMs) have led to a surge of theoretical research into language generation. A recent line of work proposes an abstract view, called language generation in the limit, where generation is seen as a game between an adversary and an algorithm: the adversary generates strings from an unknown language $K$, chosen from a countable collection of candidate languages, and after seeing a finite set of these strings, the algorithm must generate new strings from $K$ that it has not seen before. This formalism highlights a key tension: the trade-off between validity (the algorithm should only produce strings from the language) and breadth (it should be able to produce many strings from the language). This trade-off is central in applied language generation as well, where it appears as a balance between hallucination (generating invalid utterances) and mode collapse (generating only a restricted set of outputs). Despite its importance, this trade-off has been challenging to study quantitatively. We develop ways to quantify this trade-off by formalizing breadth using measures of density. Existing algorithms for language generation in the limit produce output sets that can have zero density in the true language, and this important failure of breadth might seem unavoidable. We show, however, that such a failure is not necessary: we provide an algorithm for language generation in the limit whose outputs have strictly positive density in $K$. We also study the internal representations built by these algorithms, specifically the sequence of hypothesized candidate languages they consider, and show that achieving the strongest form of breadth may require oscillating indefinitely between high- and low-density representations. Our analysis introduces a novel topology on language families, with notions of convergence and limit points playing a key role.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jon Kleinberg, Fan Wei. 2025-04-19. Density Measures for Language Generation. https://arxiv.org/abs/2504.14370

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Rooted Spider Embeddings and the Erd\H os-Sós Conjecture

Under a local density condition, we prove that every $k$-edge spider embeds at any prescribed center of degree at least $k$, unless all legs are even and the host graph has one of two specified structures. These structures contain complete bipartite subgraphs with prescribed neighborhoods. The proof uses path rerouting and three exchange lemmas that describe equality in neighborhood estimates. As a consequence, we recover the Erd\H os-Sós bound for all spiders.

math.CO

Generalized Goulden-Yong duals and signed minimal factorizations

In this paper, we give two combinatorial ways to study signed exceptional sequences. First, we show the equivalence between one-way reflections and relatively projective representations. Secondly, we construct generalized Goulden-Yong duals using reverse Garside element actions and folded chord diagrams. We then give two applications of the generalized Goulden-Yong duals: constructing generalized Prüfer codes and counting signed factorizations using the matrix-tree theorem.

math.CO

Explicit expressions for iterates of power series

We present several formulas for both the discrete and fractional iterates of an invertible power series $f$, using a new unifying approach based on umbral calculus. Known formulas are extended, and their proofs simplified, while new expressions are introduced. In particular, by employing $q$-calculus identities, we eliminate the requirement for $f'(0)$ to equal $1$ and the resulting general expressions for the iterative logarithm are obtained as well.

math.CO