Search arXiv⌕ Search

arXiv · 2609.39081

Coding Agents for Coding Theory

Abstract

We spent five weeks using an LLM coding agent on open problems in coding theory: finding large sets of four-letter words, such as DNA barcodes, that stay far apart in edit distance. The agent wrote the verifiers and search code; a human chose the problem and set the verification protocol. Restricting the search to codes with a prescribed symmetry, a classical technique, shrank the problem about fourfold and raised the best known code of length 6 and minimum edit distance 3 from 114 to 120 words ($E_4(6,3) \geq 120$). The same pipeline improved twelve further lower bounds at lengths 6 to 9 and distances 3 to 6. We give the failures equal space. Our own search stopped at 116 and recorded the last symmetry class as topping out at 112; a second agent session, running the same search with a better operator, found the 120. A later verdict that the method did not carry over to length 7 was wrong for the same reason, and an earlier instance cost three weeks. Each time, an intermediate result was written down, never rechecked, and treated as a fact that ruled out further search. Checking final outputs, as our protocol required, does not catch such errors.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Abraham Yeung. 2026-09-30. Coding Agents for Coding Theory. https://arxiv.org/abs/2609.39081

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Secret key-distribution over networks with node-based adversarial errors

We study the multiple key-cast problem in network coding under active node-based adversaries. In multiple key-cast, a source generates independent secret keys to be securely and reliably delivered to designated terminal subsets. The network adversary can observe \(\ell_o\) nodes, inject additive or overwrite errors into \(\ell_e\) nodes, and simultaneously observe and corrupt \(\ell_{oe}\) nodes, while having full knowledge of the topology and coding operations. Adversarial models of similar nature, however, where corruption and eavesdropping is done on edges instead of nodes, have seen previous studies in the context of secure multicast network-coding. The work at hand builds on and extends these studies to address the challenges in node-based adversaries in the context of (multiple) key distribution. For single-source networks where every node is d-vertex connected from the source, we show that perfectly secure multiple key-cast under additive and overwrite error models is asymptotically achievable at the key-capacity of \(d-\ell_o-\ell_e-2\ell_{oe}\). We then extend our analysis to networks where only terminal nodes satisfy this connectivity requirement, while intermediate nodes may be only partially connected. For these topologies, we develop coding schemes that achieve secure and reliable multiple key-cast capacities determined by the source vertex-connectivity and additional structural properties of the network. Finally, we show that our results generalize to multi-source settings, ensuring perfect secrecy even if the adversary observes all but one source node, and establish that our constructions apply directly to secure multicast network coding and to network secret-sharing scenarios. As part of our studies, we improve the security guarantee of a central scheme in [Zhang et al., IEEE Trans. Comm., 2023] addressing parallel-edge networks, from weak-security to perfect-security.

cs.IT↗

A Mathematical Theory of Pragmatic Information

We propose a mathematical theory of pragmatic information that connects communication, control, and decision-making. Its central notion is the isoteleia mapping, which formalizes equifinality: distinct semantic paths that lead to the same optimal action are treated as pragmatically equivalent. This mapping yields a three-tier hierarchy of syntactic, semantic, and pragmatic information, in which each successive abstraction removes distinctions that are irrelevant to the task. We then define pragmatic entropy, up/down pragmatic mutual information, channel capacity, and rate-distortion, and prove lossless source coding, channel coding, and rate-distortion theorems that extend Shannon's results. These measures quantify decision uncertainty, reliable transmission, and task-oriented compression at the level of terminal actions. We further introduce pragmatic value of information (VoI) and pragmatic cost of information (CoI) as decision-theoretic duals to rate-distortion and capacity, and develop a Lagrangian dual framework for cross-layer optimization. The resulting pragmatic efficiency bound $\mathcal{E}_p(λ)=\sup_R[Φ_p(R)-λ\mathrm{CoI}_p(R)]$ characterizes the maximum net utility attainable by a resource-constrained intelligent system under a given resource price, yielding a behavioral capacity that extends Shannon's symbol-level capacity to goal-directed action. Extensions to continuous messages provide closed-form expressions for Gaussian channels and sources, while dynamic settings are addressed through a Bellman equation for sequential decision-making. The framework supports task-oriented communication, networked control, autonomous systems, and embodied AI by shifting emphasis from symbol fidelity to the effectiveness of information in guiding actions. In this way, it offers a common language for systems that extract value from information under resource constraints.

cs.IT↗

Exact Second-Order Asymptotics for Covert Communication over DMCs with Variational Distance Constraints

We determine the exact second-order asymptotics of covert communication over binary-input discrete memoryless channels when covertness is measured by variational distance. Previous work by Tahmasbi and Bloch [IEEE Trans. Inf. Theory, Apr. 2019] characterized the first-order asymptotics and derived achievability and converse bounds on the second-order term, but these bounds do not match. The gap arises from an additional penalty of order n^{1/4} in the achievability bound. We show that this penalty can be removed through a sharper analysis of the distribution of the warden's output induced by pulse-position modulation. Specifically, we express the variational distance through the Bhattacharyya coefficient of two distributions and the expectation of a continuous function of the log-likelihood ratio. Because the resulting expectation involves a continuous function rather than the probability of a likelihood-ratio event, an analysis of the characteristic function combined with a Gaussian smoothing argument reduces the approximation error from O(n^{-1/4}) (derived from the Berry--Esseen bound in prior work) to O(n^{-1/2}). With this better controlled approximation error, we manage to derive a matching achievability result to the existing converse result, thus establishing the exact second-order asymptotics.

cs.IT↗