arXiv · 2310.18491
Publicly-Detectable Watermarking for Language Models
Abstract
We present a publicly-detectable watermarking scheme for LMs: the detection algorithm contains no secret information, and it is executable by anyone. We embed a publicly-verifiable cryptographic signature into LM output using rejection sampling and prove that this produces unforgeable and distortion-free (i.e., undetectable without access to the public key) text output. We make use of error-correction to overcome periods of low entropy, a barrier for all prior watermarking schemes. We implement our scheme and find that our formal claims are met in practice.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jaiden Fairoze, Sanjam Garg, Somesh Jha, Saeed Mahloujifar, Mohammad Mahmoody, Mingyuan Wang. 2023-10-27. Publicly-Detectable Watermarking for Language Models. https://arxiv.org/abs/2310.18491
Cite the original work for its findings. Save a collection to share your selection of sources.