Extractable Memorization From First Principles
Recent work on extractable memorization in language models suffers from two contrasting validity problems. Some studies overstate extraction, for example, by using sequences too short to distinguish memorization from predictability. Others imply that extraction is unreliable evidence because models can reproduce real-world text they weren't explicitly trained on. Both overlook what makes a valid extraction claim: the model must generate a training sequence with high enough probability to indicate memorization. To determine what's high enough, one has to perform a matched comparison, measuring generation probabilities of training and comparable non-training sequences. Since non-training sequences can't have been memorized, their probabilities are a baseline for predictability; exceeding this baseline is evidence of memorization. We formalize matched comparisons with (1) a conformal test calibrated to a chosen false-positive rate when sequences are sampled from populations, and (2) a single-document census that calibrates against a matched non-training document. Matched comparisons enable rigorous, calibrated memorization claims and clarify where prior setups have validity issues. On Wikipedia, OLMo 2 32B reproduces non-training 10-token suffixes roughly 24% as often as training ones: that share reflects false positives, not memorization. For Llama 3.1 70B on books, calibrated census thresholds reach as low as 10^(-27), supporting memorization claims for sequences no feasible sampling budget would extract. We therefore refine "extractable memorization" to require both a valid memorization claim and near-certain generation within a realistic budget.