Steal the Knowledge, Inherit the Mark: TwinMark for Distillation Watermarking via Second Moments and Class Multiplexing
Knowledge distillation can copy a deployed model by training a student on its logits or features. The student inherits a watermark only through the output it imitates. Both outputs therefore need a mark, yet most distillation watermarks cover only one. TwinMark instead writes one secret payload into what each distillation objective preserves, the second moment of normalized features and the class-mean logits. Both marks use linear readouts that give sufficient conditions for bit recovery on fixed audit inputs, while detection is tested separately under a declared null. In a ten-seed CIFAR-100 study, each mark is detected in every student that imitates its output and stays at chance otherwise. Marking costs the teacher 1.3 accuracy points and the audited embedding of feature-distilled students 4.4 nearest-neighbor points. Second-moment bounds certify 50 to 60 of 64 bits in every feature-distilled student, whereas pointwise bounds and the evaluated logit bounds certify none. Class multiplexing, which signs the payload per class, raises the logit decoder's rank ceiling and lowers bit errors at equal energy, and even at 10x over-encoding, 1,024 bits on 100 classes stay detectable in every student. Detection extends to 39 of 40 students of four other architectures and to segmentation, object detection and satellite-navigation jammer detection and classification. These results establish output-matched inheritance under the evaluated protocols, not resistance to all utility-preserving transformations.