arXiv · 2610.12137
Poster: A Preliminary Study of LLM Distillation Inference
Abstract
Unauthorized model distillation, in which a model is trained on the outputs of a proprietary large language model (LLM), is a growing threat to model providers. We study distillation inference: determining whether a suspect model was distilled from another model or trained independently. We formulate this problem as a hypothesis test and estimate the behavior expected under each hypothesis by training shadow models: distilled shadow models learn from the teacher's reasoning traces, whereas independent shadow models learn only from reference answers. The auditor measures how closely each model predicts the teacher's reasoning outputs and then uses the shadow models to convert the suspect's score into a calibrated p-value. In a preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for the suspects, our test achieves a true positive rate of 1.0 at a significance level of 0.02. These results demonstrate the feasibility of using distillation inference to detect distillation attacks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Edward Chen, Yuntao Du. 2026-10-08. Poster: A Preliminary Study of LLM Distillation Inference. https://doi.org/10.1145/3830454.3846416
Cite the original work for its findings. Save a collection to share your selection of sources.