Search arXiv⌕ Search

arXiv subjects

Yi Chen Liu

Publications and source records attributed to Yi Chen Liu.

2 recordsLinked to original sources

Tracing Decoder Artifacts for Compact Synthetic Speech Screening

Recent advances in speech synthesis and voice cloning have increased the need for reliable synthetic-speech detection, yet high-accuracy detectors increasingly rely on large pretrained models that are costly to invoke on every recording. Rather than replacing such detectors, we investigate a compact front-end screen that processes all inputs cheaply and forwards only suspicious recordings for more expensive analysis. To enable lightweight screening without a large learned encoder, we exploit spectral traces introduced by speech-generation operations. We analyze how learned upsampling and inverse short-time Fourier transform synthesis can produce predictable spectral artifacts and measure their presence directly in generated waveforms. Because the strength of these artifacts varies across generators, we combine decoder-guided spectral measurements with complementary descriptors of short-time spectral shape and temporal variation in a compact gradient-boosted tree. Across seven speech generators and two human-speech sources, the proposed screen achieves an equal error rate of 0.021\% with an estimated model storage of 151 KiB. When used as the first stage of a simulated cascade with a 1.15-billion-parameter detector, it reduces estimated detection energy by 84.4\% while operating at a 0.050\% synthetic-speech miss rate, demonstrating the potential of decoder-guided acoustic evidence for low-cost front-end screening.

cs.SD↗

Auditing Training Data in Generative Music Models via Black-Box Membership Inference

Recent advances in text-to-music generation enable high-fidelity synthesis of structured musical audio, raising growing concerns about data provenance, consent, and training transparency. These models are typically trained on large-scale corpora with little disclosure, leaving no practical mechanism to verify whether a particular audio sample was included in training. In this paper, we investigate black-box membership inference for generative music models, aiming to determine whether a candidate music sample was used during training, given only query access to the deployed system. Our key insight is that training membership induces systematically stronger semantic and structural alignment between a candidate sample and the model's generation conditioned on its caption. We query the target model with the associated caption and measure the relationship between the candidate audio and the generated output in a learned feature space. To capture features that separate members from non-members, we construct paired examples consisting of each track and its caption-conditioned generation from shadow models, and train a music auditor to classify membership. The auditor captures alignment patterns characteristic of training membership and generalizes to unseen target models in a fully black-box setting without access to model parameters or training metadata. Across multiple state-of-the-art music generators, our method achieves up to 98.6% accuracy, with false-positive and false-negative rates as low as 1.9% and 1.0%, demonstrating that reliable training-data auditing is feasible in realistic deployment scenarios.

cs.LG↗