Publication Date
Spring 2026
Degree Type
Master's Project
Degree Name
Master of Science in Computer Science (MSCS)
First Advisor
Faranak Abri
Second Advisor
Amith Kamath Belman
Third Advisor
Saptarshi Sengupta
Keywords
audio deepfake detection, self-supervised speech models, variational information bottleneck, per-layer-group bottleneck, adversarial domain adaptation, cross-domain generalization
Abstract
Audio deepfake detectors trained on one corpus often generalize poorly to unfamiliar recording channels, codecs, and synthesizers. This project examines where in a self-supervised speech model the channel and codec nuisance is concentrated, and introduces GLADA, an architecture that aggregates SSL hidden states by layer group, applies a Variational Information Bottleneck and an adversarial domain discriminator to each group, and encourages group diversity through a CKA penalty. A multi-seed cross-dataset study on In-the-Wild, ASVspoof 2021 LA, ASVspoof 2021 DF, WaveFake, and CodecFake+ finds that restricting the bottleneck to the lowest layer group outperforms both a single post-pooling bottleneck and the full per-group placement on every secondary out-of-domain dataset. Per-group linear probes and per-group CKA against handcrafted acoustic descriptors both localize codec and channel signal to that lowest group, which supports the placement.
Recommended Citation
Kauffmann, Jakob, "Targeted Bottlenecking of Self-Supervised Speech Layers for Generalizable Audio Deepfake Detection" (2026). Master's Projects. 1804.
DOI: https://doi.org/10.31979/etd.kmc6-t6vf
https://scholarworks.sjsu.edu/etd_projects/1804