Publication Date
Spring 2026
Degree Type
Thesis
Degree Name
Master of Science (MS)
Department
Applied Data Science
Advisor
Vishnu Pendyala; Guannan Liu; Mohammad Masum
Abstract
The rapid surge of deepfake audio presents a significant challenge, driving the need for effective detection techniques. Large Language Models (LLMs) have demonstrated significant cross-domain adaptability and can be fine-tuned for a wide range of applications, including audio deepfake detection. This research explores the use of multimodal LLMs for detecting deepfake audio, with a particular emphasis on the usage of explainable artificial intelligence methods, such as Local Interpretable Model-agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP). These techniques facilitate the identification of key acoustic features that could influence the model’s classifications, thereby improving the explainability and reliability of the detection framework. Experimental results indicate that the fine-tuned models exhibit strong correlations with key acoustic features, particularly loudness attributes and spectral variation, in their predictive behavior.
Recommended Citation
Manoharan, Janani Kripa, "Explainable Deepfake Audio Detection Using Fine-Tuned Foundation Models" (2026). Master's Theses. 5764.
DOI: https://doi.org/10.31979/etd.54rk-euuc
https://scholarworks.sjsu.edu/etd_theses/5764