Publication Date
Spring 2026
Degree Type
Master's Project
Degree Name
Master of Science in Computer Science (MSCS)
Department
Computer Science
First Advisor
Amith Kamath Belman
Second Advisor
Mark Stamp
Third Advisor
Thomas Austin
Keywords
Speech Emotion Recognition, Multimodal Fusion, Temporal Con- volutional Network, Prosody, Cross-Modal Attention, HuBERT, DeBERTa, Deep Learning
Abstract
Trying to understand emotion from speech is a problem that is present in human computer interaction. Nevertheless, there are still some shortcomings in current SER methods. Text-based systems may miss vital vocal cues, such as sarcasm, tone changes, and delivery. On the other hand, purely audio-based systems are prone to noise and unstable acoustic features. The combination of linguistic and acoustic features in multimodal approaches partially solves this problem, but many existing approaches use inflexible multimodal fusion techniques that cannot adjust their behaviors according to the quality of input signals. In this work, we propose a multimodal approach based on late fusion, which is capable of solving aforementioned problems using the following four solutions. First, a Prosodic Temporal Convolutional Network (TCN) model that captures temporal changes of pitch, energy, and voice quality in emotional speech. Second, we adopt a multi-stream paralinguistic approach that fuses linguistic features and acoustic features extracted from CNN-BiLSTM spectral analysis, prosody model, and self-supervised HuBERT representation using Squeeze-and-Excitation channel attention. Third, a cross-modal gated attention fusion module that learns how to weigh the linguistic and acoustic input representations, obtaining a noticeable gating weight to indicate the relative importance of different modalities. Finally, we use a modality dropout consistency learning strategy to ensure good performance even if one modality is corrupted. We experimentally validate our proposed approach on two datasets, achieving 85.6% accuracy on RAVDESS+CREMA-D and 41.2% on MELD
Recommended Citation
Munshi, Shubhankar Sameer, "Multimodal Emotion Detection System" (2026). Master's Projects. 1801.
DOI: https://doi.org/10.31979/etd.kdz9-cqvv
https://scholarworks.sjsu.edu/etd_projects/1801