Publication Date

Spring 2026

Degree Type

Master's Project

Degree Name

Master of Science in Computer Science (MSCS)

Department

Computer Science

First Advisor

Amith Kamath Belman

Second Advisor

Mark Stamp

Third Advisor

Thomas Austin

Keywords

Speech Emotion Recognition, Multimodal Fusion, Temporal Con- volutional Network, Prosody, Cross-Modal Attention, HuBERT, DeBERTa, Deep Learning

Abstract

Trying to understand emotion from speech is a problem that is present in human computer interaction. Nevertheless, there are still some shortcomings in current SER methods. Text-based systems may miss vital vocal cues, such as sarcasm, tone changes, and delivery. On the other hand, purely audio-based systems are prone to noise and unstable acoustic features. The combination of linguistic and acoustic features in multimodal approaches partially solves this problem, but many existing approaches use inflexible multimodal fusion techniques that cannot adjust their behaviors according to the quality of input signals. In this work, we propose a multimodal approach based on late fusion, which is capable of solving aforementioned problems using the following four solutions. First, a Prosodic Temporal Convolutional Network (TCN) model that captures temporal changes of pitch, energy, and voice quality in emotional speech. Second, we adopt a multi-stream paralinguistic approach that fuses linguistic features and acoustic features extracted from CNN-BiLSTM spectral analysis, prosody model, and self-supervised HuBERT representation using Squeeze-and-Excitation channel attention. Third, a cross-modal gated attention fusion module that learns how to weigh the linguistic and acoustic input representations, obtaining a noticeable gating weight to indicate the relative importance of different modalities. Finally, we use a modality dropout consistency learning strategy to ensure good performance even if one modality is corrupted. We experimentally validate our proposed approach on two datasets, achieving 85.6% accuracy on RAVDESS+CREMA-D and 41.2% on MELD

Share

COinS