Parallel Stream Transformer Based Architecture for Multimodal User Verification

Publication Date

1-1-2026

Document Type

Conference Proceeding

Publication Title

2026 IEEE World AI Iot Congress Aiiot 2026

DOI

10.1109/AIIoT68874.2026.11569556

First Page

934

Last Page

940

Abstract

This paper presents a Dual-Stream Transformer-based architecture for multimodal user verification, leveraging both keyboard and mouse dynamics to capture complementary behavioral patterns. While conventional sequential models such as Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) networks have shown effectiveness in modeling temporal dependencies, they primarily focus on localized sequential granularity and often struggle to capture long-range contextual relationships. To address this limitation, the proposed framework employs two parallel Transformer-based encoders, each dedicated to one behavioral modality. Each stream integrates temporal convolutional layers for local feature extraction and self-attention mechanisms for modeling global temporal dependencies, allowing the system to learn fine-grained and high-level behavioral representations concurrently. A dot-product fusion mechanism aligns and combines the modality-specific embeddings to enhance cross-modal interaction while maintaining modality independence. Experimental evaluations demonstrate that the proposed framework effectively distinguishes between genuine and impostor inputs across diverse users. These findings highlight the potential of Transformer-based multimodal fusion as a robust and scalable solution for continuous and unobtrusive user authentication.

Funding Number

25-RSG-08-134

Keywords

Authentication, Convolutional Layers, Keystroke Dynamics, Machine Learning, Mouse Dynamics, Multimodal, Transformer-based encoders

Department

Computer Science

Share

COinS