Author

Publication Date

Spring 2026

Degree Type

Master's Project

Degree Name

Master of Science in Computer Science (MSCS)

Department

Computer Science

First Advisor

Navrati Saxena

Second Advisor

William Andreopoulos

Third Advisor

Sai Sashank Peddibhotla

Keywords

millimetre-wave bandwidth allocation; deep reinforcement learning;Vision Transformer; multimodal state representation; proactive resource managememy

Abstract

While millimetre-wave (mmWave) wireless networks offer substantial bandwidth capacity, their susceptibility to blockage induces abrupt and unpredictable channel degradation. In this paper we investigate the use of co-located camera sensing at a mmWave base station to enhance proactive bandwidth allocation via reinforcement learning. We propose a multimodal framework in which a Dueling Double Deep Q- Network (D3QN) with Prioritized Experience Replay (PER) operates on a joint state representation comprising a 768-dimensional scene embedding extracted from a frozen pre-trained Vision Transformer (ViT) and a 128-dimensional PCA-compressed channel state information (CSI) feature vector. The agent performs a three-class bandwidth allocation task on the ViWi co-located camera dataset. To analyze the contribution of architectural and feature design choices, we conduct a 2×2 ablation study across algorithm variants (DDQN, D3QN) and input modalities (CSI-only, ViT+CSI). Our proposed framework achieves 66.4% accuracy on held-out user positions, outperforming a naive baseline by 17.5%, a DDQN by 15.2% points, and its CSI-only counterpart by 16.5%. Notably, these gains are much higher in weak-channel detection. Our proposed D3QN attains 85.5% accuracy for conservative bandwidth allocation compared to 24.5% for the CSI-only variant. Ablation results indicate that visual features degrade DDQN performance but substantially enhance D3QN, yielding the largest performance gain among all configurations. Further analysis attributes the effectiveness of the proposed approach to PER, which biases learning toward infrequent but critical weak-channel conditions where visual context is most informative.

Available for download on Sunday, May 23, 2027

Share

COinS