Publication Date
Spring 2026
Degree Type
Thesis
Degree Name
Master of Science (MS)
Department
Computer Engineering
Advisor
Kaikai Liu; Bernardo Flores; Vidyacharan Bhaskar
Abstract
Autonomous-driving vision-language models describe scenes fluently but reason poorly about metric, safety-critical spatial relationships because they do not read sensor geometry in a grounded way. This thesis presents MoRAL (Multimodal Reasoning for Autonomous Language Models), a two-stage fine-tuning pipeline that teaches a compact 2-billion-parameter VLM to decode a physics-encoded Bird’s Eye View (BEV) representation – LiDAR distance as color, object class as cluster shape, radar Doppler velocity as directional wedges – and then trains it to reason over that representation for driving decisions. Stage 2 then fine-tunes on 57,696 teacher-generated chain-of-thought examples across eight question types, using Cosmos-Reason2-8B as teacher on the nuScenes dataset. Three claims are validated. First, Stage 1 evaluation on 808 held-out frames shows the BEV vocabulary is learnable: the fine-tuned model achieves Zone F1 of 0.89, while both zero-shot baselines fail to produce reliable output (0% parseable for the 2B model, 79.2% empty for the 8B model). Second, on 2,304 held-out validation frames evaluated by a human-calibrated LLM judge, the fine-tuned 2B model achieves a normalized composite score of 0.565, outperforming the zero-shot 8B baseline (0.439) on seven of eight question types despite using four times fewer parameters. Third, grounding improves safety-relevant behavior: EMERGENCY BRAKE recall rises from 10.8% to 47.8%, and output degeneration falls from 94.1% to 20.8%. An edge deployment test confirms the full pipeline runs on a consumer 8 GB GPU at 4.61 GB peak VRAM and 42 tok/s without quantization. Most critical cases still go undetected, so the system is not deployment-ready, but the results show that targeted sensor-grounded fine-tuning substantially improves spatial reasoning in compact VLMs under open-loop evaluation.
Recommended Citation
Govindarajulu Kaliamurthi, Ambarish, "Moral: Multimodal Reasoning for Autonomous Language Models with Sensor-Grounded Spatial Bev Rendering" (2026). Master's Theses. 5752.
DOI: https://doi.org/10.31979/etd.2bbf-xvpf
https://scholarworks.sjsu.edu/etd_theses/5752