Off-campus SJSU users: To download campus access theses, please use the following link to log into our proxy server with your SJSU library user name and PIN.

Publication Date

Summer 2026

Degree Type

Thesis - Campus Access Only

Degree Name

Master of Science (MS)

Department

Computer Engineering

Advisor

Kaikai Liu; Bernardo Flores; Fabio Di Troia

Abstract

End-to-end driving policies trained by imitation learning are optimized to reproduce recorded trajectories rather than to respect the rules of the road, and their failures compound in closed loop. This work asks whether reward terms derived from a high-definition (HD) map can improve the closed-loop behavior of a large vision-language-action driving model. We convert Argoverse 2 sensor logs into photorealistic 3D Gaussian Splatting reconstructions, embed each log’s HD vector map, and fine-tune the 10B-parameter Alpamayo 1.5 driving model in closed loop with Group Relative Policy Optimization (GRPO). On a progress-and-safety backbone, we add a route-corridor lane-keeping reward and two families of map-derived legality rewards: a solid-line crossing penalty and, to our knowledge, the first crosswalk-compliance rewards for closed-loop RL—a fail-to-yield term and a blocking-the-crosswalk term, both calibrated on 149 human driving logs so that human behavior incurs near-zero penalty. We ablate a trunk of increasingly informed rewards against three legality branches, sweeping the KL anchor under standard and variance-decoupled GRPO advantage estimators. On a 44-scene held-out test split, each reward family transfers the property it was designed to improve: the crosswalk-trained policy records zero RL-induced off-road events, lane-keeping rewards yield the best correct-lane occupancy and lane centering, and no fine-tuned policy ever fails to yield to a pedestrian. These gains come at a consistent price in conservatism: the zero-shot base model retains the best composite closed-loop score. We therefore analyze these trades per axis rather than through the composite alone. All scorers are validated by replaying ground-truth human trajectories through the same pipeline.

Share

COinS