Publication Date

Spring 2026

Degree Type

Thesis

Degree Name

Master of Science (MS)

Department

Computer Engineering

Advisor

Kaikai Liu; Haonan Wang; Jun Liu; Wencn Wu

Abstract

This thesis focuses on real-time task execution and object detection for autonomous robots through dataset augmentation. We propose a data augmentation approach to address dataset imbalance in Vision-Language-Action (VLA) models across both image and text modalities during the fine-tuning process. The proposed method takes an image as input and generates a structured textual description using a prompt engineering strategy to augment the textual input. The generated augmented text includes key elements such as the task goal, scene description, reasoning, and execution plan, along with other relevant contextual information. This enriched representation improves the quality of the training data and supports more effective robotic task understanding and execution during fine-tuning. The Vision-Language-Action (VLA) models, such as NVIDIA’s GR00T N1.5, rely heavily on visual inputs and often struggle with limited or poorly structured task commands. Our goal is to convert visual information into VLA-readable textual representations to improve robotic understanding and decision-making. Vision-Language Models (VLMs), such as Qwen-VL, demonstrate strong capabilities in image reasoning and prediction, which can significantly enhance a robot’s perception and motion planning. However, their inference process is relatively slow compared to the real-time response requirements of robotic systems. Traditional language models primarily operate on textual inputs, whereas many VLA models depend largely on visual inputs. This creates a gap between vision-based perception and text-based reasoning. To address this challenge, we leverage a VLM combined with prompt engineering to transform images into structured textual descriptions, making them suitable for VLA training during the fine-tuning process.

Share

COinS