Publication Date

6-29-2026

Document Type

Conference Proceeding

Publication Title

Proceedings I3d Companion 2026 ACM Symposium on Interactive 3D Graphics and Games

DOI

10.1145/3807895.3807934

Abstract

The Vision Language Models (VLMs) demonstrate remarkable performance in semantic reasoning tasks. However, it is still challenging for them to fully understand 3D spatial environment. Trained on 2D data, the model does not understand the relationship between positions and the first-person viewpoint, leading to viewpoint confusion, which is especially evident during egocentric tasks. The VLMs use 2D image Red-Green-Blue channel (RGB) pattern matching instead of actual point cloud coordinates for distance calculations. This modality could prevent the model’s 3D space geometry understanding, and leads to metric hallucination, as the model guesses the distance without mathematical calculations. In addition, the industrial environment poses a challenge for consumer-grade depth sensors and SLAM libraries. In this work, we propose a Semantic Geometric Decoupling framework to address these challenges. This framework decouples the visual reasoning from the statistical depth filter, providing the necessary deterministic link for high-performance safety and reliability. The aim is to develop a universal spatial cursor that reliably anchors VLM detections into physical 3D space. This approach enables robust real-time spatial understanding for eXtended Reality (XR) applications and opens new research directions at the intersection of VLMs and XR.

Keywords

Extended Reality, Geometric Decoupling, Object Detection

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution-Noncommercial-No Derivative Works 4.0 License.

Department

Computer Science

Share

COinS