Empowering Large Language Models with 3D Situation Awareness

Yuan, Zhihao; Peng, Yibo; Ren, Jinke; Liao, Yinghong; Han, Yatong; Feng, Chun-Mei; Zhao, Hengshuang; Li, Guanbin; Cui, Shuguang; Li, Zhen

Computer Science > Computer Vision and Pattern Recognition

arXiv:2503.23024 (cs)

[Submitted on 29 Mar 2025]

Title:Empowering Large Language Models with 3D Situation Awareness

Authors:Zhihao Yuan, Yibo Peng, Jinke Ren, Yinghong Liao, Yatong Han, Chun-Mei Feng, Hengshuang Zhao, Guanbin Li, Shuguang Cui, Zhen Li

View PDF HTML (experimental)

Abstract:Driven by the great success of Large Language Models (LLMs) in the 2D image domain, their applications in 3D scene understanding has emerged as a new trend. A key difference between 3D and 2D is that the situation of an egocentric observer in 3D scenes can change, resulting in different descriptions (e.g., ''left" or ''right"). However, current LLM-based methods overlook the egocentric perspective and simply use datasets from a global viewpoint. To address this issue, we propose a novel approach to automatically generate a situation-aware dataset by leveraging the scanning trajectory during data collection and utilizing Vision-Language Models (VLMs) to produce high-quality captions and question-answer pairs. Furthermore, we introduce a situation grounding module to explicitly predict the position and orientation of observer's viewpoint, thereby enabling LLMs to ground situation description in 3D scenes. We evaluate our approach on several benchmarks, demonstrating that our method effectively enhances the 3D situational awareness of LLMs while significantly expanding existing datasets and reducing manual effort.

Comments:	Accepted by CVPR 2025
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2503.23024 [cs.CV]
	(or arXiv:2503.23024v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2503.23024

Submission history

From: Zhihao Yuan [view email]
[v1] Sat, 29 Mar 2025 09:34:16 UTC (4,398 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Empowering Large Language Models with 3D Situation Awareness

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Empowering Large Language Models with 3D Situation Awareness

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators